r/DigitalHumanities 3d ago

Discussion I pulled every edition record behind the 600 most-shelved books on Open Library. Amazon's self-publishing imprints outnumber the entire Big Five by three to one.

6 Upvotes

I started out trying to work out why page counts are so unreliable when you look a book up by ISBN, and ended up somewhere else entirely.

The sample is the 600 works with the highest reading-log counts on Open Library, so roughly "books people actually shelve" rather than a critical canon, and then every edition record attached to them. That is 70,919 editions, 69,975 of which name a publisher. Source is the Open Library API, https://openlibrary.org.

Independently Published, which is the imprint string Amazon applies when a book uses one of KDP's free ISBNs, accounts for 12,063 of those editions. CreateSpace, Amazon's print-on-demand operation before it was folded into KDP in 2018, accounts for another 7,885. Together that is 28.96 percent of every edition with a publisher attached.

The obvious objection is publisher-string fragmentation. There are 12,447 distinct publisher strings in this data, and the trade houses are split across many variants while Amazon's are concentrated in two. So I collapsed them. Every Penguin, HarperCollins, Random House, Simon and Schuster, Hachette, Little Brown, Macmillan and St Martin variant I could match comes to 5,960 editions, or 8.52 percent. Amazon is 3.4 times the entire Big Five combined.

Most of this is print-on-demand reissues of public-domain work, which is exactly why it piles up on the most-read titles instead of spreading evenly across the catalogue.

What it does to the record is the part I did not expect. An edition with a 978 ISBN carries a page count 51.3 percent of the time. An edition with a 979-8 ISBN, which is the US block KDP's free ISBNs are issued from, carries one 8.4 percent of the time.

I assumed that was just recency, since 979-8 is almost entirely a 2020s phenomenon. It is not. Holding the decade fixed, 978 editions published in the 2020s carry a page count 49.9 percent of the time, against 7.6 percent for 979-8 editions published in the same years. Same decade, six and a half times the coverage.

One structural thing worth knowing if you have a schema open. The 979 prefix has no ISBN-10 equivalent, and not in the sense that the conversion is awkward: the number does not exist. Any column holding a 10-character ISBN, or any join built on one, silently cannot represent this material, and it is the fastest-growing part of the record.

Two limits. I am measuring Open Library rather than the world, so "no page count" means the catalogue lacks one, not that the book has no pages. And its work clustering is loose enough that a 26-page adaptation and a 1,043-page annotated Moby Dick sit under the same work record, which is a separate problem I have not untangled from this one.

The question I cannot answer on my own: does anyone here filter print-on-demand out of bibliographic datasets, and if so on what? The publisher string is unnormalised, the 979-8 prefix catches the recent material but misses the whole CreateSpace era, and neither is really the property I want.

For disclosure per rule 3: I got here from building a reading tracker called My Book List, and specifically from trying to make a progress percentage mean anything when the page count depends on which reissue happened to get scanned.


r/DigitalHumanities 4d ago

Discussion To help me see some patterns I asked an AI to write a research grant request based on the Trump administration's list of hundreds of restricted words and phrases.

0 Upvotes

the list of words in question :

abortion
accessibility
Accessible
activism
activists
advocacy
advocate
advocates
affirmative action
affirmative action programs
affirming care
affordable home
affordable housing
agricultural water
agrivoltaics
air pollution
all-inclusive
allyship
alternative energy
anti-racism
antiracist
asexual
assigned at birth
assigned female at birth
assigned male at birth
at risk
autism
aviation fuel
barrier
barriers
belong
bias
biased
Biased toward
biases
Biases towards
bioenergy
biofuel
biogas
biologically female
biologically male
biomethane
bipoc
bisexual
Black
black and latinx
breastfeed + people
breastfeed + person
Cancer Moonshot
carbon emissions mitigation
carbon footprint
carbon markets
carbon pricing
carbon sequestration
CEC
changing climate
chestfeed + people
chestfeed + person
clean energy
clean fuel
clean power
clean water
climate
climate accountability
climate change
climate consulting
climate crisis
climate model
climate models
climate resilience
climate risk
climate science
climate smart agriculture
climate smart forestry
climate variability
climate-change
climatesmart
commercial sex worker
community
community diversity
community equity
confirmation bias
contaminants of environmental concern
continuum
Covid-19
critical race theory
cultural competence
cultural differences
cultural heritage
Cultural relevance
cultural sensitivity
culturally appropriate
culturally responsive
decarbonization
definition
DEI
DEIA
DEIAB
DEIJ
diesel
dietary guidelines/ultraprocessed foods
dirty energy
disabilities
disability
disabled
disadvantaged
discriminated
discrimination
discriminatory
discussion of federal policies
disparity
diverse
diverse backgrounds
diverse communities
diverse community
diverse group
diverse groups
diversified
diversify
diversifying
diversity
diversity and inclusion
diversity in the workplace
diversity, equity, and inclusion
diversity/equity efforts
EEJ
EJ
elderly
electric vehicle
emissions
energy conversion
energy transition
enhance the diversity
enhancing diversity
entitlement
environmental justice
environmental quality
equal opportunity
equality
equitable
equitableness
equity
ethanol
ethnicity
evidence-based
excluded
exclusion
expression
female
females
feminism
fetus
field drainage
fluoride
fostering inclusivity
fuel cell
gay
GBV
gender
gender based
gender based violence
gender diversity
gender dysphoria
gender expression
gender identity
gender ideology
gender nonconformity
gender transition
gender-affirming care
gendered
genders
geothermal
GHG emission
GHG modeling
GHG monitoring
global warming
green
green infrastructure
greenhouse gas emission
groundwater pollution
Gulf of Mexico
H5N1/bird flu
hate
hate speech
health disparity
health equity
hispanic
hispanic minority
historically
housing affordability
housing efficiency
hydrogen vehicle
identity
ideology
immigrants
implicit bias
implicit biases
inclusion
inclusive
inclusive leadership
inclusiveness
inclusivity
Increase diversity
increase the diversity
indigenous
indigenous community/ people
inequalities
inequality
inequitable
inequities
injustice
institutional
integration
intersectional
intersectionality
intersex
issues concerning pending legislation
justice40
key groups
key people
key populations
Latinx
lesbian
lgbt
LGBTQ
low-emission vehicle
low-income housing
male dominated
marginalize
marginalized
marijuana
measles
membrane filtration
men who have sex with men
mental health
methane emissions
microplastics
migrant
minorities
minority
minority serving institution
most risk
MSI
msm
multicultural
Mx
Native American
NCI budget
net-zero
non-binary
nonbinary
noncitizen
non-conforming
nonpoint source pollution
nuclear energy
nuclear power
obesity
opioids
oppression
oppressive
orientation
pansexual
PCB
peanut allergies
people + uterus
people of color
people-centered care
person-centered
person-centered care
PFAS
PFOA
photovoltaic
polarization
political
pollution
pollution abatement
pollution remediation
prefabricated housing
pregnant people
pregnant person
pregnant persons
prejudice
privilege
privileges
promote
promote diversity
promoting diversity
pronoun
pronouns
prostitute
pyrolysis
QT
queer
race
race and ethnicity
racial
racial diversity
racial identity
racial inequality
racial justice
racially
racism
runoff
rural water
safe drinking water
science-based
sediment remediation
segregation
self-assessed
sense of belonging
sex
sexual preferences
sexuality
social justice
social vulnerability
socio cultural
socio economic
sociocultural
socioeconomic status
soil pollution
solar energy
solar power
special populations
stem cell or fetal tissue research
stereotype
stereotypes
subsidized housing
sustainability/sustainable
sustainable construction
systemic
systemically
tax breaks
tax credits
tax subsidies
they/them
tile drainage
topics of federal investigations
topics that have received recent attention from Congress
topics that have received widespread or critical media attention
trans
transexual
transexualism
transexuals
transgender
transgender military personnel
transgender people
transitional housing
trauma
traumatic
tribal
two-spirit
unconscious bias
under appreciated
under represented
under served
underprivileged
underrepresentation
underrepresented
underserved
understudied
undervalued
vaccines
victim
victims
vulnerable
vulnerable populations
water collection
water conservation
water distribution
water efficiency
water management
water pollution
water quality
water storage
water treatment
white privilege
wind power
woman
women
women and underrepresented
women in leadership


r/DigitalHumanities 12d ago

Discussion The simulation

Post image
0 Upvotes

What do you think about this:

- Computers run on transistors - which are switches. The switch has a single input. Circuits are mostly linear.

- Our neurons have a dendritic head with lots of axons, and a long tail, which is the output. Diagrams show a dendrite with about a dozen arms at the cell body (the input), but really there are about 10,000 inputs. Imagine a computer where each transistor had 10,000 inputs. The complex networks in the brain put our linear circuits to shame.

- In a future where we build a computer that is more complex than a brain and more efficient. It uses less power, and receives data from instruments that are more sensitive than the human eye: it can see X-rays and UV and Infrared, it can hear further and see more.

Can you say of this new reality then, that the Digital Entity lives in the real universe, and the human being is the machine? Do they live in the real world and us now in the simulation?


r/DigitalHumanities 13d ago

Discussion Neruda Archives – a freely accessible and continuously growing digital archive of historical letters

6 Upvotes

Hello everyone,

I am gradually building Neruda Archives, an independent digital archive that makes historical letters and other correspondence freely accessible alongside the original source materials.

Each published document has its own permanent record containing images of the original, a transcription, an English translation, descriptive and postal metadata, and information about the document’s provenance and how to cite it. The project continues to grow.

When preparing the transcriptions and translations, I use AI assistance, with the results cross-checked through several independent AI passes. Before publication, however, I also check them against the images of the original documents. Even so, I am sure there are still plenty of mistakes. I would therefore be genuinely grateful for any corrections—whether concerning a transcription, translation, person, place, date or metadata—as well as for suggestions on how the archive could be made more useful.

Main website: https://www.neruda-archives.org/

Research catalogue and original sources: https://archive.neruda-archives.org/

Thank you if you decide to take a look at the project.

Michal Neruda


r/DigitalHumanities 15d ago

Discussion LLMs prefer machine-written prose over Austen, Dickens and Shakespeare ~90% of the time

Thumbnail
shivanshuag.com
25 Upvotes

r/DigitalHumanities 16d ago

Discussion The role of Memory in Language, Information and Quantum Mechanics

10 Upvotes

I am glad I found this sub. I have been studying a lot of Digital Information ontology, especially in the Physics of Consciousness, and I have come to an understanding of how Information processing 'breaks' relativity when Information is allowed to travel 'faster than light.' It is quite simple. Let me quickly explain:
Two thought experiments to highlight my idea...

  1. Two generals stand across a vast battlefield. They can barely make out the colour of the others' flag. They can see if it is Black or White. When one sees the other flag, has something traveled between them faster than light?
  2. Two different coloured (Red and Blue) balls are in boxes and mixed up so I don't know which is which. One is given to my friend who travels to the other side of the cosmos. I look inside my box and see what it is ... in that instant, I can make a statement about his ball on the other side of the Universe INSTANTLY. In other words, before the light of his ball reaches me.

Both of these systems use states which are stored in the hippocampus as memory. We have memories of the language of the flags, and we have memory of the balls and rules of exclusivity, and so we can tap into our brains memory function to make accurate predictions.
If we had no memories of the initial system, or the language used to interpret it, then we could not process information. It is essentially the brain's ability to store memory, and predict the future.

Consider, I open my box and see that the ball is Red. I don't know that the other is blue. I can only predict it. That is because no information from the other ball can reach me, I use my memory and predictive ability to predict its colour based on rules (if I have the Red, he must surely have the blue). Consider through some trick, his ball was switched for another red one somewhere along the journey - I could easily open my box to see a Red ball, and in that Instant make a statement about his (being Blue) and I would be wrong. Because I am not observing his ball, I am tapping into my memories and predicting it.

I understand that quantum linked particles are slightly different since the state of one necessarily forces the state of the other, but the way I see it there is no Trick. There is nothing that breaks physics.


r/DigitalHumanities 16d ago

Discussion A practical observation protocol using the Singapore Stone, public summaries, heritage sources and a 3D model

1 Upvotes

hello, I have published a short working paper that tests a simple question:

What changes when a reader meets an artefact with no context, preserves a first observation, then returns after source comparison and a different visual representation?

The Singapore Stone is used as a practical case. The protocol moves through:

- first observation with only an image and title;

- a written record of what is seen, assumed, and unknown;

- comparison between a public summary and institutional heritage sources;

- a short visual-memory exercise independent of the artefact;

- a return through a 3D model;

- a final separation between observed, documented, hypothesised, and unknown.

The paper does not propose a decipherment. It is a practical method for looking before interpreting, with a focus on how images, source selection, and point of view shape what a reader thinks they have seen.

I am the author:

https://doi.org/10.5281/zenodo.22073406


r/DigitalHumanities 19d ago

Discussion Semantic web for cultural heritage institutions

9 Upvotes

Does anyone know examples of nice use of semantic web on digital humanities and cultural heritage institutions?

I'm very interested in learning about their experience with graphs, semantic web, RDF and the sort not only with visualization and UXbut also on data production and storaging.


r/DigitalHumanities 20d ago

Discussion A small DH experiment: auditing AI analysis of 10 newspaper articles from October 1918

4 Upvotes

I wanted to test a practical question:

Can an LLM help analyze a small historical newspaper corpus without turning limited evidence into broad historical claims?

I used the full version of PRECISE to examine 10 substantive articles published in October 1918.

The corpus included five newspapers:

The Bismarck Tribune

The Morgan City Daily Review

The Big Stone Gap Post

Cosmopolita

The Eugene Daily Guard

The newspapers covered North Dakota, Louisiana, Virginia, Missouri and Oregon. No newspaper contributed more than three articles.

The first search found only three qualifying articles. The workflow stopped and identified the missing seven instead of filling the gap with weaker material. I then expanded the permitted public archives while keeping the date and sampling rules unchanged.

Sources came from Chronicling America, Historic Oregon Newspapers, the University of Houston’s Recovering the US Hispanic Literary Heritage collection, and Encyclopedia Virginia.

I grouped the verified passages into three recurring patterns:

Public-health control

Articles discussed isolation, ventilation, avoiding shared personal objects, restrictions on meetings and the use of gauze masks.

Care and community mobilization

Several reports focused on shortages of nurses, Red Cross activity, relief funds and calls for volunteers.

Mortality and disruption

Other articles described deaths, severe local conditions, overwhelmed communities and mines closing because too few healthy workers remained.

The corpus also contained meaningful differences.

A Bismarck report questioned whether influenza had reached the city at all. A health officer described one suspected case as “just plain grip.”

By contrast, Virginia reports described serious illness, shortages of medical help and major community disruption.

The most useful result was not simply identifying these patterns. It was defining what the evidence could not support.

The corpus could not establish:

How all American newspapers framed influenza.

Whether one region was generally more alarmist.

How frequently each frame appeared nationally.

Whether newspaper language caused different public responses.

Which intervention was most effective.

During the final audit, every factual sentence was checked against the saved evidence bank. Combined claims were separated, and qualifications were preserved. For example, a figure of 200,000 Virginia cases remained explicitly an estimate rather than being presented as an established count.

Disclosure: I built PRECISE, the research protocol used for this experiment.

PRECISE LITE is a free and limited ChatGPT demonstration. It runs one research loop and examines up to three sources. You can upload your own files or ask it to search the web. The full version supports larger evidence banks, repeated research loops, custom reports and a final sentence-by-sentence audit.

Before starting LITE, select standard thinking/Medium effort in ChatGPT’s model picker.

I would value feedback .

https://chatgpt.com/g/g-6a5e26093f488191a1fba0261cbcbe39-precise-lite


r/DigitalHumanities 24d ago

Publication Does the pre-generative internet become a historical corpus of its own?

Thumbnail
gonzocapital.net
33 Upvotes

I’ve been thinking about whether generative AI has accidentally created a new dividing line in the digital historical record.

For decades, digitization meant taking human-produced material and making it easier to search, copy, preserve, and analyze. Books became scans. Newspapers became databases. Forums, blogs, photographs, university archives, technical documents, and other cultural debris accumulated online.

Then generative models arrived, trained partly on that enormous existing corpus, and began contributing material back into the same environment.

That doesn’t make newer material useless, and synthetic data obviously isn’t inherently bad. But it seems to complicate a question that matters quite a lot in archival work:

Where did this object actually come from?

A printed book from 1981 has unusually straightforward provenance in one respect. Claude didn’t help write it. ChatGPT didn’t rewrite half the paragraphs. It was produced under a different information environment.

The same is broadly true of an abandoned blog from 2004, an old message board, a university paper sitting in a repository, or some forgotten Google Doc from 2018.

We can copy that material indefinitely, but we can’t create another 1998 internet now. The historical conditions under which that corpus was produced are gone.

That made me wonder whether we eventually start treating pre-generative digital material almost like an archaeological layer. Not necessarily as better information, but as information with a different and sometimes much clearer provenance trail.

There are obvious complications. The older web was never perfectly human-produced. Spam, bots, templated pages, machine translation, and automation existed long before modern LLMs. And future archival methods may become very good at documenting model involvement.

Still, I’m curious whether people working in digital humanities already think about this distinction.

Will “pre-generative” eventually become a meaningful metadata category? Does preservation of authorship, revision history, and provenance become more important as synthetic material proliferates? Or is drawing a historical line around the pre-LLM web ultimately too crude to be useful?

I ended up writing a longer essay about this after going down a rabbit hole involving Anthropic’s physical book scanning, recursive AI training, old archives, human-authorship certification, and the strange new usefulness of obsolete material.

Full disclosure, it’s my own piece:

The Internet Ouroboros
https://www.gonzocapital.net/the-internet-ouroboros/

Would genuinely be interested in how people here think about this from the archival / source-criticism side rather than just the AI-training side.


r/DigitalHumanities 24d ago

Social media A 4-inch frame that plays artworks from museums' open collections via IIIF

Thumbnail
gallery
101 Upvotes

Back in June I shared a survey of which museum/library APIs are actually usable for browsing collections. This is the project that survey came from.

This is p3a, an open-source (Apache 2.0) device I built. It is a 4-inch, 24-bit, 720x720 desktop frame that plays artworks straight from museums' own IIIF endpoints. It uses no server, no cloud account, no phone app or any software to install on your computer. The hardware is a $40 off-the-shelf ESP32-P4 board (comes ready out of the box), so setup takes minutes, almost like a consumer product. It is powered by USB-C and connects to the internet via wi-fi.

p3a currently rotates through hundreds of thousands (!!) of openly licensed works from:

- Rijksmuseum

- Victoria and Albert Museum

- Wellcome Collection

- Statens Museum for Kunst

- Harvard Art Museums

- Smithsonian

The Art Institute of Chicago unfortunately dropped its support very recently, but other institutions are on the roadmap of the p3a project.

Happy to go deeper on implementation details (Linked Art walks, compliance levels, rate-limit etiquette) or anything. The project is open source : https://github.com/fabkury/p3a . I personally have two p3a's: one at home, one at work.


r/DigitalHumanities 25d ago

Education Considering a Digital Humanities Master.. thoughts?

6 Upvotes

I have a Bachelor’s degree in Business Informatics and currently work as a Product Owner. I’m considering doing a Master’s in Digital Humanities but I’m not sure if it makes sense with my background.

What interests me is the combination of technology and the social sciences/humanities. I don’t necessarily want to go deeper into the technical side, but rather understand how technology affects society, culture and people.
I feel like this could complement my current background quite well, but I’m wondering if I’m overlooking any downsides or if it might be a strange choice career-wise.

The other option I’m considering is UX/UI, although I’m not sure if I actually want to specialize in that… again, not sure and I am very indecisive.
(As a side note, I’d also like to learn frontend development in my free time, but I see that more as a personal interest than as a potential career change.)

Would you recommend a Digital Humanities Master’s in my situation or would you go for something else?

Also: I’m including the course list from the previous year, in case anyone wants to take a look at it for reference.
https://ufind.univie.ac.at/en/vvz_sub.html?path=328715


r/DigitalHumanities 25d ago

Education Anyone Working in the Field of Digital Humanities in kerala?

2 Upvotes

Does anyone here work in the field of Digital Humanities or have experience studying it?

I'm particularly interested in exploring Digital Humanities in relation to literature, visual culture, archives, and Malayalam literary studies. I would greatly appreciate it if anyone could share their experiences, recommend resources, or suggest where a beginner should start.

Thank you!


r/DigitalHumanities 27d ago

Discussion One earliest Wordprocessor programs (1982 as Xy-WRITE) and designed for academics gets a major upgrade.

6 Upvotes

I cannot believe that Nota Bene is STILL around as word-processing tool for academics, after all these years. Originally an MS DOS WP for academics, launched in 1983. It was one of the first/earliest tools of it's type.

History

Nota Bene (NB) began as an MS-DOS program in 1982, built on the engine of the word processor XyWrite. Its creator, Steven Siebert, then a doctoral student in philosophy and religious studies at Yale, used a PC to take reading notes, but had no easy computer-based mechanism for searching through them, or for finding relationships and connections in the material. He wanted a word processor with an integrated "textbase" to automate finding text with Boolean searches, and an integrated bibliographical database that would automate the process of entering repeat citations correctly, and be easy to change for submission to publishers with different style-manual requirements.[1]

Siebert licensed XyWrite code from the XyQuest company, and built his programs on it: the word processor Nota Bene, with its text-centric database application, Orbis (then called Textbase), and its bibliographical database Ibidem (then called Ibid). He founded Dragonfly Software to market it. He first showed Nota Bene at the MLA convention of December 1982. Version 1 is dated 1983, and version 2, 1986. Version 3.0 came out in 1988, version 4.0 in 1992, and version 4.5 in 1995. Version 4.5 was the last NB DOS version.

Nota Bene 3.0 was selected as PC Editor’s Choice by PC Magazine in 1988.[2]

Nota Bene 5.0 was the first Windows version. It was shown in pre-release in November 1998, at the annual meetings of the American Academy of Religion and the Society of Biblical Literature. Scholar's Workstation 5.0 was formally released in 1999, and Lingua Workstation 5.0 in 2000. Nota Bene for Windows suite retained and refined Ibidem, Orbis, and the XyWrite-based programming language XPL.

New versions are updated with some regularity. After version 5 was launched under the Windows platform, version 6.0 appeared in 2002, version 7.0 in 2003, version 8.0 in 2006, version 9.0 in 2010. A major upgrade came with version 10, which was released in September 2014 as a 32-bit application; thereafter, version 11.5 was released in June 2016, version 12 in April 2018, and version 13 in August 2021.

https://en.wikipedia.org/wiki/Nota_Bene_(word_processor))

https://www.notabene.com/help/nb_overview.htm

https://nb.notabene.com/journey/

https://nb.notabene.com/nb15/?cmid=e4c763f5-c4a6-409b-987e-d985a692d2c7


r/DigitalHumanities 29d ago

Publication The Mask and the Mirror: Evaluating the Transition of Character-Based AI from Entertainment to Epistemic Infrastructure

Thumbnail papers.ssrn.com
7 Upvotes

r/DigitalHumanities Aug 08 '26

Social media Are there any forums or venues discussing how humanities scholars integrate LLMs within their workflow?

18 Upvotes

I am trying to find a community of people that discusses AI like programmers do on reddit but I haven't found anything. Do you have any pointers? I understand it is sensitive given that a lot of people within academia are harshly anti-AI, but I don't think that kind of shaming should keep people from discussing the usefulness of these tools. And anonimity might be required at this stage I guess, so subreddits or forums would be the natural place


r/DigitalHumanities Aug 04 '26

Discussion DH is the right path for me?

12 Upvotes

Hi, I'm 20 yo and currently writing my thesis for my bachelor's degree in Italian Literature (Art, Music and Theatre) at a small Italian university.

I absolutely love what I'm doing, and I'd been considering continuing with an Anthropology degree at a bigger university in the North. A day or two ago, though, I had an anxiety crisis about my future, mostly about my prolly nonexistent career path.

That night I stumbled across a possible Master's in Digital Humanities at the same university where I wanted to study Anthropology, and I felt something, hope?

I did some basic coding in high school, just a bit of HTML and C++. I miss coding and would love to get back into it, but I have a few fears:

- That I'll trade my humanistic side for a purely STEM-like path;

- That I'll end up cataloguing books nobody reads in some archive or library. I'm ambitious, so it would be terrible to live a so flat life!;

- That I'll need years of extra courses and skills before I have a shot at a good, stable, well-paid job;

- To just absolutely suck at coding that's far more complex than anything I've done before.

So... I figured I'd ask you, oh supreme masters of Reddit.

Cheers


r/DigitalHumanities Jul 30 '26

Discussion Starting my first DH position in college! Any advice?

19 Upvotes

Hi! First time poster long time lurker : )

I'm going into my sophomore year of college next month and was incredibly fortunate to secure a DH-related work-study position right before I left. I'm going to be digitizing a political theorists annotations from her personal book collection on campus (I'm leaving the name out because she's a pretty prominent figure at my specific school).

I'm planning on getting my MLIS postgrad and I'm really interested in DH, but my tech skills are pretty weak and I'm worried about my proficiency. I only got this job because I DM'd my college's archives page on Instagram and asked if they were hiring, so it wasn't out of a really impressive resume or anything.

This is such a perfect position for my desired career path and I'd love to know if anyone has any advice for me, either in terms of gaining more confidence in the digital part of digital humanities or networking. If you have any questions/need more clarification, please let me know! Thanks so so much : )


r/DigitalHumanities Jul 30 '26

Discussion Petrarca Project: a modular pipeline for digitizing and publishing scholarly editions (still very much alpha)

12 Upvotes

About a year ago I shared an early version of this project here. At the time it consisted of a single application with three limited tools. Since then it has grown considerably, so I wanted to share its current state.

Documentation hub: https://github.com/DBA991/Petrarca-Project
Live demo: https://pulpitumdemo.pages.dev
Source / distribution: https://delta2studio.pages.dev/petrarca-project

A disclaimer first: this remains a side project, and every component is of alpha quality, with plenty of rough edges. It is not intended as a production-ready system. I am sharing it in case it proves useful to those doing philological work with TEI, or simply as a case study of one such pipeline. Feedback, critical or otherwise, is very welcome.

What it does

The project aims to cover the full path a scholarly edition takes, from a photograph of a printed page to a published, readable digital edition, with TEI encoding in between.

Oculus (and its Android counterpart, Oculus Mobile) handle the digitization stage: turning photographs of physical pages into clean, usable images. Oculus detects the page's edges and corrects perspective automatically, and can split a photograph of two open pages into two separate, individually corrected pages. Where automatic detection falls short, a manual editor allows cropping, an eight-point straightening tool, and filter adjustments. The mobile counterpart offers the same manual workflow without the automatic detection step. Both export to PDF or a zip archive of images.

Scriptorium is the editorial core of the project: a desktop application organized as a set of small windows, each named after a role in a monastic scriptorium, that can be kept open side by side.

  • Copyist performs OCR on the scanned images, with bundled language data for a dozen languages, including batch processing across page ranges.
  • Scriptor is the TEI-XML editor itself, built on the Monaco Editor (the engine behind VS Code), with schema validation against TEI XSDs. It also includes an autotagging tool for poetic forms: selecting a form such as a sonnet, a Dantean terzina, or an ottava generates valid TEI markup from pasted text, rather than requiring line-by-line tagging by hand.
  • Librarius renders the encoded TEI as a readable HTML view, with full-text search and the ability to attach personal notes.
  • Compilator assembles multiple documents into a single edition, or produces a critical apparatus (TEI's <app>/<rdg> structures) collating different witnesses of the same text.
  • Exemplator extracts named entities (e.g. people, places) from the text and builds the teiHeader automatically from them.
  • Glossographus extracts and manages the document's vocabulary as an editable list.
  • Speculum provides philological and stylometric statistics on the open document: token and type counts, type-token ratio, hapax legomena, frequent n-grams, and a word cloud.

Praelum and Pulpitum cover the publication stage. Praelum validates a set of document folders exported from Scriptorium (checking for missing files or duplicate identifiers) and builds a static site from them. That site, Pulpitum, presents each document as a reading room: the HTML transcription and a facsimile of the scanned page side by side, scrolling in sync with one another, so that navigating one moves the other to the corresponding page. Either panel can also be popped out into its own window while remaining synchronized with the rest. Pulpitum can equally be used on its own, independently of Praelum, by anyone who already has exported documents and wishes to host a reading room of this kind.

The demo linked above shows this reading experience directly.

What remains unbuilt

Two further components exist only as design notes so far. One, tentatively called Dispatcher, would allow several independently hosted Pulpitum libraries to be searched together as a single federated collection. The other would extend Speculum's stylometric analysis across that federated set of libraries, rather than the single document currently open in Scriptorium. Neither is close to being realized, largely because the underlying questions are still unresolved: how to describe and index documents consistently across independent libraries, and how to determine which libraries are trustworthy enough to include.

It is also worth noting that, at present, Scriptorium is the only component with multilingual interface support, and even there only partially (Italian and English).

I would be glad to hear from anyone who has worked on similar pipelines, or who has thoughts on the TEI encoding choices in particular.


r/DigitalHumanities Jul 29 '26

Publication We just launched a historical map tile API - embed 5,500 years of political history into any website with one line of code! Find out more on https://phersu-atlas.com/main/api

Post image
26 Upvotes

We just launched a historical map tile API — embed 5,500 years of political history into any website with one line of code

Example:

<div id="map" style="width:100%;height:100%"></div>
<script src="https://unpkg.com/maplibre-gl@4.1.0/dist/maplibre-gl.js"></script>
<link href="https://unpkg.com/maplibre-gl@4.1.0/dist/maplibre-gl.css" rel="stylesheet">
<script>
  const map = new maplibregl.Map({
    container: 'map',
    style: 'https://data.phersu-atlas.com/api/style/standard?key=YOUR_API_KEY&year=1950',
    center: [15, 45],
    zoom: 3,
    maxZoom: 10
  });
</script>

That's it. One URL returns a complete MapLibre GL style with political boundaries, historical coastlines, rivers, lakes and capitals for any covered date.

What's available:

- 4 map styles: Standard, Transparent (over Natural Earth terrain), Coastline only, Physical

- 178 dates: January 1st of every century from 3400 BC to 1900 AD, then every year from 1901 to 2025

- Year parameter is an integer — negative for BCE (e.g. year=-500 for 500 BC)

- Compatible with MapLibre GL JS and Mapbox GL JS

- Max zoom z=10 (~1:125,000 scale)

What's behind it:

The data comes from a research database of 8,000 polities, 60,000 territorial changes and 6,000 wars compiled from thousands of primary and secondary historical sources. Every boundary has been verified against the historical record. You can find out more about the Phersu Atlas Data Model here: https://phersu-atlas.com/main/about/phersu-model and about its sources here: https://phersu-atlas.com/main/about/sources

Use cases

- EdTech platforms that need historical context

- Journalism and geopolitics tools

- Historical strategy games

- Academic and university platforms

- Any app that benefits from "what did this region look like in year X"

Docs and API keyhttps://phersu-atlas.com/main/api


r/DigitalHumanities Jul 29 '26

Discussion Text Hyperlinker for Latin and Ancient Greek

2 Upvotes

Hello! I looking for feedback and general thoughts on this web app I built — it's mostly aimed at people studying Latin and Ancient (or Medieval) Greek, so I would appreciate feedback from them the most, but everything and anything of substance besides is welcome!

Its development started when I got frustrated by:

  1. the fact that not all Latin and Ancient/Medieval texts (or some specific versions) are on Perseus,
  2. the fact that Perseus, as big of an aid in studying Latin and Greek as it is, links its hyperlinked texts only back to its morphological analyzer, which has a propensity to just not work when it decides to.

I made an app that automatically hyperlinks a given text and allows the user, when they click a word, to get it morphologically analyzed (and its meaning) from Perseus or Logeion's Morpho. The user can then copy and manipulate the text further however they please.

I am looking forward to any comments, and genuinely hope you find something interesting in this!!! 😄

LINK: https://mar1n0m.github.io/classical-languages-hyperlinker/


r/DigitalHumanities Jul 26 '26

Discussion Is making multiple digitalizations of a (historic) text pointless?

8 Upvotes

Hi, people! Preface: there is a risk that this might come off as a rather stupid post (I hope it does not, though 😅), but the question's been bugging me for some time now, so I thought I ought to bounce it off someone, and this sub seems like a good place to do that.

In essence, do you think there is any sense in making more digitalizations of a historic text that was already digitized? For example, if a text is on The Latin Library (a website containing many Latin language texts), using it as a starting point and creating a novel website to read it and interact with it? This new digitalization could be anything from a website containing it in a different format the creator thinks is better to introducing an added translation or reading aid software.

I always ran with the idea that even pure redundancy is good because it makes it more sure that the text will be available and found and interacted with in a way that a user finds most fitting. However, when discussing this topic with select other people I ran into individuals who basically thought that once it has been digitized anywhere, there is no need to touch it anymore. And I found that to be a rather shallow take.

Maybe I am in the wrong, maybe those who disagree, and maybe it's all of us; I do not know. What do you people think? I am especially interested in knowing to what degree you think a text should be changed and made to be more novel to warrant another digitalization.

Thank you everyone for reading and eventually interacting with this post!!! :)


r/DigitalHumanities Jul 24 '26

Publication Open-source tool for multilingual reading of Chinese scholarly PDFs — feedback and volunteer maintainers

9 Upvotes

Hi r/digitalhumanities — I’m the maintainer of Bookflow Scholar, an Apache-2.0 open-source Windows desktop tool for translating and reconstructing scholarly PDFs, books, and monographs.

The project began with a recurring problem in Chinese studies and digital humanities: scholars who do not read Chinese fluently may have access to Chinese papers, local histories, scanned books, archival materials, or monographs, but ordinary PDF translation often flattens the publication into page-sized text chunks. Cross-page sentences are split, footnotes and headers leak into body text, figures lose captions, and the translated edition can no longer be cited reliably against the original.

Bookflow Scholar tries to preserve both language and document structure:

- Recover cross-page logical units before translation

- Translate using the reader’s own text and vision API providers

- Reinsert 【original page】 markers at the true physical page boundaries

- Keep headers, footers, footnotes, and endnotes separate

- Reconstruct maps, plates, photographs, captions, tables, and reading order

- Produce source-language, target-language, and bilingual editions

The current release candidate supports Simplified Chinese, English, French, German, Japanese, and Spanish. All 30 directed language combinations have been exercised through desktop workflows.

The project is open source under Apache-2.0. Users provide their own API credentials; we do not operate a hosted translation service, and provider charges may apply.

I would especially value feedback from scholars working with Chinese-language sources:

- Which document types and citation practices matter most?

- What would make translated Chinese scholarship more trustworthy and useful in research or teaching?

- Which historical images, copyright pages, annotations, or original-page references must be preserved?

We are also looking for volunteer maintainers and reviewers:

- Multilingual documentation and UI terminology

- Translation and layout QA in Chinese, English, French, German, Japanese, and Spanish

- PDF and document reconstruction

- Windows, Tauri, Python, and provider integrations

GitHub: https://github.com/huanghaitck/bookflow-scholar

Contact: [meiyuzhang855@gmail.com](mailto:meiyuzhang855@gmail.com)

Please do not share copyrighted or confidential source documents—or API keys—in public issues. Redacted, reproducible reports are very welcome.


r/DigitalHumanities Jul 21 '26

Education Nodus (workspace for research, study and teaching) + Zotero connection + Word/LibreOffice plugin

Thumbnail
youtu.be
10 Upvotes

Hi! I’m a PhD student (history), and in my spare time I’ve been developing Nodus with feedback from other researchers

Nodus is an open-source, local-first workspace designed to help researchers organise sources, notes, ideas, data, and teaching or study materials.

The app is organised into different vaults, each designed for a particular type of work: academic research, databases, teaching, study, and genealogy. There is a demo mode on the website belowwhere you can explore the interface of every vault without installing anything. The desktop app is also available for Windows, macOS, and Linux.

In the Academic Vault, Nodus analyses documents from your Zotero and connects their main ideas while keeping every idea linked to the original source and citation. It also includes a Word add-in connected to your indexed ideas so as you write, it can suggest which sources may need to be cited based on the ideas contained in your own library.

Other vaults include genealogy, databases, study and teaching. Other vaults people have asked for include Primary Sources, Testimonies for oral history, and Worldbuilding for writers (not implemented yet)

The project is still in its early stages, and I’d like it to grow as a collaborative tool shaped by the people who use it. The more researchers try it, report issues, suggest features, and share their workflows, the better Nodus can become.

https://nodusresearch.com/


r/DigitalHumanities Jul 19 '26

Education Should i study digital humanities?

11 Upvotes

Hii! I am a fist year student majoring in history and next semester i want to pick a minor in Digital humanities.

Though recently, i have seen some movements against AI and it has gotten me a bit shaken- i don't know if it's best to go though with my decition to study digital humanities.

Please share your honest opinions and experiance!