Info: Call for Graduate Fellows!

Applications for the CDH Graduate Fellowship are open through Oct 20, 2026. Apply now.

Transcribe Lab

Built by CDH

Open, scholar-shaped AI infrastructure for transforming handwritten and complex historical sources into machine-readable text, and for learning to use these methods critically.

Automated Text Recognition
Data Curation
Data Development
Digital Research Infrastructures
Handwritten Text Recognition (HTR)
Software Development
transcribe lab placeholder2

Code

Code

Rebecca Sutton Koeser and Christine Roughan, Htr2hpc, Python, v. 0.5, August 22, 2024; Center for Digital Humanities at Princeton, released October 21, 2025, https://github.com/Princeton-CDH/htr2hpc.

Publications and Presentations

Publications and Presentations

Christine Roughan and Rebecca Sutton Koeser, “Integrating ATR Software with University HPC Infrastructure: Balancing Diverse Compute Needs,” paper presented at US Research Software Engineering Conference 2025 (USRSE’25), USRSE25 Conference Proceedings, October 1, 2025, https://doi.org/10.5281/zenodo.17245595.

Christine Roughan, “Integrating ATR Software with University HPC Infrastructure,” SCOOP: Source Codes of the Past: Launching an international ATR/HTR Network for Manuscript Analysis, Princeton, June 13, 2025, https://cdh.princeton.edu/events/2025/06/scoop-source-codes-of-the-past-launching-an-international-atrhtr-network-for-manuscript-analysis

Research challenge
One of the great promises of the digital turn was that humanists would be able to work with their sources at scale. Mass digitization has vastly expanded access to archival and manuscript collections, but access to page images is not the same as access to text. For handwritten documents, early print, non-Latin scripts, and damaged or degraded materials, automated text recognition remains unreliable or unavailable. Much of the historical record therefore stays beyond the reach of computational inquiry. Where AI tools do exist, they are increasingly concentrated in proprietary platforms. These platforms were not designed around humanistic research questions, and they offer scholars little insight into, or control over, how their data is processed and their results are produced.

An infrastructure for scholars, shaped by scholars
The Transcribe Lab offers an alternative: open, transparent infrastructure for developing, training, and deploying AI models suited to the distinctive challenges of humanities sources. Its pipelines and workflows are deliberately designed around scholarly practice. Researchers retain ownership of their data, models, and transcriptions, and every stage of the process remains open to inspection. This commitment is reflected in our choice of tools. The lab is built on eScriptorium, an open-source platform developed within the academic community, rather than on proprietary, subscription-based alternatives. This choice preserves scholars' methodological independence and keeps the resulting models and datasets available as shared, reusable research assets rather than locked into a commercial service.

Learning with AI, not just using it
The lab's aim is not only to produce transcriptions but to build critical AI literacy. Through hands-on work with their own materials, faculty and graduate students learn how these models are trained, where they succeed and fail, and how recognition errors and training biases can shape downstream analysis. They also learn how to document and evaluate machine-generated text with the same rigor they bring to any other source. Scholars leave the lab not just with results, but with the judgment to know when and how to trust them.

A virtuous cycle
As researchers transcribe, annotate, and correct model output on their collections, they generate the training data needed to improve recognition for underrepresented scripts and complex materials. Each project strengthens the next. Over time, the lab builds a growing library of models and datasets tuned to humanities research at Princeton, lowering the barrier to entry for scholars with less technical experience while deepening expertise across the community.

Initial phase (2026–2028)
In its first phase, the lab will deploy eScriptorium as a production-grade platform for training data creation and handwritten text recognition, and will evaluate the computational requirements of humanities AI work on Princeton's high-performance computing cluster. In collaboration with Princeton's Data Driven Social Science, we will also develop support for vision-language models, which interpret images and text together and open new possibilities for complex page layouts and mixed materials.

Collaborative origin
The Transcribe Lab is part of the broader Princeton Open HTR (Handwritten Text Recognition) Initiative and builds on work completed in the HTR2HPC research partnership. Led by CDH staff in collaboration with a core group of faculty from History, Near Eastern Studies, English, Classics, and East Asian Studies, the initiative will catalyze research innovation and foster community among researchers working at the intersection of the humanities and computation. It will also position Princeton as a contributor to the growing international network of open, community-governed infrastructure for computational humanities.

Associated Projects

Bringing HTR to the HPC

Customizing the eScriptorium HTR software for use on Princeton high performance computing hardware

Built by CDH
HTR2HPC cover

Princeton Open HTR Initiative

Establishing research infrastructures to support Princeton use of HTR for manuscripts and archival documents in a variety of languages and scripts

eScriptorium Syriac cover

Team

Co-PI: Research Lead

Research Software Engineer

Faculty Advisor

Project Advisor

Technical Advisor

Grants

2026–

Research Partnership