Transkribus Datasets
Build and Share Research Datasets
for Historical Documents
Datasets is where transcribed and annotated pages in Transkribus become training data for text recognition, layout analysis and information extraction. Curate pages from any collection, structure them into splits, freeze versions you can point to, and publish them for other researchers to build on.
A Dataset Is a Curated Set of Pages, Ready to Train On
Collections hold everything you have transcribed. A dataset holds the pages you have chosen for a purpose: one script, one period, one document type, checked and structured so a model can learn from it.
Curated
Pull pages from any collection you have access to: a whole collection, one document, or hand-picked pages. Duplicates are flagged before they are added, and items can move between datasets as your selection sharpens.
Structured
Assign pages to training, validation and test splits with adjustable ratios, by page or in bulk. Open any page to inspect the image and transcription before a model learns from it.
Versioned
Freeze a version to pin its exact contents, then keep editing the draft. Every addition, removal, split and freeze is recorded with who did it and when.
Built for research groups and archives that train their own models. The data stays yours, every step is recorded, and publishing is a decision you make per dataset, on the terms you choose.
Datasets Published by the Community
Ground truth other researchers have transcribed, checked and published. Find training data for your script, language or period, and read every record without an account.
Dataset_Angeliki
Pages 1-34 of manuscript Athens EBE 263. Medieval Greek handwriting. Created for the Paleography - HTR assignment.
DERMANIS_ACHILLEAS
Konstantina Kyriakou
This dataset is for the lesson "Digital Technologies in Papyri and Manuscripts" of the MSc in Digital Methods for the Humanities of the Athens University of Economics and Business
Sample Across Collections, Reproducibly
Picking training pages by hand stops scaling at a few hundred. The sampling wizard draws a random or stratified sample across several collections at once, keeps only pages with ground truth, and skips pages already in the dataset. You see the result before anything is written, and the seed is stored so the same sample can be reproduced.
- Chancery records 1750-18002,140 with GT
- Parish registers, Tyrol1,380 with GT
- Private letters, 19th c.480 with GT
- Only pages with ground truth
- Exclude pages already in this dataset
- Chancery records 1750-1800214
- Parish registers, Tyrol138
- Private letters, 19th c.48
Nothing is written until you confirm. The seed is stored with the dataset, so this exact sample can be drawn again.
Structure Your Data the Right Way
Good training data is a matter of structure as much as size: pages the model learns from, pages that keep it honest during training, and pages you hold back to check the finished model on pages it has never seen. Set the ratios once and apply them across the whole set.
Work Methodically. Version Your Training Data.
Freeze a version before training and keep working on the next. Every version captures exactly which pages and transcriptions were included, and publishing freezes a version automatically, so later edits can never quietly change what you published or what a model was trained on. The change log records every step, with who did it and when.
You Decide How Far Your Data Travels
Publishing is not a single switch. Pick a visibility level per dataset, change it whenever you want, and invite individual colleagues as viewers or editors on top.
Private
Nothing leaves your workspace. This is where every dataset starts.
Everyone else can:
- No: See the pages and transcriptions
- No: Copy the data into their own dataset
- No: Train their own model on it
You and the people you invite keep full access.
View only
Show your work without handing it over.
Everyone else can:
- Yes: See the pages and transcriptions
- No: Copy the data into their own dataset
- No: Train their own model on it
Enforced, not requested: the copy and train actions are refused.
Public and reusable
Published for the community to build on, under the licence you choose.
Everyone else can:
- Yes: See the pages and transcriptions
- Yes: Copy the data into their own dataset
- Yes: Train their own model on it
Your name, your licence, your terms stay attached to the record.
The same three levels apply to the ground truth behind a published model, so sharing a model does not have to mean giving away the data it learned from. Visibility decides whether; the licence you attach decides on what terms: from CC0 and CC BY through to AI-specific licences and your own custom wording.
Built for the Way You Work
What a finished dataset is for, and who you can bring along.
Curate from Anywhere
Pull the best transcribed pages from any collection into one focused dataset, by hand or by sampling across collections with a saved seed.
Never Lose Track
Freeze versions for reproducibility. Always know exactly what went into every training run.
Train in One Step
Go straight from a finished dataset to model training: text, layout, field and table recognition. The model records the exact dataset version it used.
Build on Published Data
Copy a public dataset into your own workspace, take just one split, or train on it where it sits. No new transcription needed.
License It Your Way
CC0, CC BY, ShareAlike, NonCommercial, licences written for AI training, or your own terms, shown on the public record.
Work as a Team
Invite people as viewers or editors, change or withdraw access later, and hand a dataset over to a colleague when you move on.
Cite a Dataset Like Any Other Publication
Every published dataset already has a stable public record: its own URL, licence, page count, version history and the authors named when it was published. We are preparing DOI registration on top of that record, so that a specific version of the data behind a model can be cited in a paper, listed in a data availability statement, and credited to the people who transcribed it.
Until DOIs are switched on, the record page is the reference to use.
Aigner, M., & Rainer, L. (2026). Chancery records, Tyrol 1750-1800 (Version 2) [Data set]. Transkribus. https://doi.org/10.…DOI to follow
Start With the Pages You Already Have
Curate a dataset from your existing transcriptions, or start from one another researcher has published, and train a model on data you can point to.