Skip to content
  • Pricing

Transkribus Datasets

Build and Share Research Datasets
for Historical Documents

Datasets is where transcribed and annotated pages in Transkribus become training data for text recognition, layout analysis and information extraction. Curate pages from any collection, structure them into splits, freeze versions you can point to, and publish them for other researchers to build on.

A Dataset Is a Curated Set of Pages, Ready to Train On

Collections hold everything you have transcribed. A dataset holds the pages you have chosen for a purpose: one script, one period, one document type, checked and structured so a model can learn from it.

Curated

Pull pages from any collection you have access to: a whole collection, one document, or hand-picked pages. Duplicates are flagged before they are added, and items can move between datasets as your selection sharpens.

Structured

Assign pages to training, validation and test splits with adjustable ratios, by page or in bulk. Open any page to inspect the image and transcription before a model learns from it.

Versioned

Freeze a version to pin its exact contents, then keep editing the draft. Every addition, removal, split and freeze is recorded with who did it and when.

Built for research groups and archives that train their own models. The data stays yours, every step is recorded, and publishing is a decision you make per dataset, on the terms you choose.

Sample Across Collections, Reproducibly

Picking training pages by hand stops scaling at a few hundred. The sampling wizard draws a random or stratified sample across several collections at once, keeps only pages with ground truth, and skips pages already in the dataset. You see the result before anything is written, and the seed is stored so the same sample can be reproduced.

Sample pagesStep 1 of 2
Source collections
  • Chancery records 1750-18002,140 with GT
  • Parish registers, Tyrol1,380 with GT
  • Private letters, 19th c.480 with GT
Method
RandomStratified
Sample size400 pages
  • Only pages with ground truth
  • Exclude pages already in this dataset
Seedk7Qm2VpXa9Rc
PreviewStep 2 of 2
400pages will be added, drawn in proportion to each collection
  • Chancery records 1750-1800214
  • Parish registers, Tyrol138
  • Private letters, 19th c.48

Nothing is written until you confirm. The seed is stored with the dataset, so this exact sample can be drawn again.

BackAdd 400 pages

Structure Your Data the Right Way

Good training data is a matter of structure as much as size: pages the model learns from, pages that keep it honest during training, and pages you hold back to check the finished model on pages it has never seen. Set the ratios once and apply them across the whole set.

Validation
90%
10%
Train
Validation

Work Methodically. Version Your Training Data.

Freeze a version before training and keep working on the next. Every version captures exactly which pages and transcriptions were included, and publishing freezes a version automatically, so later edits can never quietly change what you published or what a model was trained on. The change log records every step, with who did it and when.

Kurrent Kanzler v3
Kurrent Kanzler v3
420 items · 16 Mar 2026
Add note…
Kurrent Kanzler v2
380 items · 10 Mar 2026
Frozen for Kurrent model
Kurrent Kanzler v1
200 items · 2 Mar 2026
Initial collection
Changes
Split Applied
by learn@transkribus.org
23 minutes ago
Seed:Y0kiKtVHmqtg
Method:random
Ratios:train: 0.8, val: 0.1, test: 0.1
Items Added
by learn@transkribus.org
23 minutes ago
Count:8
Item Ids:8 itemsShow
Version Created
by learn@transkribus.org
25 minutes ago
Name:Kurrent Kanzler
Dataset Type:Standalone

You Decide How Far Your Data Travels

Publishing is not a single switch. Pick a visibility level per dataset, change it whenever you want, and invite individual colleagues as viewers or editors on top.

Private

Nothing leaves your workspace. This is where every dataset starts.

Everyone else can:

  • No: See the pages and transcriptions
  • No: Copy the data into their own dataset
  • No: Train their own model on it

You and the people you invite keep full access.

View only

Show your work without handing it over.

Everyone else can:

  • Yes: See the pages and transcriptions
  • No: Copy the data into their own dataset
  • No: Train their own model on it

Enforced, not requested: the copy and train actions are refused.

Public and reusable

Published for the community to build on, under the licence you choose.

Everyone else can:

  • Yes: See the pages and transcriptions
  • Yes: Copy the data into their own dataset
  • Yes: Train their own model on it

Your name, your licence, your terms stay attached to the record.

The same three levels apply to the ground truth behind a published model, so sharing a model does not have to mean giving away the data it learned from. Visibility decides whether; the licence you attach decides on what terms: from CC0 and CC BY through to AI-specific licences and your own custom wording.

Built for the Way You Work

What a finished dataset is for, and who you can bring along.

Curate from Anywhere

Pull the best transcribed pages from any collection into one focused dataset, by hand or by sampling across collections with a saved seed.

Never Lose Track

Freeze versions for reproducibility. Always know exactly what went into every training run.

v1
v2
v3

Train in One Step

Go straight from a finished dataset to model training: text, layout, field and table recognition. The model records the exact dataset version it used.

Dataset
AI Model

Build on Published Data

Copy a public dataset into your own workspace, take just one split, or train on it where it sits. No new transcription needed.

License It Your Way

CC0, CC BY, ShareAlike, NonCommercial, licences written for AI training, or your own terms, shown on the public record.

CC0CC BYYour terms

Work as a Team

Invite people as viewers or editors, change or withdraw access later, and hand a dataset over to a colleague when you move on.

+2
Coming next

Cite a Dataset Like Any Other Publication

Every published dataset already has a stable public record: its own URL, licence, page count, version history and the authors named when it was published. We are preparing DOI registration on top of that record, so that a specific version of the data behind a model can be cited in a paper, listed in a data availability statement, and credited to the people who transcribed it.

Until DOIs are switched on, the record page is the reference to use.

Cite this versionTextBibTeXRIS

Aigner, M., & Rainer, L. (2026). Chancery records, Tyrol 1750-1800 (Version 2) [Data set]. Transkribus. https://doi.org/10.…DOI to follow

Licence
CC BY 4.0
Version
2, frozen 10 Mar 2026
Pages
1,240

Start With the Pages You Already Have

Curate a dataset from your existing transcriptions, or start from one another researcher has published, and train a model on data you can point to.