Skip to content
  • Pricing
sindlev · PyLaia · Published August 30, 2026

Vat. gr. 2228 – Medieval Greek HTR

Text Recognition

Description

This manuscript-specific HTR model was developed for Vaticanus graecus 2228 (Vat. gr. 2228), a 14th-century Medieval Greek paper manuscript held by the Biblioteca Apostolica Vaticana. The manuscript contains rhetorical treatises and Byzantine commentaries and is characterized, in the sections targeted by the model, by a comparatively regular layout and consistent scholarly minuscule handwriting. The model is based on manually curated line-level ground truth with expert transcriptions of Medieval Greek. The transcriptions preserve the text closely, including polytonic diacritics and the resolution of abbreviations, tachygraphic signs, ligatures, and other palaeographical features. Text normalization was limited to removing invisible characters, standardizing whitespace, applying Unicode NFC normalization, and mapping visually indistinguishable upper-right tick encodings to U+2019 RIGHT SINGLE QUOTATION MARK (`’`). The model and its training data were developed within the E-Rhetoric project, funded by the Carlsberg Foundation (grant CF24-1999). The work was carried out by Nicklas Sindlev Andersen, Byron MacDougall, Ugo Valori, Tariq Yousef, and Aglae Pizzone. For a more detailed description of the model development, ground truth curation, and evaluation, see Andersen et al. (2026), From Manuscript to Model: Developing HTR for Medieval Greek. (DOI: 10.63317/47tkobv5mgbu)

Try this model

Drag an image here

Select a file...

PNG or JPG up to 10 Mb

Wolpi
AI Assistant

By uploading an image, you accept our terms and privacy policy.

Use this modelOpen in Transkribus
Very low error rate3.33% CER

Character Error Rate (CER) measures the percentage of characters incorrectly recognised. Lower is better. This model scored 3.33% on its validation set. As a rule of thumb, a CER below 10% is considered good for most handwritten material.

Measured on the model's own validation data. Results on your documents may differ depending on handwriting style, document condition, language, and how closely your material resembles the training data.

Words111,364
Lines9,148
Training Pages227
Model ID627649
Languages
Greek Ancient (to 1453)
Centuries
14th c.