Smart Extract
From page image to structured data – in one step
A single AI model reads your document and produces text, tables, named entities, and metadata. No pipeline. No configuration.
Drag an image here
Select a file...PNG or JPG up to 10 Mb
Extracted text will appear here
One model replaces the pipeline
Traditional document processing chains multiple stages – layout detection, line segmentation, text recognition – each passing errors to the next. Smart Extract replaces the entire chain with a single model that reads a page image and produces structured output directly.
Atlas
Atlas is the first Smart Extract model. It reads a complete page image and returns structured data in a single step: text, layout, tables, and named entities, all in reading order.
It handles complex layouts without configuration – forms, multi-column pages, nested tables, and tables of contents – across modern and historical documents. Atlas also serves as the base model for fine-tuning: custom Smart Extract models are trained on top of it.
Structured output, not plain text
The result is typed, structured data – ready for your database, your search index, or your spreadsheet.
Text & layout
Reading order across columns and around images. Paragraphs, marginalia, footnotes, and page numbers as distinct elements.
Tables
Structured grids with rows, cells, and merged spans. Nested tables and tables of contents handled automatically.
Named entities
People, places, and dates tagged in place within the transcribed text.
Document classification
Language and script type identified for every page. Printed, handwritten, or mixed.
Complex layouts, no setup
Forms, multi-column pages, nested table structures, and marginalia – handled out of the box. The output is designed for direct downstream use without manual correction of the document structure.
Fine-tune for your project
The base model handles a wide range of documents out of the box. Fine-tune it on your own material to create a custom output schema – your own elements, entity types, and data fields. Whatever you can consistently label on a page, the model can learn to produce.
Fine-tuning will initially be available for selected projects only. Interested? Get in touch to discuss your use case.
More at the Transkribus User Conference
We will share more about Smart Extract, custom models, and what comes next at the Transkribus User Conference. Stay tuned for the full announcement.