Integrating Field Models into Transkribus-based HTR Pipelines for Early East Slavic Cyrillic Manuscripts Cover Image

Integrating Field Models into Transkribus-based HTR Pipelines for Early East Slavic Cyrillic Manuscripts
Integrating Field Models into Transkribus-based HTR Pipelines for Early East Slavic Cyrillic Manuscripts

Author(s): Olga Kalashnikova
Subject(s): Language studies, Language and Literature Studies, Theoretical Linguistics, Applied Linguistics, Computational linguistics, Philology
Published by: Институт за литература - БАН
Keywords: Transkribus; Handwritten Text Recognition (HTR); Field Models; Machine Learning; East Slavic Cyrillic manuscripts; 11th-12th centuries; Church Slavonic

Summary/Abstract: This article addresses the lack of domain-specific HTR tools for 11th-12th century East Slavic Church Slavonic (CS) manuscripts by presenting EastSlavCyrillic11-12_v2, a specialized HTR model trained on eight East Slavic Gospel codices (50013 word tokens) using the Transkribus software platform. The model was developed through an iterative workflow. Having achieved a validation CER of 3,41%, it outperformed the most robust generic CS Transkribus model on the same test materials. Moreover, the effect of integrating a tailored polygon-based Field Model (MAP 80,79%) into the HTR pipeline as a preprocessing step was also evaluated by demonstrating that it measurably improves transcription output on out-of-domain materials, particularly for lines containing superscript characters. Therefore, the article shows that layout recognition may serve as a necessary preprocessing step that can affect transcription quality for early Cyrillic uncial texts. Overall, the paper argues that accurate HTR of early East Slavic manuscripts requires both domain-specific training data and improved polygon-based line segmentation.

  • Issue Year: 2026
  • Issue No: 26
  • Page Range: 61-78
  • Page Count: 18
  • Language: English
Toggle Accessibility Mode