Semi-automatic Identification and Attribution of Biblical Quotations in Old Church Slavonic Texts: Computational Methods in the Project “BogoSlov”
Semi-automatic Identification and Attribution of Biblical Quotations in Old Church Slavonic Texts: Computational Methods in the Project “BogoSlov”
Author(s): Maksim Gavrilkov, Iris Karafillidis, Tomáš Mikulka, Martin Ruskov, Janusz SzablewskiSubject(s): Language studies, Language and Literature Studies, Applied Linguistics, Computational linguistics, Philology
Published by: Институт за литература - БАН
Keywords: Biblical Quotations; Old Church Slavonic; Text Reuse; Digital Slavic Studies; LLMs
Summary/Abstract: The identification of biblical quotations in Old Church Slavonic (OCS) texts is a longstanding philological challenge. We present a systematic evaluation of computational methods for identifying and attributing biblical quotations in OCS texts, conducted within the BogoSlov project. Working with available digitized manuscripts we evaluate three algorithmic approaches – lemmatized N-gram search, longest common subsequence (LCS), Best Matching 25 (BM25) – and two Sentence Transformers. Each method is assessed against exemplary quotations. No single method proves universally optimal. BM25 performs reliably across varied input conditions, including incomplete or (un)punctuated queries. N-grams excel when word order is preserved and lemmatization can resolve morphological variation. LCS shows strength in longer queries. Sentence Transformers underperform on OCS, but show promise for allusive material in more modern languages. The methods should be considered complementary: their respective strengths cover distinct quotation types, degrees of literality, and query conditions; a guided sequenced application yields more reliable results than any single approach alone. This constitutes a semi-automatic workflow, involving expert review of algorithmically ranked attributions. The approach is also applicable to the broader challenge of source attribution in Slavonic patristics and provides a foundation for future AI-assisted identification of quotations across OCS, Greek, and Latin traditions.
Journal: Scripta & e-Scripta
- Issue Year: 2026
- Issue No: 26
- Page Range: 105-120
- Page Count: 16
- Language: English
