Heatmap-based visualisation of the linguistic polymorphism in Transcarpathian East Slavic
Heatmap-based visualisation of the linguistic polymorphism in Transcarpathian East Slavic
Author(s): Ilia AfanasevSubject(s): Language studies, Language and Literature Studies, Foreign languages learning, Theoretical Linguistics, Applied Linguistics, Computational linguistics, Eastern Slavic Languages, Philology
Published by: Институт за литература - БАН
Keywords: visualisation; Ukrainian; Transcarpathian; Lemko; heatmap; language variation and change
Summary/Abstract: The study examines the complex phenomenon of language polymorphism (an umbrella term necessary in cases where there is no clear understanding of the source of variation and/or change: an occasional error, the influence of the native lect of the data collector, or other, more prototypical causes) through the lens of morphosyntactic tagging. It introduces a new visualisation method, based on heatmaps, that facilitates both a bird’s-eye view of the data and closer inspection. The main dataset of the research is TransVar, a historical corpus of Transcarpathian East Slavic (Ukrainian small territorial lects1 recorded in the early twentieth century: Bojko, Central Transcarpathian, Hutsul, and Lemko2). The results of its morphosyntactic tagging with Stanza constitute the case study. The analysis, conducted with the help of visualisation methods, identifies the primary source of tagger errors: copular clauses. It demonstrates how a single non-standard structure may significantly reduce tagger performance across all morphosyntactic levels. The prospects for future research include the collection of additional material and the development of more fine-grained annotation techniques.
Journal: Scripta & e-Scripta
- Issue Year: 2026
- Issue No: 26
- Page Range: 11-22
- Page Count: 12
- Language: English
