Can GPT-4o detect explicit contrast? A pilot study with Croatian corpus data
Can GPT-4o detect explicit contrast? A pilot study with Croatian corpus data
Author(s): Zrinka KolakovićSubject(s): Language studies, Language and Literature Studies, Applied Linguistics, Computational linguistics, Philology
Published by: Институт за литература - БАН
Keywords: information structure; contrast; corpus linguistics; large language models; annotation; Croatian
Summary/Abstract: Annotation of information-structural categories is widely recognized as both time-consuming and context-dependent, posing challenges for large-scale empirical research. This pilot study, based on 125 Croatian sentences extracted from the Riznica corpus, investigates the potential of LLM-assisted annotation of contrast, understood as the explicit highlighting of alternatives in discourse-embedding. Sentences and their discourse context were processed and annotated for the presence of contrast by GPT-4o. Qualitative inspection of the assigned annotation labels indicates that the model can capture explicitly marked contrasts, with structures introduced by ordinal numbers being significantly easier to detect. However, in some instances, GPT-4o struggles to adhere to the operational definition of contrast: it assigns contrastiveness to sentences with only implicit contrast and remains unreliable in more context-dependent cases. While not generalizable beyond this specific tested model and case, the results suggest that to benefit both formalization of theoretical IS concepts and automatic annotation, further experimentation requires improved prompt design, refined annotation protocols, data formatting, and testing of additional LLMs.
Journal: Scripta & e-Scripta
- Issue Year: 2026
- Issue No: 26
- Page Range: 79-88
- Page Count: 10
- Language: English
