Understanding LLM Performance Gaps - Strategic Implications of Stance Detection and Sentiment Analysis in Small Languages Cover Image

Understanding LLM Performance Gaps - Strategic Implications of Stance Detection and Sentiment Analysis in Small Languages
Understanding LLM Performance Gaps - Strategic Implications of Stance Detection and Sentiment Analysis in Small Languages

Author(s): Jurgita Kapočiūtė-Dzikienė, Mantas Vaškevičius, Tadas Sadzevičius
Subject(s): Language studies, Applied Linguistics, Communication studies, Baltic Languages, Security and defense
Published by: NATO Strategic Communications Centre of Excellence
Keywords: Large Language Models (LLMs); Stance Detection; Sentiment Analysis; Multilingual NLP; Low-Resource Languages; Strategic Communications;
Summary/Abstract: Large language models (LLMs) have become essential tools for tasks critical to understanding the information environment, including analysing public discourse, detecting information manipulation, and understanding sentiment towards geopolitically sensitive topics. However, their performance varies significantly across languages. Two previous studies by the NATO Strategic Communications Centre of Excellence progressively documented these disparities. The first, Narrative Detection and Topic Modelling in the Baltics (Barbu, Banerjee, Isupova, and Zeng, 2024), established that natural language processing (NLP) capabilities for Baltic languages remain significantly underdeveloped, finding that while named entity recognition (NER) was reasonably supported, critical downstream tasks such as relationship extraction and plot discovery remained largely unexplored. The second study, AI in Support of StratCom: The Use and Evaluation of Large Language Models in Less Widely Used Official EU Languages (Barbu, Banerjee, Lim, and Zīvere, 2025), expanded the scope considerably by evaluating five contemporary LLMs (GPT-4.0, Mistral Nemo, Mistral Large, Llama 3.1, and Gemini Pro) on three strategic NLP tasks: narrative detection, topic modelling, and aspect-based sentiment analysis (ABSA), across both English and Latvian. The report confirmed a persistent performance gap between high-resource and low-resource languages and revealed that while English outputs showed higher fluency, coherence, and structural accuracy, they were still not sufficient to replace human annotators, especially in zero-shot settings. In Latvian, model outputs were generally less complete, with frequent entity misclassification, such as mislabelling geopolitical organisations like NATO as locations rather than actors, and weaker thematic depth across all evaluated layers. The present report builds directly upon the findings of both preceding studies by shifting the focus to empirically measuring how modern LLMs perform on two additional operationally critical tasks (stance detection and sentiment analysis) across English, Lithuanian, and Russian, targeting politically sensitive entities: Ukraine, Russia, NATO, the USA, and China. This report also responds to conclusions from preceding studies by evaluating next-generation adaptation strategies such as fine-tuning and retrieval-augmented generation (RAG) to determine whether targeted model adaptation can close the performance gaps that both earlier studies identified as a structural limitation of working with less-resourced languages. The findings of this study reveal substantial performance disparities, with direct implications for strategic communication operations in the Baltic region and Eastern Europe. Consistent with the 2025 report’s observation that LLM capabilities degrade markedly for less-resourced languages, this study documents that Russian-language analysis underperforms English by up to 9 percentage points, and that fine-tuned lightweight models can outperform larger proprietary systems relying on prompting alone. The 2025 report also noted that models such as Mistral Large and GPT-4.0 performed most consistently in English but showed marked inconsistencies in Latvian, particularly in sentiment analysis where models struggled to align specific sentiments with their respective aspects. This study extends those findings to stance detection, confirming the same pattern across an additional language pair. Together, the three reports establish a clear progression: from identifying what NLP tools and resources exist for Baltic languages (2024), to benchmarking LLM performance on foundational StratCom tasks in English and Latvian (2025), to empirically quantifying how well current LLMs perform on stance detection and sentiment analysis across English, Lithuanian, and Russian, and demonstrating that targeted adaptation strategies can meaningfully close the gaps that all three studies have identified.

  • E-ISBN-13: 978-9934-619-83-0
  • Print-ISBN-13: 978-9934-619-83-0
  • Page Count: 16
  • Publication Year: 2026
  • Language: English
Toggle Accessibility Mode