OPTIMIZING INTER-RATER RELIABILITY IN FOREIGN LANGUAGE CONSTRUCTED-RESPONSE ASSESSMENTS
OPTIMIZING INTER-RATER RELIABILITY IN FOREIGN LANGUAGE CONSTRUCTED-RESPONSE ASSESSMENTS
Author(s): Jose Fabian Elizondo-GonzalezSubject(s): Social Sciences
Published by: Visoka škola strukovnih studija za vaspitače "Mihailo Palov"
Keywords: Constructed-response assessment; Costa Rica; EFL; language assessment; reliability
Summary/Abstract: This study examines inter-rater reliability in a constructed-response English proficiency test developed by the Foreign Language Assessment Program (PELEx) in Costa Rica. Thirty university instructors completed three writing tasks aligned with A2, B2, and C1 CEFR bands, each scored by two trained raters. Inter-rater reliability was estimated using percent agreement, Cohen’s weighted kappa, intraclass correlation coefficients (ICCs), and Generalizability Theory (G-Theory). While traditional estimators suggested good to excellent reliability, G-Theory revealed additional sources of error not accounted for by kappa or ICC, particularly prompt-related variability. For example, in the C1 task, person-by-prompt interaction accounted for over 20% of total variance. These findings suggest that while rater training remains important, prompt-related variability must also be addressed to ensure fairness and score comparability. Incorporating calibrated prompts and structured scoring protocols across proficiency levels may strengthen the reliability of constructed-response tasks, especially in high-stakes settings.
Journal: Research in Pedagogy
- Issue Year: 15/2025
- Issue No: 2
- Page Range: 471-482
- Page Count: 12
- Language: English
