Evaluating AI-Based Automated Essay Scoring Through Signal Detection Theory: Beyond Aggregate Agreement Metrics

Xiaoliang Zhou

Applied Psychological Measurement · 2026

The rapid adoption of Large Language Models (LLMs) in educational assessment has reshaped scoring practices, yet evaluation remains tethered to aggregate reliability metrics like Quadratic Weighted Kappa, which obscure discrimination and rater effects. This study applies Signal Detection Theory to evaluate eight state-of-the-art LLMs (including Claude 3.5 Haiku, DeepSeek-V3, Gemini 3 Flash, GPT-4o, and Grok 4.1) against expert human raters across 1,726 essays. By decoupling discrimination from response criteria, I provide a diagnostic analysis of AI scoring behavior.

Results indicate that human raters exhibit significantly superior evaluative precision, with average discrimination estimates approximately double those of the AI models. Furthermore, LLMs are prone to pronounced centrality effects and score compression, systematically failing to award the highest rubric tiers. These findings demonstrate that low human-machine agreement stems from both a deficit in discriminative accuracy and systematic shifts in response criteria.

Ultimately, this research provides a robust framework for calibrating and selecting AI scoring systems based on specific pedagogical goals and fairness requirements.

📄 이 논문을 인용한 Paperis 글

이 논문이 근거 목록에 올라 있는 Paperis 글입니다.