Multi model deliberation improves the stability of automated scoring by large language models across genres and assessment contexts

Xiaoying ZHENG, Benhui Chen, Yuqing Chen, Yun Yang

Research Square · 2026

Abstract Large language models show considerable promise for automated scoring of subjective questions, yet their scoring instability and limited cross-context evidence constrain reliable application in educational assessment. This study proposes a multi-model deliberation framework that employs three heterogeneous Chinese large language models (Baichuan2, Qwen, and ChatGLM) in a three-round process: independent scoring, information exchange, and final confirmation. Each model reassesses its judgement after seeing the other models' scores and rationales.

The framework is evaluated on two complementary datasets covering cross-genre breadth and within-essay stability. On the ASAP-AES corpus, covering 8 prompts across argumentative, source-dependent, and narrative genres (5 essays × 30 trials per prompt; n = 7,350 trait-level paired observations), inter-model disagreement decreased significantly from R1 to R3 in every prompt (Holm-corrected p < 0.001). The pooled effect size was large (Cohen's d z = 1.08, rank-biserial r rb = 0.92) and uniform across genres and rubric scales.

On a Chinese university journalism corpus (10 authentic responses × 100 trials; 9,000 scoring records), the coefficient of variation decreased from 7.29% to 2.78%—a 61.9% relative improvement (paired t(9) = 7.98, p < 0.001, Cohen's d = 2.52). The deliberation effect on scoring precision replicated consistently across the two datasets, which differ in language, genre, assessment scale, rubric structure, and rater type. Complementary analyses revealed a task-dependent pattern in alignment with human reference scores: rank-order correspondence was preserved on the journalism corpus but decreased modestly on the multi-prompt ASAP corpus.

Both corpora exhibited a systematic post-deliberation upward shift, which is shown to be recoverable through standard human-anchored linear calibration. The framework therefore enhances scoring precision and can be paired with anchor-based calibration when absolute alignment with human ratings is required.

📄 이 논문을 인용한 Paperis 글

이 논문이 근거 목록에 올라 있는 Paperis 글입니다.