Generalizability Theory for LLM-as-Evaluator Reliability: Univariate and Multivariate Variance Decomposition Across Models, Prompts, and Temperatures in Text Classification

Jin Liu

SSRN Electronic Journal · 2026

Large language models (LLMs) are increasingly deployed as text annotators, yet most studies report a single agreement coefficient under one prompting condition, conflating item difficulty, model choice, prompt design, and temperature into a single reliability estimate. We apply univariate and multivariate generalizability theory (G-theory) to decompose LLM annotation variance across three instrumentation facets (evaluator, prompt type, temperature) with seed replication. Under a fully crossed design (four evaluators × three prompt types × six temperatures × three seeds; 64,800 observations across hatespeech, mental-health, and drug-review tasks), item difficulty dominated total variance (74.6-93.0%), the item × evaluator interaction was the primary nuisance component (2.65-14.75%), and temperature contributed effectively zero variance.

Generalizability coefficients ranged from Eρ 2 = 0.940 to 0.988; D-studies showed that two to three evaluators, two prompt types, and one temperature suffice for Eρ 2 > 0.90, a 92% cost reduction. Facet severity analysis revealed systematic evaluator biases invisible to agreement metrics, and rubric prompts produced greater inter-model consistency than alternatives but not always higher human-label alignment, exposing a reliability-validity distinction that standard metrics conflate. Multivariate composite scoring from psychiatric comorbidity flags exceeded holistic reliability (∆Eρ 2 c = 0.019, CI excluding zero), while drug-review holistic scores outperformed all composites, depending on whether dimensions represent independent constructs or derived decompositions.

The drug-review cost dimension illustrated this: measurement-degenerate under G-theory yet a strong predictor in the companion study. Cross-task validation showed that G-theory captures distinct sources of evaluator disagreement, from boundary uncertainty on a single severity signal to mixed information from comorbidity and multi-criteria trade-offs.

📄 이 논문을 인용한 Paperis 글

이 논문이 근거 목록에 올라 있는 Paperis 글입니다.