A Review of a Hybrid Evaluation Framework for Clinical AI: Integrating Interrater Reliability with the "LLM as a Judge" Methodology

Dhruv Limaye, Rutwik Shah, Chandra Jonelgadda

Research Square · 2026

Abstract Background The rapid integration of Large Language Models (LLMs) into clinical workflows is driven by studies claiming performance that rivals or exceeds human experts. However, these claims frequently rely on "Gold Standard" datasets generated by human annotators, assuming these annotations represent an objective truth. This assumption overlooks the inherent variability and noise in human clinical judgment.

Without quantifying this variability through Inter-Rater Reliability (IRR), it is impossible to define the upper bound of performance against which AI should be measured. Objective To systematically evaluate the methodological rigor of clinical LLM studies, specifically focusing on the reporting of IRR and the establishment of robust human performance baselines. Methods We conducted a systematic review of 66 original research articles published between 2024 and 2025 assessing LLMs (including GPT-4, LLaMA, and Med-PaLM) across diverse clinical domains, including Radiology, Surgery, and Ophthalmology.

We extracted data regarding study design, ground truth generation, conflict resolution strategies, and the specific IRR metrics reported (e.g., Cohen’s Kappa, Fleiss’ Kappa, ICC). Results Our analysis reveals substantial heterogeneity in evaluation standards. A significant subset of studies failed to report any IRR metric, relying instead on single-rater annotations or unverified consensus, effectively treating subjective judgment as objective fact.

Where reported, human agreement was frequently imperfect, with Kappa values often ranging from moderate (0.40) to substantial (0.80), but rarely achieving perfection. This indicates a "Human Standard" significantly lower than 100%, suggesting that many LLMs achieving "expert-level" accuracy are approximating the noise inherent in the expert standard. Furthermore, methodology for resolving expert disagreement was often opaque, limiting reproducibility.

Conclusion Current evaluation standards for clinical LLMs are insufficient to substantiate claims of performance above that of humans. The lack of standardized IRR reporting obscures the distinction between true algorithmic superiority and statistical noise. We propose a standardized reporting framework for clinical AI studies that mandates the calculation of IRR to explicitly define the true human standard.

Adopting this framework is critical for ensuring that LLMs are evaluated against a transparent, reproducible, and statistically valid standard before deployment in patient care.

📄 이 논문을 인용한 Paperis 글

이 논문이 근거 목록에 올라 있는 Paperis 글입니다.

Paperis - A Review of a Hybrid Evaluation Framework for Clinical AI: Integrating Interrater Reliability with the "LLM as a Judge" Methodology