When the Judge is Wrong: Measuring LLM-as-Judge Reliability Against Graph-Verified Ground Truth in Financial Documents

Agus Sudjianto, Wingyan Lau

SSRN Electronic Journal · 2026

LLM-as-judge is widely adopted for evaluating AI system outputs, yet its reliability for structured tasks remains poorly understood because ground truth is typically established by the same class of system being evaluated. We break this circularity by introducing a graph-verified evaluation framework: we construct provably correct ground truth from financial documents using a Geometric Memory System (GMS), then measure how accurately LLM judges assess answer correctness. Across 258 structured retrieval questions spanning five financial report types, we find that even a strict LLM judge with access to ground truth disagrees with GMSverified truth in 5.9% of cases, with a false acceptance rate of 2.9%-approving wrong answers as correct.

A lenient judge inflates this to 7.9% and a blind judge (without ground truth) reaches 22.3%. A grounded judge-given the full source document but no ground truthachieves only 54.3% agreement, with a 33.1% false acceptance rate-approving 1 in 3 wrong answers even with the source material in context. False acceptances concentrate in exact recall (50%), multi-hop (46%), and threshold (33%) categories-precisely the structured tasks where accuracy matters most in financial regulation.

These findings demonstrate that LLM-as-judge is unreliable for verifying structured information extraction and that graph-based verification provides a necessary alternative.

📄 이 논문을 인용한 Paperis 글

이 논문이 근거 목록에 올라 있는 Paperis 글입니다.

Paperis - When the Judge is Wrong: Measuring LLM-as-Judge Reliability Against Graph-Verified Ground Truth in Financial Documents