Can LLMs Judge Legal Accuracy? Reliability of LLM Evaluators for High-Stakes Insurance QA in a Low-Resource Language

David Beauchemin, Richard Khoury

Research Square · 2026

Abstract Large language models (LLMs) are increasingly used as automated evaluators, a setting known as LLM-as-a-Judge, but their reliability outside general-domain English tasks is still poorly understood. We test this in a demanding case: judging the factual correctness of answers about Quebec automobile insurance, a specialized legal domain expressed in Quebec French. Evaluating 52 open-source and commercial judges against an expert-authored benchmark, we find that even the strongest is only moderately reliable and that different judges frequently disagree.

Crucially, a judge’s overall quality does not predict how often it accepts a wrong answer as correct, the error that matters most in this setting; enabling step-by-step reasoning helps, and, contrary to the AI-favoritism reported in general domains, the judges are, if anything, stricter on AI-generated answers than on human ones. Selecting judges by their ability to catch errors, supplying an authoritative reference, and combining a few strong judges each reduces these dangerous acceptances. We conclude that a single LLM judge is not trustworthy enough to evaluate high-stakes legal answers on its own, and argue for cost-aware, human-in-the-loop use in deployment.

📄 이 논문을 인용한 Paperis 글

이 논문이 근거 목록에 올라 있는 Paperis 글입니다.