Retrieval-Augmented Large Language Models for Clinical Decision Support: A Systematic Review of Hallucination Mitigation and Evidence Grounding

Sumit Barua, Charles Barnabas Rodgers

Research Square · 2026

Abstract

Background: Large language models (LLMs) are increasingly explored for clini­cal decision support systems (CDSS), yet hallucinations—plausible but factually unsupported outputs—pose direct patient safety risks. Retrieval-augmented gen­eration (RAG), which conditions LLM outputs on external evidence at inference time, has emerged as a proposed mitigation strategy. However, the extent to which RAG improves clinical safety and rigor in evaluation remains unclear.

Objective: To systematically examine architectural patterns of RAG in clini­cal CDSS, evaluate empirical evidence for hallucination mitigation, and assess whether current evaluation practices align with clinically meaningful safety standards.

Methods: Following PRISMA guidelines, we reviewed studies published between January 2023 and December 2025 evaluating retrieval-augmented LLM systems in clinical contexts. Inclusion required inference-time retrieval, empirical evalua­tion, and assessment of correctness, grounding, or safety. Data were extracted on architectural design, knowledge sources, evaluation methodology, and validation practices. Owing to heterogeneity in clinical tasks and reported metrics, findings were synthesized narratively.

Results: Thirty studies met the inclusion criteria. Four architectural patterns were identified: single-step retrieval, iterative or agentic retrieval, graph-based retrieval, and hybrid symbolic–neural systems. Across comparative evaluations, RAG approaches were associated with improved factual grounding and reduced unsupported claims relative to non-retrieval LLMs, particularly in guideline- and protocol-based tasks. However, evaluation practices were predominantly retro­spective and metric-driven, with limited clinician involvement and noprospective clinical outcome studies. Recurrent failure modes included retrieval errors, incom­plete or selective integration of evidence, and residual hallucinations despite retrieved context.

Conclusions: RAG architectures represent a substantive advance in improving evidentiary grounding of clinical language models under controlled conditions. Nevertheless, the current evidence base reflects architectural innovation without parallel maturation in clinical validation. Standardized safety-oriented evaluation frameworks, prospective validation, and regulatory-aligned governance mecha­nisms are necessary before RAG-based systems can be considered ready for high-stakes clinical deployment.

📄 이 논문을 인용한 Paperis 글

이 논문이 근거 목록에 올라 있는 Paperis 글입니다.