Retrieval-Augmented Large Language Models for Clinical Decision Support: A Systematic Review of Hallucination Mitigation and Evidence Grounding
Sumit Barua, Charles Barnabas Rodgers
Research Square · 2026
Abstract
Background: Large language models (LLMs) are increasingly explored for clinical decision support systems (CDSS), yet hallucinations—plausible but factually unsupported outputs—pose direct patient safety risks. Retrieval-augmented generation (RAG), which conditions LLM outputs on external evidence at inference time, has emerged as a proposed mitigation strategy. However, the extent to which RAG improves clinical safety and rigor in evaluation remains unclear.
Objective: To systematically examine architectural patterns of RAG in clinical CDSS, evaluate empirical evidence for hallucination mitigation, and assess whether current evaluation practices align with clinically meaningful safety standards.
Methods: Following PRISMA guidelines, we reviewed studies published between January 2023 and December 2025 evaluating retrieval-augmented LLM systems in clinical contexts. Inclusion required inference-time retrieval, empirical evaluation, and assessment of correctness, grounding, or safety. Data were extracted on architectural design, knowledge sources, evaluation methodology, and validation practices. Owing to heterogeneity in clinical tasks and reported metrics, findings were synthesized narratively.
Results: Thirty studies met the inclusion criteria. Four architectural patterns were identified: single-step retrieval, iterative or agentic retrieval, graph-based retrieval, and hybrid symbolic–neural systems. Across comparative evaluations, RAG approaches were associated with improved factual grounding and reduced unsupported claims relative to non-retrieval LLMs, particularly in guideline- and protocol-based tasks. However, evaluation practices were predominantly retrospective and metric-driven, with limited clinician involvement and noprospective clinical outcome studies. Recurrent failure modes included retrieval errors, incomplete or selective integration of evidence, and residual hallucinations despite retrieved context.
Conclusions: RAG architectures represent a substantive advance in improving evidentiary grounding of clinical language models under controlled conditions. Nevertheless, the current evidence base reflects architectural innovation without parallel maturation in clinical validation. Standardized safety-oriented evaluation frameworks, prospective validation, and regulatory-aligned governance mechanisms are necessary before RAG-based systems can be considered ready for high-stakes clinical deployment.