Research map: Evaluation Methods for Inference-Time Retrieval-Augmented and Graph Retrieval-Augmented Large Language Models in Health Care: Scoping Review
Back to the article
Papers in this map
PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and Explanation
· Andrea C. Tricco · 2018 · 44107 citations · Cited by this paper
Reliability and construct validity of a multi-model LLM-as-judge evaluator battery for clinician-facing clinical question answering: a controlled adversarial benchmark study (Preprint)
· Henry Bergman · 2026 · Cites this paper
Model and Task-Aware Test-Time Scaling Strategies for Large Language and Vision-Language Models in Medicine: Evaluation Study
· 2026 · Related
Updated methodological guidance for the conduct of scoping reviews
· Micah D.J. Peters · 2020 · 7658 citations · Cited by this paper
How Social Media and Chatbot Bans Could Backfire
· 2026 · Related
PRISMA-S: an extension to the PRISMA Statement for Reporting Literature Searches in Systematic Reviews
· Melissa Lyle Rethlefsen · 2021 · 3412 citations · Cited by this paper
Retrieval-augmented generation in medicine: A scoping review of technical implementations, clinical applications, and ethical considerations
· 2026 · Related
A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions
· Lei Huang · 2024 · 2105 citations · Cited by this paper
Retrieval augmented large language model system for comprehensive drug contraindications
· 2026 · Related
Toward expert-level medical question answering with large language models
· K. K. Singhal · 2025 · 958 citations · Cited by this paper
JADE: jawbone lesion diagnosis and decision supporting system
· 2026 · Related
A new sociotechnical model for studying health information technology in complex adaptive healthcare systems
· Dean F. Sittig · 2010 · 893 citations · Cited by this paper
Nursing Retrieval-Augmented Generation: Retrieval augmented generation for nursing question answering with large language models
· 2025 · Related
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
· Lianmin Zheng · 2023 · 794 citations · Cited by this paper
Development and Evaluation of a Retrieval-Augmented Generation-Based Electronic Medical Record Chatbot System
· 2025 · Related
Testing and Evaluation of Health Care Applications of Large Language Models
· Suhana Bedi · 2024 · 619 citations · Cited by this paper
Multimodal Knowledge Graph–Guided RAG-LLM for Clinical Decision Support in Pediatric Leukemia
· 2026 · Related
Reporting guideline for the early stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI
· Baptiste Vasey · 2022 · 572 citations · Cited by this paper
The TRIPOD-LLM reporting guideline for studies using large language models
· Jack Gallifant · 2025 · 530 citations · Cited by this paper
The application of large language models in medicine: A scoping review
· Xiangbin Meng · 2024 · 407 citations · Cited by this paper
Almanac — Retrieval-Augmented Language Models for Clinical Medicine
· Cyril Zakka · 2024 · 400 citations · Cited by this paper
FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation
· Sewon Min · 2023 · 348 citations · Cited by this paper
A systematic review of large language model (LLM) evaluations in clinical medicine
· Sina Shool · 2025 · 330 citations · Cited by this paper
All That’s ‘Human’ Is Not Gold: Evaluating Human Evaluation of Generated Text
· Elizabeth Clark · 2021 · 281 citations · Cited by this paper
Large language models in patient education: a scoping review of applications in medicine
· Serhat Aydın · 2024 · 276 citations · Cited by this paper
Evaluation framework to guide implementation of AI systems into healthcare settings
· Sandeep Reddy · 2021 · 231 citations · Cited by this paper
Optimization of hepatological clinical guidelines interpretation by large language models: a retrieval augmented generation-based framework
· Simone Kresevic · 2024 · 206 citations · Cited by this paper
Enabling Large Language Models to Generate Text with Citations
· Tianyu Gao · 2023 · 187 citations · Cited by this paper
Trust in Artificial Intelligence–Based Clinical Decision Support Systems Among Health Care Workers: Systematic Review
· Hein Minn Tun · 2025 · 179 citations · Cited by this paper
Applying generative AI with retrieval augmented generation to summarize and extract key clinical information from electronic health records
· Mohammad Alkhalaf · 2024 · 150 citations · Cited by this paper
Development of a liver disease–specific large language model chat interface using retrieval-augmented generation
· Jin Ge · 2024 · 147 citations · Cited by this paper
Evaluation of Retrieval-Augmented Generation: A Survey
· Hao Yu · 2025 · 147 citations · Cited by this paper
Retrieval augmented generation for 10 large language models and its generalizability in assessing medical fitness
· Yu He Ke · 2025 · 140 citations · Cited by this paper
Exploring the risks of automation bias in healthcare artificial intelligence applications: A Bowtie analysis
· Moustafa Abdelwanis · 2024 · 136 citations · Cited by this paper
Biomedical knowledge graph-optimized prompt generation for large language models
· Karthik Soman · 2024 · 134 citations · Cited by this paper
Seven Failure Points When Engineering a Retrieval Augmented Generation System
· Scott A. Barnett · 2024 · 131 citations · Cited by this paper
ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems
· Jon Saad-Falcon · 2024 · 125 citations · Cited by this paper
MedRAG: Enhancing Retrieval-augmented Generation with Knowledge Graph-Elicited Reasoning for Healthcare Copilot
· Xuejiao Zhao · 2025 · 106 citations · Cited by this paper
Medical Graph RAG: Evidence-based Medical Large Language Model via Graph Retrieval-Augmented Generation
· Junde Wu · 2025 · 76 citations · Cited by this paper
Source: OpenAlex (CC0)