Research map: Evaluation Methods for Inference-Time Retrieval-Augmented and Graph Retrieval-Augmented Large Language Models in Health Care: Scoping Review

Back to the article

Papers in this map

  1. PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and Explanation · Andrea C. Tricco · 2018 · 44107 citations · Cited by this paper
  2. Reliability and construct validity of a multi-model LLM-as-judge evaluator battery for clinician-facing clinical question answering: a controlled adversarial benchmark study (Preprint) · Henry Bergman · 2026 · Cites this paper
  3. Model and Task-Aware Test-Time Scaling Strategies for Large Language and Vision-Language Models in Medicine: Evaluation Study · 2026 · Related
  4. Updated methodological guidance for the conduct of scoping reviews · Micah D.J. Peters · 2020 · 7658 citations · Cited by this paper
  5. How Social Media and Chatbot Bans Could Backfire · 2026 · Related
  6. PRISMA-S: an extension to the PRISMA Statement for Reporting Literature Searches in Systematic Reviews · Melissa Lyle Rethlefsen · 2021 · 3412 citations · Cited by this paper
  7. Retrieval-augmented generation in medicine: A scoping review of technical implementations, clinical applications, and ethical considerations · 2026 · Related
  8. A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions · Lei Huang · 2024 · 2105 citations · Cited by this paper
  9. Retrieval augmented large language model system for comprehensive drug contraindications · 2026 · Related
  10. Toward expert-level medical question answering with large language models · K. K. Singhal · 2025 · 958 citations · Cited by this paper
  11. JADE: jawbone lesion diagnosis and decision supporting system · 2026 · Related
  12. A new sociotechnical model for studying health information technology in complex adaptive healthcare systems · Dean F. Sittig · 2010 · 893 citations · Cited by this paper
  13. Nursing Retrieval-Augmented Generation: Retrieval augmented generation for nursing question answering with large language models · 2025 · Related
  14. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena · Lianmin Zheng · 2023 · 794 citations · Cited by this paper
  15. Development and Evaluation of a Retrieval-Augmented Generation-Based Electronic Medical Record Chatbot System · 2025 · Related
  16. Testing and Evaluation of Health Care Applications of Large Language Models · Suhana Bedi · 2024 · 619 citations · Cited by this paper
  17. Multimodal Knowledge Graph–Guided RAG-LLM for Clinical Decision Support in Pediatric Leukemia · 2026 · Related
  18. Reporting guideline for the early stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI · Baptiste Vasey · 2022 · 572 citations · Cited by this paper
  19. The TRIPOD-LLM reporting guideline for studies using large language models · Jack Gallifant · 2025 · 530 citations · Cited by this paper
  20. The application of large language models in medicine: A scoping review · Xiangbin Meng · 2024 · 407 citations · Cited by this paper
  21. Almanac — Retrieval-Augmented Language Models for Clinical Medicine · Cyril Zakka · 2024 · 400 citations · Cited by this paper
  22. FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation · Sewon Min · 2023 · 348 citations · Cited by this paper
  23. A systematic review of large language model (LLM) evaluations in clinical medicine · Sina Shool · 2025 · 330 citations · Cited by this paper
  24. All That’s ‘Human’ Is Not Gold: Evaluating Human Evaluation of Generated Text · Elizabeth Clark · 2021 · 281 citations · Cited by this paper
  25. Large language models in patient education: a scoping review of applications in medicine · Serhat Aydın · 2024 · 276 citations · Cited by this paper
  26. Evaluation framework to guide implementation of AI systems into healthcare settings · Sandeep Reddy · 2021 · 231 citations · Cited by this paper
  27. Optimization of hepatological clinical guidelines interpretation by large language models: a retrieval augmented generation-based framework · Simone Kresevic · 2024 · 206 citations · Cited by this paper
  28. Enabling Large Language Models to Generate Text with Citations · Tianyu Gao · 2023 · 187 citations · Cited by this paper
  29. Trust in Artificial Intelligence–Based Clinical Decision Support Systems Among Health Care Workers: Systematic Review · Hein Minn Tun · 2025 · 179 citations · Cited by this paper
  30. Applying generative AI with retrieval augmented generation to summarize and extract key clinical information from electronic health records · Mohammad Alkhalaf · 2024 · 150 citations · Cited by this paper
  31. Development of a liver disease–specific large language model chat interface using retrieval-augmented generation · Jin Ge · 2024 · 147 citations · Cited by this paper
  32. Evaluation of Retrieval-Augmented Generation: A Survey · Hao Yu · 2025 · 147 citations · Cited by this paper
  33. Retrieval augmented generation for 10 large language models and its generalizability in assessing medical fitness · Yu He Ke · 2025 · 140 citations · Cited by this paper
  34. Exploring the risks of automation bias in healthcare artificial intelligence applications: A Bowtie analysis · Moustafa Abdelwanis · 2024 · 136 citations · Cited by this paper
  35. Biomedical knowledge graph-optimized prompt generation for large language models · Karthik Soman · 2024 · 134 citations · Cited by this paper
  36. Seven Failure Points When Engineering a Retrieval Augmented Generation System · Scott A. Barnett · 2024 · 131 citations · Cited by this paper
  37. ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems · Jon Saad-Falcon · 2024 · 125 citations · Cited by this paper
  38. MedRAG: Enhancing Retrieval-augmented Generation with Knowledge Graph-Elicited Reasoning for Healthcare Copilot · Xuejiao Zhao · 2025 · 106 citations · Cited by this paper
  39. Medical Graph RAG: Evidence-based Medical Large Language Model via Graph Retrieval-Augmented Generation · Junde Wu · 2025 · 76 citations · Cited by this paper

Source: OpenAlex (CC0)