Research map: Effect of evaluation prompt strategies on LLM-as-a-judge reliability in critical care

Back to the article

Papers in this map

  1. A Guideline of Selecting and Reporting Intraclass Correlation Coefficients for Reliability Research · Terry K. K. Koo · 2016 · 29493 citations · Cited by this paper
  2. Comparison of LTS-D and i-gel in non-paralyzed pediatric patients under general anesthesia: a randomized trial · 2025 · Related
  3. PhysioBank, PhysioToolkit, and PhysioNet · Ary L. Goldberger · 2000 · 14916 citations · Cited by this paper
  4. Overtube-assisted removal of an enteral feed bezoar: a case report · 2025 · Related
  5. Survey of Hallucination in Natural Language Generation · Ziwei Ji · 2022 · 4543 citations · Cited by this paper
  6. Anabolic androgenic steroids and illicit drugs as potential modulating factors in malignant hyperthermia: a case series · 2025 · Related
  7. BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence · Wei, Jason · 2022 · 4331 citations · Cited by this paper
  8. Standards of anaesthesia for total knee and hip arthroplasty procedures. A survey-based study.Part II: Anaesthetic management · 2025 · Related
  9. MIMIC-IV, a freely accessible electronic health record dataset · Alistair E. W. Johnson · 2023 · 3332 citations · Cited by this paper
  10. Emergency surgery and post-STEMI dual antiplatelet therapy. Looking for the sweet spot · 2026 · Related
  11. DISCERN: an instrument for judging the quality of written consumer health information on treatment choices. · David Charnock · 1999 · 3308 citations · Cited by this paper
  12. Consensus statement of the Paediatric Anaesthesiology and Intensive Therapy Section of the Polish Society of Anaesthesiology and Intensive Therapy on the use of VV ECMO in paediatric patients for the treatment of acute respiratory failure · 2026 · Related
  13. An overview of clinical decision support systems: benefits, risks, and strategies for success · Reed Taylor Sutton · 2020 · 3103 citations · Cited by this paper
  14. GLP-1 agonists: a new hope for patients, a new challenge for anaesthetists · 2025 · Related
  15. Training Language Models to Follow Instructions with Human Feedback · Long Ouyang · 2022 · 1182 citations · Cited by this paper
  16. Effectiveness of antidepressants in neuropathic pain management: a retrospective, multicenter cross-sectional analysis · 2025 · Related
  17. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena · Lianmin Zheng · 2023 · 794 citations · Cited by this paper
  18. Advocating for universal access to epidural analgesia for women during childbirth: a scientific review Polish National Social Campaign “Hear the voice of pain” · 2025 · Related
  19. The TRIPOD-LLM reporting guideline for studies using large language models · Jack Gallifant · 2025 · 530 citations · Cited by this paper
  20. A framework for human evaluation of large language models in healthcare derived from literature review · Thomas Yu Chow Tam · 2024 · 382 citations · Cited by this paper
  21. A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity · Yejin Bang · 2023 · 360 citations · Cited by this paper
  22. Barriers to and Facilitators of Artificial Intelligence Adoption in Health Care: Scoping Review · Masooma Hassan · 2024 · 321 citations · Cited by this paper
  23. Large pre-trained language models contain human-like biases of what is right and wrong to do · Patrick Schramowski · 2022 · 298 citations · Cited by this paper
  24. An Empirical Evaluation of Prompting Strategies for Large Language Models in Zero-Shot Clinical Natural Language Processing: Algorithm Development and Validation Study · Sonish Sivarajkumar · 2024 · 208 citations · Cited by this paper
  25. Large Language Models are not Fair Evaluators · Peiyi Wang · 2024 · 108 citations · Cited by this paper
  26. A survey on LLM-as-a-judge · Jiawei Gu · 2026 · 105 citations · Cited by this paper
  27. G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment · Yang Liu · 2023 · 88 citations · Cited by this paper
  28. Challenges in Implementing Artificial Intelligence in Breast Cancer Screening Programs: Systematic Review and Framework for Safe Adoption · Serene Si Ning Goh · 2024 · 45 citations · Cited by this paper
  29. Large Language Models Are Not Robust Multiple Choice Selectors · Chujie Zheng · 2023 · 23 citations · Cited by this paper
  30. Does Prompt Formatting Have Any Impact on LLM Performance? · Jia He · 2024 · 23 citations · Cited by this paper
  31. Canon Suppression in Corpus-Loaded Assessment: A Two-Assessor Study · Arjun Panickssery · 2024 · 18 citations · Cited by this paper
  32. RULER: What's the Real Context Size of Your Long-Context Language Models? · Cheng-Ping Hsieh · 2024 · 13 citations · Cited by this paper
  33. Large Language Models Sensitivity to The Order of Options in Multiple-Choice Questions · Pouya Pezeshkpour · 2023 · 9 citations · Cited by this paper
  34. Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates · Hui Li Wei · 2024 · 8 citations · Cited by this paper
  35. Self-Preference Bias in LLM-as-a-Judge · Koki Wataoka · 2024 · 5 citations · Cited by this paper
  36. Why Does the Effective Context Length of LLMs Fall Short? · Chenxin An · 2024 · 3 citations · Cited by this paper
  37. Enhancing large language model clinical support information with machine learning risk and explainability: a feasibility study · Yu‐Chang Yeh · 2026 · 2 citations · Cited by this paper
  38. Untitled · Cited by this paper

Source: OpenAlex (CC0)