Research map: Effect of evaluation prompt strategies on LLM-as-a-judge reliability in critical care
Back to the article
Papers in this map
- A Guideline of Selecting and Reporting Intraclass Correlation Coefficients for Reliability Research · Terry K. K. Koo · 2016 · 29493 citations · Cited by this paper
- Comparison of LTS-D and i-gel in non-paralyzed pediatric patients under general anesthesia: a randomized trial · 2025 · Related
- PhysioBank, PhysioToolkit, and PhysioNet · Ary L. Goldberger · 2000 · 14916 citations · Cited by this paper
- Overtube-assisted removal of an enteral feed bezoar: a case report · 2025 · Related
- Survey of Hallucination in Natural Language Generation · Ziwei Ji · 2022 · 4543 citations · Cited by this paper
- Anabolic androgenic steroids and illicit drugs as potential modulating factors in malignant hyperthermia: a case series · 2025 · Related
- BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence · Wei, Jason · 2022 · 4331 citations · Cited by this paper
- Standards of anaesthesia for total knee and hip arthroplasty procedures. A survey-based study.Part II: Anaesthetic management · 2025 · Related
- MIMIC-IV, a freely accessible electronic health record dataset · Alistair E. W. Johnson · 2023 · 3332 citations · Cited by this paper
- Emergency surgery and post-STEMI dual antiplatelet therapy. Looking for the sweet spot · 2026 · Related
- DISCERN: an instrument for judging the quality of written consumer health information on treatment choices. · David Charnock · 1999 · 3308 citations · Cited by this paper
- Consensus statement of the Paediatric Anaesthesiology and Intensive Therapy Section of the Polish Society of Anaesthesiology and Intensive Therapy on the use of VV ECMO in paediatric patients for the treatment of acute respiratory failure · 2026 · Related
- An overview of clinical decision support systems: benefits, risks, and strategies for success · Reed Taylor Sutton · 2020 · 3103 citations · Cited by this paper
- GLP-1 agonists: a new hope for patients, a new challenge for anaesthetists · 2025 · Related
- Training Language Models to Follow Instructions with Human Feedback · Long Ouyang · 2022 · 1182 citations · Cited by this paper
- Effectiveness of antidepressants in neuropathic pain management: a retrospective, multicenter cross-sectional analysis · 2025 · Related
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena · Lianmin Zheng · 2023 · 794 citations · Cited by this paper
- Advocating for universal access to epidural analgesia for women during childbirth: a scientific review Polish National Social Campaign “Hear the voice of pain” · 2025 · Related
- The TRIPOD-LLM reporting guideline for studies using large language models · Jack Gallifant · 2025 · 530 citations · Cited by this paper
- A framework for human evaluation of large language models in healthcare derived from literature review · Thomas Yu Chow Tam · 2024 · 382 citations · Cited by this paper
- A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity · Yejin Bang · 2023 · 360 citations · Cited by this paper
- Barriers to and Facilitators of Artificial Intelligence Adoption in Health Care: Scoping Review · Masooma Hassan · 2024 · 321 citations · Cited by this paper
- Large pre-trained language models contain human-like biases of what is right and wrong to do · Patrick Schramowski · 2022 · 298 citations · Cited by this paper
- An Empirical Evaluation of Prompting Strategies for Large Language Models in Zero-Shot Clinical Natural Language Processing: Algorithm Development and Validation Study · Sonish Sivarajkumar · 2024 · 208 citations · Cited by this paper
- Large Language Models are not Fair Evaluators · Peiyi Wang · 2024 · 108 citations · Cited by this paper
- A survey on LLM-as-a-judge · Jiawei Gu · 2026 · 105 citations · Cited by this paper
- G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment · Yang Liu · 2023 · 88 citations · Cited by this paper
- Challenges in Implementing Artificial Intelligence in Breast Cancer Screening Programs: Systematic Review and Framework for Safe Adoption · Serene Si Ning Goh · 2024 · 45 citations · Cited by this paper
- Large Language Models Are Not Robust Multiple Choice Selectors · Chujie Zheng · 2023 · 23 citations · Cited by this paper
- Does Prompt Formatting Have Any Impact on LLM Performance? · Jia He · 2024 · 23 citations · Cited by this paper
- Canon Suppression in Corpus-Loaded Assessment: A Two-Assessor Study · Arjun Panickssery · 2024 · 18 citations · Cited by this paper
- RULER: What's the Real Context Size of Your Long-Context Language Models? · Cheng-Ping Hsieh · 2024 · 13 citations · Cited by this paper
- Large Language Models Sensitivity to The Order of Options in Multiple-Choice Questions · Pouya Pezeshkpour · 2023 · 9 citations · Cited by this paper
- Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates · Hui Li Wei · 2024 · 8 citations · Cited by this paper
- Self-Preference Bias in LLM-as-a-Judge · Koki Wataoka · 2024 · 5 citations · Cited by this paper
- Why Does the Effective Context Length of LLMs Fall Short? · Chenxin An · 2024 · 3 citations · Cited by this paper
- Enhancing large language model clinical support information with machine learning risk and explainability: a feasibility study · Yu‐Chang Yeh · 2026 · 2 citations · Cited by this paper
- Untitled · Cited by this paper
Source: OpenAlex (CC0)