Research map: Validation of large language models for multiple-choice assessment in cranio-maxillofacial surgery – Part A: In-silico accuracy on AI- and human-authored questions

Back to the article

Papers in this map

  1. Large language models encode clinical knowledge · Karan Singhal · 2023 · 3892 citations · Cited by this paper
  2. Pulmonary embolism after radial forearm free flap reconstruction for oral squamous cell carcinoma: a multicenter retrospective cohort study · 2026 · Related
  3. Large language models in medicine · Arun James Thirunavukarasu · 2023 · 3892 citations · Cited by this paper
  4. Distinct synovial proinflammatory cytokine and hyaluronic acid profiles in temporomandibular joint of patients with Class Ⅱ and Class Ⅲ dentofacial deformities · 2026 · Related
  5. How Does ChatGPT Perform on the United States Medical Licensing Examination (USMLE)? The Implications of Large Language Models for Medical Education and Knowledge Assessment · Aidan Gilson · 2023 · 2165 citations · Cited by this paper
  6. Superficial temporal artery perforator flap: an anatomical study and topographic mapping of cutaneous perforators · 2026 · Related
  7. A Review of Multiple-Choice Item-Writing Guidelines for Classroom Assessment · Thomas M. Haladyna · 2002 · 963 citations · Cited by this paper
  8. Accuracy of fiducial-based augmented reality in auricular reconstruction · 2026 · Related
  9. Programmatic assessment: From assessment of learning to assessment for learning · Lambert Schuwirth · 2011 · 902 citations · Cited by this paper
  10. Comparison of audiometric outcomes between modified two-flap and Furlow's palatoplasty techniques: A retrospective cohort study · 2026 · Related
  11. Guidelines for clinical trial protocols for interventions involving artificial intelligence: the SPIRIT-AI extension · Samantha Cruz Rivera · 2020 · 732 citations · Cited by this paper
  12. The position of the mental foramen on panoramic radiographs of patients with neurofibromatosis type 1 · 2026 · Related
  13. 2018 Consensus framework for good assessment · John J. Norcini · 2018 · 338 citations · Cited by this paper
  14. Deep learning-automatic 3D analysis of regional condylar remodeling and skeletal relapse following bimaxillary surgery: a two-year follow-up study · 2026 · Related
  15. Performance of ChatGPT and GPT-4 on Neurosurgery Written Board Examinations · Rohaid Ali · 2023 · 216 citations · Cited by this paper
  16. Outcomes of the pectoralis major muscle flap for covering a mandibular reconstruction plate: Our experience at the Leiden University Medical Center · 2026 · Related
  17. ChatGPT versus human in generating medical graduate exam multiple choice questions—A multinational prospective study (Hong Kong S.A.R., Singapore, Ireland, and the United Kingdom) · Billy Ho Hung Cheung · 2023 · 202 citations · Cited by this paper
  18. Announcements · 2026 · Related
  19. Reporting guidelines for clinical trials of artificial intelligence interventions: the SPIRIT-AI and CONSORT-AI guidelines · Hussein Fadil Ibrahim · 2021 · 177 citations · Cited by this paper
  20. PEEK plates for the fixation of mandibular condylar base fractures: a finite element analysis comparison of single and double plate configurations with titanium and resorbable systems · 2026 · Related
  21. Performance of Generative Large Language Models on Ophthalmology Board–Style Questions · LOUIS Z. CAI · 2023 · 141 citations · Cited by this paper
  22. ChatGPT-4 Omni Performance in USMLE Disciplines and Clinical Skills: Comparative Analysis · Brenton T. Bicknell · 2024 · 106 citations · Cited by this paper
  23. A Framework for Improving the Quality of Multiple-Choice Assessments · Marie Tarrant · 2012 · 81 citations · Cited by this paper
  24. The impact and opportunities of large language models like ChatGPT in oral and maxillofacial surgery: a narrative review · Behrus Puladi · 2023 · 64 citations · Cited by this paper
  25. ChatGPT 3.5 fails to write appropriate multiple choice practice exam questions · Alexander Ngo · 2023 · 58 citations · Cited by this paper
  26. The Accuracy and Capability of Artificial Intelligence Solutions in Health Care Examinations and Certificates: Systematic Review and Meta-Analysis · William J. Waldock · 2024 · 44 citations · Cited by this paper
  27. Exploring the Performance of ChatGPT Versions 3.5, 4, and 4 With Vision in the Chilean Medical Licensing Examination: Observational Study · Marcos Rojas · 2024 · 36 citations · Cited by this paper
  28. How does artificial intelligence master urological board examinations? A comparative analysis of different Large Language Models’ accuracy and reliability in the 2022 In-Service Assessment of the European Board of Urology · Lisa Kollitsch · 2024 · 34 citations · Cited by this paper
  29. Evaluating Artificial Intelligence Chatbots in Oral and Maxillofacial Surgery Board Exams: Performance and Potential · Reema Mahmoud · 2024 · 25 citations · Cited by this paper
  30. Comparative performance evaluation of ChatGPT-4 Omni and Gemini Advanced in the Turkish Dentistry Specialization Exam · Makbule Buse Dündar Sarı · 2026 · 5 citations · Cited by this paper
  31. Artificial Intelligence Chatbots Taking American Board of Endodontics Simulated Oral Board Examination · Poorya Jalali · 2026 · 3 citations · Cited by this paper

Source: OpenAlex (CC0)