Sudan Journal of Medical Sciences · Published 2026-06-30 · DOI 10.18502/sjms.v21i2.19868
Baleegh Elsir Ahmed Mohammed, Tanzeel Elsir Ahmed Mohammed, Anthony Paul Breitbach
Background: Constructing high-quality multiple-choice questions (MCQs) is critical yet resource-intensive in medical education. Large language models (LLMs), such as ChatGPT, have become promising tools for generating MCQs in medical education; however, their effectiveness and reliability remain underexplored. This study aimed to explore the reported effects and characteristics of using LLMs to generate MCQs in medical education and to describe the nature of the available evidence. Methods: Arksey and O’Malley’s framework was used to conduct this scoping review. Three databases were searched: ScienceDirect, MEDLINE, and Web of Science. Studies published from 2022 onward were included and screened using eligibility criteria. Data were extracted and analyzed descriptively, and qualitative findings were thematically synthesized using Braun and Clarke’s approach. Results: Eighteen studies published between 1/1/ 2022 and 11/1/2025 were included. Half of the studies employed quantitative methods, while others used qualitative (27.8%) or mixed-methods (22.2%) designs. ChatGPT was used in 17 studies (94.4%). Regarding the effects of using LLMs for generating MCQs, these tools have shown potential to produce MCQs with satisfactory psychometric properties within seconds, thereby reducing educators’ workload. However, reported limitations include inaccurate output and the generation of MCQs with low complexity. Other reported limitations also included cost and ethical concerns. To address these limitations, studies suggest using high-quality prompts and recommend expert review and refinement of the generated MCQs. Conclusion: LLMs present a promising solution for generating MCQs in medical education; however, some limitations persist. High-quality research across diverse geographical regions is needed to determine their effectiveness.
Abstract from DOAJ. Public domain (CC0 1.0).
Read the article at the publisher →
Mohammed, B., Mohammed, T., Breitbach, A. (2026). The Effects of Large Language Models on the Generation of Multiple-choice Questions in Medical Education: A Scoping Review. Sudan Journal of Medical Sciences. https://doi.org/10.18502/sjms.v21i2.19868