Evaluation of large language models in a national orthopaedic proficiency examination: Implications for health informatics and medical education

Health Informatics Journal · Published 2026-07-01 · DOI 10.1177/14604582261470637

Free full text

Authors (1)

Bünyamin Arı

Abstract

Objective This study evaluates the performance of large language models (LLMs)—ChatGPT-4.0, Gemini 2.0 Pro, o3-mini, Doctor GPT and DeepSeek-V3—in a national orthopaedic proficiency examination and explores their implications for health informatics and medical education. The responses of these models were analysed to assess accuracy rates and differences between models. Method A total of 100 multiple-choice questions from the 2024 TOTEK examination were administered to each AI model under identical conditions. Correct and incorrect responses were recorded, and differences in performance were evaluated using chi-square testing and frequency analysis. Question categories were also compared to identify domain-specific variations. Results o3-mini achieved the highest accuracy rate (79%), while Gemini 2.0 showed the lowest (68%); all models exceeded the 60% pass threshold. A statistically significant difference between models was identified in the Surgical Procedures category, in which Gemini 2.0 answered fewer questions correctly (23/36) than the other models (30–32/36) (χ 2 = 9.87, df = 4, p = 0.043). No significant differences were observed in the remaining categories (all p > 0.05), and the overall difference in accuracy between models did not reach statistical significance (χ 2 = 4.01, df = 4, p = 0.405). Clinical decision-making and visual content-based questions were the most challenging for all models. Conclusion AI models demonstrate generally high accuracy in medical examinations; however, they struggle with interpreting clinical context, recognising atypical medical scenarios and answering questions involving visual content.

Abstract from DOAJ. Public domain (CC0 1.0).

Read the article at the publisher →

Publication details

Year
2026

Citation

Arı, B. (2026). Evaluation of large language models in a national orthopaedic proficiency examination: Implications for health informatics and medical education. Health Informatics Journal. https://doi.org/10.1177/14604582261470637

Related articles