Evaluation of large language models in clinical neuroanatomy: a comparative scoring analysis based on accuracy, concordance, insight, and anatomical terminology accuracy

Folia Morphologica (Poland) · Published 2026-02-06 · DOI 10.5603/fm.110023

Free full text

Authors being retrieved — see the publisher record. https://doi.org/10.5603/fm.110023

Abstract

BACKGROUND: Large language models (LLMs), such as ChatGPT-4 and Gemini 2.5, are increasingly being evaluated for clinical reasoning and medical education. However, their performance in structured, neuroanatomical diagnostic tasks and their use of accurate anatomical terminology remain underexplored. MATERIALS AND METHODS: This study assessed the diagnostic performance of ChatGPT-4 and Gemini 2.5 across 20 novel clinical neuroanatomy cases. Each model was tested under three prompting conditions: no prompt, short prompt, and long prompt. Responses were independently scored by two evaluators using the ACI (accuracy, concordance, insight) framework and a separate ATA (anatomical terminology accuracy) Scale. Statistical comparisons were conducted using the Friedman test and Wilcoxon signed-rank tests. Inter-rater reliability was assessed via intraclass correlation coefficients (ICCs). RESULTS: ChatGPT-4 consistently outperformed Gemini 2.5 in total ACI + ATA scores under all prompt conditions (p < 0.05). The highest component scores were observed in diagnostic accuracy, while Insight scores were generally lower, particularly for Gemini. ATA scores were significantly higher for ChatGPT-4, reflecting better use of standardized anatomical terminology. Prompt specificity influenced Gemini’s performance more than ChatGPT’s. Notably, long prompts improved overall scores in some cases but led to performance decline in others, suggesting prompt-induced narrowing. CONCLUSIONS: ChatGPT-4 demonstrates strong and stable diagnostic performance in neuroanatomical cases, with high accuracy and precise anatomical language. Gemini 2.5 shows potential, but is more sensitive to prompt variations and performs inconsistently in complex scenarios. Structured scoring frameworks like ACI and ATA offer valuable tools for evaluating LLMs in both clinical and educational settings.

Abstract from DOAJ. Public domain (CC0 1.0).

Read the article at the publisher →

Publication details

Year
2026

Related articles