Evaluation of Large Language Models for Automated Simple CPT Coding in Foot and Ankle Surgery

Foot & Ankle Orthopaedics · Published 2026-04-01 · DOI 10.1177/24730114261448207

Free full text

Authors being retrieved — see the publisher record. https://doi.org/10.1177/24730114261448207

Abstract

Background: Accurate Current Procedural Terminology ( CPT ) coding is crucial for appropriate billing and reimbursement in foot and ankle surgery but is often time-consuming and prone to error. Large language models (LLMs) offer an approach to automate coding and reduce administrative burden, yet their performance in orthopaedic subspecialties remains limited. This study sought to evaluate the accuracy of 5 publicly available LLMs—ChatGPT-5 Mini, Google Gemini 2.5 Flash, Claude 4.0 Sonnet, Deepseek V3, and Perplexity—in correctly generating CPT codes for single-code (simple) foot and ankle procedures. Methods: Twenty-one common procedures identified by a single CPT code were selected. Each LLM was queried 4 times with standardized prompts requesting CPT codes. Accuracy was assessed based on the correct identification of CPT codes, with statistical analyses comparing performance across models and trials. Results: Perplexity achieved the highest accuracy (92.9%, 95% CI = 87.4%-98.4%), whereas Deepseek V3 performed worst (48.2%, 95% CI = 37.5%-58.9%). Global χ 2 testing showed significant differences in coding accuracy among models ( P  < .001). Pairwise comparisons revealed Perplexity outperformed ChatGPT-5 Mini, Claude 4.0 Sonnet, and Deepseek V3; Google Gemini 2.5 Flash outperformed Deepseek V3 and ChatGPT-5 Mini; and Claude 4.0 Sonnet outperformed Deepseek V3. Conclusion: LLM performance in CPT coding for simple foot and ankle procedures was highly variable, with some models demonstrating acceptable accuracy whereas others performed poorly. These findings highlight that current LLMs are not sufficiently reliable for independent clinical use. Select models may serve as first-pass aids when combined with careful human verification.

Abstract from DOAJ. Public domain (CC0 1.0).

Read the article at the publisher →

Publication details

Year
2026

Related articles