Prediction of Postoperative Vomiting Within 24 Hours Using Machine Learning With Large Language Model–Enhanced Interpretability: Development and Validation Study

JMIR Medical Informatics · Published 2026-07-31 · DOI 10.2196/84260

Free full text

Authors (5)

Huan-Jun Wang, Wei-Po Lee, Tz-Ping Gau, Kuang-I Cheng, Cheng-Ru Wei

Abstract

Abstract BackgroundPostoperative nausea and vomiting are common complications after anesthesia. However, vomiting represents a clinically distinct and objectively measurable endpoint. ObjectiveThis study aimed to develop and internally validate predictive models for postoperative vomiting within 24 hours using structured perioperative data and unstructured clinical text, while introducing a structured framework that separates feature construction from interpretability using large language models (LLMs). MethodsWe analyzed 33,460 anesthesia records from a single center (2019‐2022). Two temporally defined prediction tasks were constructed to reflect real-world clinical decision-making and prevent information leakage: a preoperative model using variables available before anesthesia induction, and a perioperative model using variables available up to the end of surgery. Structured data were modeled using machine learning algorithms (logistic regression, Extreme Gradient Boosting, Light Gradient Boosting Machine [LightGBM]). Unstructured clinical text was incorporated through a deterministic, concept-driven preprocessing pipeline, where LLMs were used solely for normalization (temperature=0) without feature generation, followed by rule-based concept mapping and feature encoding. Post hoc interpretability was further supported using an LLM-based Question Answering Chain module. Model performance was evaluated using receiver operating characteristic-area under the curve (AUC), precision-recall AUC, calibration metrics, and threshold-based operating characteristics. Classification thresholds were selected using the Youden J statistic, and all metrics were reported with 95% CIs derived from bootstrap resampling. Decision curve analysis was performed to assess clinical utility. ResultsA total of 33,460 surgical procedures were included, of which 3607 (10.8%) experienced postoperative vomiting within 24 hours. In the preoperative task, LightGBM achieved an AUC of 0.729 (95% CI 0.706‐0.749), compared with 0.610 (95% CI 0.588‐0.632) for the Apfel score. In the end-of-surgery task, LightGBM achieved an AUC of 0.735 (95% CI 0.714‐0.757). At the Youden-optimal threshold, the negative predictive value exceeded 0.95 across all models. Decision curve analysis demonstrated positive net benefit across clinically relevant threshold probabilities. Incorporating text-derived features provided modest improvements, while LLM-based explanation modules generated structured, natural-language explanations intended to enhance interpretability without substantially improving predictive performance. ConclusionsMachine learning models can effectively predict postoperative vomiting within 24 hours using perioperative data. The proposed framework demonstrates that LLMs can be integrated in a controlled and reproducible manner—restricted to deterministic normalization and post hoc reasoning—to generate natural-language explanations intended to enhance the interpretability of model predictions, without introducing information leakage or altering predictive modeling. As no formal clinician-based evaluation was conducted, this interpretability benefit cannot yet be objectively confirmed, and the generated explanations should be regarded as a useful interpretability aid to be validated in future clinician-centered studies. External, multicenter validation is required before broader clinical applicability can be assumed.

Abstract from DOAJ. Public domain (CC0 1.0).

Read the article at the publisher →

Publication details

Year
2026

Citation

Wang, H., Lee, W., Gau, T., et al. (2026). Prediction of Postoperative Vomiting Within 24 Hours Using Machine Learning With Large Language Model–Enhanced Interpretability: Development and Validation Study. JMIR Medical Informatics. https://doi.org/10.2196/84260

Related articles