Knowledge-driven synthetic data generation framework for building large language models to classify rare disease subtype of spondyloarthritis

Intelligent Medicine · Published 2026-07-01 · DOI 10.1016/j.imed.2026.07.006

Free full text

Authors (15)

Tao Li, Jing Dong, Xiaojian Ji, Xiaoli Liu, Yimin Song, Ying Lei, Anan Wang, Xiaoyi Liu, Yuhan Guo, Jiaxin Bai, Zheng Zhao, Feng Huang, Zhengbo Zhang, Jian Zhu, Kunpeng Li

Abstract

ABSTRACT: Background: Spondyloarthritis (SpA) is a rare chronic inflammatory disease consisting of subtypes such as ankylosing spondylitis (AS), psoriatic arthritis (PsA), reactive arthritis (ReA), inflammatory bowel disease-associated arthritis (IBDA), and juvenile spondyloarthritis (JSpA). Early diagnosis and precise subtype identification are crucial for improving prognosis. However, in clinical practice, the overlapping clinical manifestations among subtypes, the lack of a single specific diagnostic biomarker, and the extremely limited clinical data for rare subtypes (e.g., IBDA, and JSpA) make early differential diagnosis challenging. Consequently, patients often face the risks of delayed diagnosis, overdiagnosis, and misclassification, thereby adversely affecting individualized treatment decisions and long-term outcomes. Currently, although general-purpose medical large language models (LLMs) have shown potential in certain domains, their performance remains inadequate in the vertical, disease-specific scenario of SpA subtyping, which demands high-quality data. Faced with scarcity and extreme imbalance across multiple categories, these models struggle to meet the practical clinical need for precise classification. Methods: This study proposes a paradigm shift from passively “relying on data” to actively “creating data”. Based on electronic health records (EHRs) of SpA inpatients hospitalized at our institution between 2000 and 2019, we constructed a real-world dataset with five subtype groups (AS, PsA, ReA, IBDA, and JSpA). The dataset encompasses critical disease features such as sacroiliac joint imaging results, Human Leukocyte Antigen B27 (HLA-B27) status, and C-reactive protein levels. A human-machine comparative trial was conducted using the McNemar test. We achieved this through a three-stage framework: (1) Template extraction: High-quality real-world records were extracted from EHRs to establish a diagnostic benchmark for SpA subtypes. (2) Synthetic data generation: We developed a knowledge-guided synthetic data engine that deeply integrates the assessment of spondyloarthritis international society (ASAS) classification criteria and expert priors into a LLM, enabling the targeted generation of clinical case pairs covering rare subtypes with logical consistency. (3) Adaptive optimization: An adaptive weighted sampling strategy was introduced, which dynamically adjusts the training data distribution based on the initial model’s accuracy across categories. This optimizes the model’s learning on weak links, ultimately yielding the final selected model (RobotGPT-SpA). Results: The experiments demonstrated that the proposed method significantly enhances the performance of SpA subtype classification. RobotGPT-SpA achieved an overall accuracy of 0.843, exhibiting balanced and superior diagnostic capabilities across all five subtypes. Notably, it improved performance by 30% on the data-scarce IBDA subtype and by 27% on ReA. Compared with general-purpose medical LLMs with larger parameter counts, our domain-specific model demonstrated higher clinical utility and precision, thereby validating the superiority of the “domain-specific fine-tuning + high-quality data” strategy. In the differential diagnostic comparative experiments, the RobotGPT-SpA model significantly outperformed junior and midlevel physicians, demonstrating higher accuracy, consistency, as well as better sensitivity and specificity. Conclusion: This study confirms that knowledge-guided data engineering is the key to building high-performance, deployable artificial intelligence (AI) models for specialized diseases. Through systematic data generation and balancing strategies, this approach effectively overcomes the application bottlenecks of medical AI in rare and complex chronic diseases. It provides a reliable technical pathway and a universal methodology for achieving early and precise SpA subtype diagnosis in resource-constrained scenarios.

Abstract from DOAJ. Public domain (CC0 1.0).

Read the article at the publisher →

Publication details

Year
2026

Citation

Li, T., Dong, J., Ji, X., et al. (2026). Knowledge-driven synthetic data generation framework for building large language models to classify rare disease subtype of spondyloarthritis. Intelligent Medicine. https://doi.org/10.1016/j.imed.2026.07.006

Related articles