A pulmonary nodule is worth 8 × 8 × 8 words: Computed tomography-based 3D vision transformer predicts early-stage high-grade lung adenocarcinoma of micropapillary and/or solid subtypes

Intelligent Medicine · Published 2025-11-21 · DOI 10.1016/j.imed.2025.11.001

Free full text

Authors being retrieved — see the publisher record. https://doi.org/10.1016/j.imed.2025.11.001

Abstract

Background: Early-stage high-grade lung invasive adenocarcinoma (IAC) has poor prognosis and is hard to identify using conventional radiological assessment. Current reliance on postoperative histology to identify high-grade subtypes delays risk-adapted surgical planning. Three-dimensional (3D) vision transformers (ViTs) may improve prediction by modeling long-range dependencies in computed tomography (CT) scans. We aimed to develop and validate 3D-ViT and Swin Transformer (SwinT) for preoperative CT-based prediction of early-stage high-grade IAC subtypes (micropapillary/solid), benchmarking against ResNet. Methods: A multicenter cohort of 1028 patients with surgically confirmed early-stage lung adenocarcinoma was divided into training (n = 806), validation (n = 100), and external test (n = 122) sets. 3D-ViT, SwinT, and ResNet models were trained on CT to classify nodules harboring high-grade histologic patterns. A novel decision-aid tool for IAC surgery was provided. Performance was evaluated using area under the curve (AUC), accuracy, sensitivity, specificity, and precision. Attention mapping was performed to interpret 3D-ViT decision-making. Results: The 3D-ViT model achieved AUC values of 0.856 (95% CI: 0.845–0.877) (validation) and 0.806 (95% CI: 0.790–0.816) (testing), compared to 0.854 (95% CI: 0.841–0.872) (validation) and 0.760 (95% CI: 0.743–0.776) (testing) for the ResNet baseline. 3D-ViT showed balanced accuracy, sensitivity, specificity, and precision in the validation set. In external testing, 3D-ViT significantly outperformed ResNet in all metrics with P < 0.01. The SwinT-based AlignSen model from decision-aid tool prioritized sensitivity for high-grade IAC (88.0% validation, 93.4% testing), while maintaining specificity (76.0% validation, 54.1% testing) which significantly outperformed ViT and ResNet-based AlignSen models’ specificity (60.0% and 54% validation, 52.5% and 24.6% testing, respectively). Attention maps highlighted 3D nodule heterogeneity and peripheral irregularities. Conclusion: The 3D-ViT model demonstrated robust accuracy and generalizability in predicting high-grade IAC subtypes using CT images. Integration of preoperative SwinT into clinical workflows may offer a viable alternative to intraoperative pathology subtyping, potentially reducing reliance on frozen sections while optimizing surgical planning.

Abstract from DOAJ. Public domain (CC0 1.0).

Read the article at the publisher →

Publication details

Year
2025

Related articles