Intelligent Medicine · Published 2025-09-27 · DOI 10.1016/j.imed.2025.08.003
Background: The development of accurate artificial intelligence (AI) models for liver cancer diagnosis using contrast-enhanced computed tomography (CT) is often hindered by patient privacy regulations and considerable data variations between hospitals. These variations—in CT scanners, patient populations, and disease prevalence—can reduce the performance of standard collaborative training methods such as federated learning (FL). This retrospective study evaluated whether a novel machine learning framework, Hepa-FedBoost, can overcome these challenges to improve diagnostic accuracy for liver cancer classification across a simulated multi-center network without sharing raw patient images. Methods: We developed a new benchmark dataset to better represent real-world clinical diversity. This was done by merging 23,583 CT images of 11 abdominal organs from the public OrganCMNIST subset of the MedMNIST v2 dataset with 11,000 liver lesion images from 500 patients with histopathology-confirmed liver cancer. These cancer-related data were obtained retrospectively from the National Hepatobiliary Standard Database of China (initiated by Beijing Tsinghua Changgung Hospital, with data collected since December 2019). The use of data from the public and private sources did not require authorization from the patients or owners. This combined dataset of 34,583 images was used to simulate a network of 12 hospitals with significant imbalances in data quantity and class labels. The Hepa-FedBoost framework, a type of clustered FL model, was trained on this network. The model works by having each simulated hospital share only compact, anonymized data summaries (prototypes) instead of model parameters or raw data. A central server uses these summaries to group hospitals with similar data and guide the training process to improve overall accuracy. The primary endpoint was the macro area under the receiver operating characteristic curve (mAUC). Furthermore, we compared Hepa-FedBoost’s performance against 6 other established FL methods. Results: On the simulated hospital network, Hepa-FedBoost demonstrated superior diagnostic performance, improving the liver cancer classification mAUC by 4.9 absolute points and overall accuracy by 2.4 absolute points compared with the next best method. It reached a threshold of clinical-grade accuracy in just 17 training rounds, 43% faster than the standard FedAvg approach, while reducing the total data transmission required by approximately half (to 800 MB). Conclusion: The Hepa-FedBoost framework markedly improved the accuracy and efficiency of multi-center liver cancer classification on CT images. By enabling robust collaborative model training while maintaining patient privacy and minimizing IT resource requirements, it represents a practicable AI solution for real-world hospital networks.
Abstract from DOAJ. Public domain (CC0 1.0).
Read the article at the publisher →