Healthcare Informatics Research · Published 2026-04-30 · DOI 10.4258/hir.2026.32.2.134
Objectives This study aimed to identify predictors of health literacy (HL), focusing on nonlinear relationships and interaction effects in a nationally representative population. Methods This cross-sectional study analyzed data from 8,630 Korean adults participating in the Korea Health Panel Survey. HL was assessed using the Korean version of the European Health Literacy Survey Questionnaire 16-item (HLS-EU-Q16) and categorized as sufficient or insufficient. An extreme gradient boosting algorithm (XGBoost) was applied, incorporating survey weights. Model performance was evaluated using standard metrics, including the area under the receiver operating characteristic curve (AUC) and Brier score. Shapley additive explanations (SHAP) values were calculated to quantify individual feature importance and identify interaction effects among 69 features. Results The XGBoost model achieved good discrimination (AUC = 0.840) and calibration (Brier score = 0.161). Age (24.4%) and education level (19.0%) were the most influential predictors. SHAP interaction analysis identified a meaningful interaction between age and education level (mean |interaction value| = 0.078), with interaction plots indicating positive patterns among adults aged 60–80 with lower educational attainment. The contribution of online health information use varied by age, showing a negative association among younger adults but a positive association among older adults. Conclusions HL is shaped by nonlinear and interactive effects across sociodemographic and health-related factors. Explainable machine learning approaches can facilitate the identification of high-priority populations and support the development of tailored, patient-centered educational strategies to improve HL and promote engagement with care.
Abstract from DOAJ. Public domain (CC0 1.0).
Read the article at the publisher →