Pathogen prevalence, feature composition and cross-centre generalisability of machine learning diagnostic models for multi-pathogen respiratory infection

Scritto il 01/10/2026
da Fengmiao Hu

Front Public Health. 2026 Sep 16;14:1961816. doi: 10.3389/fpubh.2026.1961816. eCollection 2026.

ABSTRACT

OBJECTIVE: To evaluate the predictive information contained in the restricted surveillance feature set (demographic, temporal and specimen variables) for respiratory pathogen identification, and to examine factors associated with model performance.

METHODS: We retrospectively analysed 24,689 acute respiratory infection cases from seven sentinel hospitals in Sichuan, China; after quality control, 21,395 samples with complete 21-pathogen testing were included. Per-pathogen multi-label binary classifiers were trained using 15 features available at presentation; logistic regression, random forest, back-propagation neural network and XGBoost were compared under an identical split and threshold-selection protocol, with thresholds derived from out-of-fold probabilities. Label-definition sensitivity, leave-one-hospital-out cross-validation and prevalence-performance association were analysed.

RESULTS: Random forest and XGBoost performed comparably (XGBoost the primary model), with Macro-F1 0.1553 (95% CI 0.104-0.2091), 0.2202 for common pathogens and Macro-AUC 0.7702; the four-algorithm spread was 0.0611. Log-prevalence correlated strongly with per-pathogen F1 (r = 0.851); the seven rare pathogens reached Macro-F1 0.0255. Leave-one-hospital-out validation gave Macro-F1 0.1073 ± 0.0289 (-30.9%) and Macro-AUC 0.6841 (-11.2%). Feature ablation showed temporal and non-temporal features were largely redundant (Macro-F1 0.1153 vs. 0.1147; AUC 0.6892/0.6865 vs. 0.7702). Treating untested as negative lowered Macro-AUC to 0.7653, whereas simulated false-negative label contamination had minimal effect.

CONCLUSION: Models built on the restricted surveillance feature set showed limited diagnostic performance and were not demonstrated suitable for clinical deployment. Performance was strongly associated with pathogen prevalence and constrained by limited feature informativeness, while algorithm choice contributed little and cross-centre generalisability was limited. Priorities include multimodal feature integration, rare-pathogen accrual and prospective multicentre external validation.

PMID:42819249 | PMC:PMC13623830 | DOI:10.3389/fpubh.2026.1961816