INTRODUCTION: Ovarian cancer remains one of the deadliest gynecological malignancies, largely due to delayed diagnosis and the limited specificity of individual tumor markers. This study aimed to develop an explainable machine learning (ML) framework for preoperative discrimination of benign and malignant ovarian tumors using routine laboratory data.
METHODS: The publicly available Soochow University dataset comprising 349 patients (178 benign, 171 malignant) with 49 clinical features was analyzed. Multi-stage missing-value imputation using multivariate imputation by chained equations (MICE) yielded a final cohort of 323 patients and 46 features. Five classifiers—Random Forest, XGBoost, LightGBM, Elastic Net, and a stacked ensemble—were trained using 5-fold repeated cross-validation and evaluated on an independent test set (n = 104). Model interpretability was assessed via SHAP analysis, calibration curves, and decision curve analysis (DCA).
RESULTS: The Elastic Net achieved the best overall performance (AUC = 0.914, 95% CI: 0.862–0.967; sensitivity = 86.8%; specificity = 74.5%). SHAP analysis identified the albumin/globulin ratio, HE4, chloride, and CO₂ combining power as the most influential predictors. DCA confirmed net clinical benefit across threshold probabilities ranging from 0.10 to 0.70.
DISCUSSION AND CONCLUSION: This explainable ML framework demonstrates that transparent models utilizing routine preoperative laboratory data can accurately differentiate benign from malignant ovarian tumors. The proposed approach may support clinical decision-making and preoperative risk stratification in patients with adnexal masses.
Keywords: ovarian cancer, machine learning, explainable artificial intelligence, SHAP, biomarkers, decision curve analysis