Title : Machine learning prediction model for depressive symptoms in pre-diabetic middle-aged and elderly individuals in communities based on the perspective of cardiorenal metabolic syndrome: A study using the CHARLS database
Abstract:
Objective: To develop and internally validate a risk prediction model for incident depressive symptoms in middle-aged and older adults with prediabetes and cardiovascular-kidney-metabolic syndrome (CKM) stages 1-2, based on longitudinal data from the China Health and Retirement Longitudinal Study (CHARLS).
Methods: Participants aged ≥45 at baseline in 2015 who met criteria for prediabetes and CKM stages 1-2, had a Center for Epidemiologic Studies-Depression Scale (CESD-10) score and had complete follow-up outcomes were included. Incident depressive symptoms were defined as a CESD-10 score ≥10 at any assessment in 2018 or 2020. Participants were stratified by follow-up outcome and randomly divided into training (70%) and testing sets (30%). All data preprocessing was performed within the training set. After correlation filtering, variables underwent fivefold cross-validated LASSO regression, and the final variable set was determined using the one-standard-error rule. Optimal parameters for the logistic regression model were identified via grid search. Model performance and interpretability were evaluated using receiver operating characteristic (ROC) curves, precision-recall (PR) curves, calibration curves, decision curve analysis (DCA), bootstrap internal validation, and Shapley additive explanations (SHAP). Subgroup analyses by age and CKM stage were conducted, and a nomogram was developed.
Results: A total of 2,340 participants were included, of whom 1,024 developed depressive symptoms during follow-up, yielding an incidence rate of 43.76%. The dataset was split into 1,638 in the training set and 702 in the test set. Of 59 candidate variables, 49 remained after correlation filtering. LASSO selected 25 and 7 variables under lambda.min and lambda.1se, respectively. The final model incorporated baseline CESD-10 total score, years of education, self-rated health, total cognitive score, hemoglobin, mean diastolic blood pressure, and uric acid. The area under the ROC curve (AUC) was 0.717 (95% CI: 0.693–0.743) in the training set and 0.746 (95% CI: 0.708–0.783) in the test set. At the Youden threshold of 0.410 derived from the training set, the test set achieved accuracy, sensitivity, specificity, positive predictive value, and F1 score of 0.667, 0.739, 0.610, 0.596, and 0.660, respectively, with an average precision of 0.680 and a Brier score of 0.204. Calibration intercept and slope in the test set were 0.049 and 1.301, respectively; the Hosmer-Lemeshow test yielded P = 0.429. The average optimistic bias across 500 bootstrap samples was 0.005, and the optimism-corrected AUC was 0.712 (95% CI: 0.688–0.735). The AUC of the univariate model constructed using only baseline CESD-10 scores was 0.703 (95% CI: 0.663–0.741). The final model showed an improvement in AUC by 0.043 compared to the univariate model, with a statistically significant difference (P = 0.002). The AUCs for the subgroups aged 45–59 years, ≥60 years, CKM stage 1, and CKM stage 2 were 0.744, 0.748, 0.687, and 0.768, respectively.
Conclusion: The logistic regression model developed based on baseline psychological symptoms, subjective health, education and cognitive function, as well as selected metabolic and physical indicators, demonstrates moderate discriminatory ability, good stability, and interpretability. It may serve as a reference for depression risk stratification in community chronic disease management, although external validation is still needed to further evaluate its value for broader application.

