Cardiovascular disease remains the leading global cause of death, emphasizing the need for improved risk stratification beyond traditional tools such as Framingham, ASCVD, QRISK, and SCORE, which show limitations in diverse modern populations. Machine learning methods applied to electronic health records can enhance prediction by capturing complex, high-dimensional, and nonlinear relationships. This systematic review (2017–2022) evaluated machine learning models for cardiovascular risk prediction using EHR data, focusing on discrimination (AUROC, AUPRC), calibration, external validation, and reporting quality including TRIPOD adherence. A PRISMA-compliant search identified peer-reviewed studies applying machine learning to EHR-based cardiovascular risk prediction. Risk of bias was assessed using PROBAST, and narrative synthesis was conducted due to heterogeneity. Twenty-nine studies were included. XGBoost, random forest, and neural networks were the most common models and generally outperformed logistic regression and traditional risk scores in discrimination. However, calibration was infrequently reported, and external validation was limited, often showing reduced performance. Machine learning models demonstrate improved predictive discrimination over conventional risk scores, but limited calibration assessment and weak external validation constrain clinical applicability. Stronger validation frameworks are needed for clinical translation.
Postoperative complications including SSI (2–20%), VTE (1–5%), and respiratory failure (1–8%) significantly increase morbidity, mortality, length of stay, and readmissions. This systematic review assessed machine learning models predicting these outcomes, their performance, external validation, and clinical deployment. A PRISMA-based search (2017–2024) identified 32 eligible studies. Models such as random forest and XGBoost showed AUROC ranges of 0.70–0.85 for SSI, 0.75–0.90 for VTE (outperforming Caprini scores), and 0.75–0.88 for respiratory failure. However, fewer than 20% of studies included external validation and less than 5% reported clinical deployment. Overall, while machine learning models show strong retrospective performance, limited validation and minimal real-world implementation remain major barriers to clinical translation.