The integration of artificial intelligence into clinical decision support systems offers improved diagnostic accuracy and efficiency, but the opacity of many machine learning models raises concerns about trust, accountability, and regulatory compliance. Explainable artificial intelligence (XAI) has been proposed to address this by making model predictions interpretable to clinicians; however, its true clinical value remains uncertain, and evaluation has not kept pace with methodological development. This systematic review aimed to identify XAI methods used in clinical decision support systems, assess how they are evaluated with clinicians, and determine whether explanations improve diagnostic accuracy, trust, mental models, and efficiency. Following PRISMA guidelines, we searched PubMed, Web of Science, IEEE Xplore, ACM Digital Library, and Scopus for studies published between 2017 and 2024. Eligible studies included original research evaluating XAI in clinical decision support systems with clinician participants and reporting quantitative or qualitative outcomes. Risk of bias was assessed using adapted QUADAS-2 and ROBIS tools, and findings were synthesized narratively with subgroup analyses. From 2,847 records, 68 studies were included. The most common XAI methods were SHAP-based feature attribution (38%), saliency or heatmap methods (29%), concept-based approaches such as TCAV (15%), and counterfactual or example-based explanations (12%). Radiology was the dominant field (54%), followed by dermatology (18%) and pathology (12%). Evaluation approaches were highly inconsistent, with few validated instruments and most studies relying on Likert-scale trust measures or qualitative feedback. Only 16% of studies showed improved diagnostic accuracy with explanations, 67% showed no significant effect, and 17% reported reduced accuracy due to over-reliance or misinterpretation. Although 82% of studies reported increased clinician trust, trust rarely correlated with actual diagnostic performance. Overall, while XAI methods are widely studied in clinical decision support, their evaluation is inconsistent and their benefits are limited. Explanations tend to increase clinician trust without reliably improving diagnostic accuracy, and may sometimes worsen performance, highlighting a trust–accuracy gap that poses important safety concerns for clinical deployment.
Postoperative complications including SSI (2–20%), VTE (1–5%), and respiratory failure (1–8%) significantly increase morbidity, mortality, length of stay, and readmissions. This systematic review assessed machine learning models predicting these outcomes, their performance, external validation, and clinical deployment. A PRISMA-based search (2017–2024) identified 32 eligible studies. Models such as random forest and XGBoost showed AUROC ranges of 0.70–0.85 for SSI, 0.75–0.90 for VTE (outperforming Caprini scores), and 0.75–0.88 for respiratory failure. However, fewer than 20% of studies included external validation and less than 5% reported clinical deployment. Overall, while machine learning models show strong retrospective performance, limited validation and minimal real-world implementation remain major barriers to clinical translation.
Suicidality and depression are major global health burdens, with over 700,000 suicide deaths annually and ~280 million people affected by major depressive disorder. Early risk prediction could support prevention, but traditional methods show limited accuracy. This PRISMA-compliant systematic review evaluated machine learning models for predicting suicidality and depression across electronic health records, social media, and wearable sensor data, focusing on performance, unimodal vs multimodal approaches, and ethical reporting. Searches of PubMed, PsycINFO, IEEE Xplore, arXiv, and ACM Digital Library identified eligible studies. EHR-based models showed AUROC 0.70–0.85 for suicide attempt prediction, social media models 0.70–0.80 for suicidal ideation, and wearable sensor models lower performance (0.65–0.75). Multimodal approaches improved performance by 5–10% over unimodal models. However, fewer than 20% of studies reported ethical considerations such as privacy, bias, or deployment safeguards. Overall, machine learning shows moderate-to-good predictive performance, with multimodal models performing best, but ethical reporting remains critically insufficient for clinical translation.