Sepsis remains a major cause of mortality in intensive care units worldwide, with an estimated 49 million cases and over 11 million deaths annually, highlighting the need for earlier detection to improve outcomes. This systematic review synthesizes evidence on machine learning models for early sepsis prediction in adult ICU patients from 2017 to 2021, focusing on prediction horizons, data modalities, and validation approaches. A comprehensive search of PubMed, Embase, IEEE Xplore, ACM Digital Library, and arXiv identified studies meeting criteria for ICU-based sepsis prediction with at least a 4-hour forecast window, following PRISMA guidelines. Of 1,478 records screened, 35 studies were included, with prediction horizons ranging from 4 to 24 hours and most relying on hourly vital sign data and internal validation. Reported performance varied widely depending on horizon length, data sampling, and validation rigor, with external validation generally producing lower but more realistic results. Overall, while machine learning models show promising predictive ability, limitations in generalizability and standardization remain, emphasizing the need for stronger validation frameworks and reporting practices to support clinical translation.
Cardiovascular disease remains the leading global cause of death, emphasizing the need for improved risk stratification beyond traditional tools such as Framingham, ASCVD, QRISK, and SCORE, which show limitations in diverse modern populations. Machine learning methods applied to electronic health records can enhance prediction by capturing complex, high-dimensional, and nonlinear relationships. This systematic review (2017–2022) evaluated machine learning models for cardiovascular risk prediction using EHR data, focusing on discrimination (AUROC, AUPRC), calibration, external validation, and reporting quality including TRIPOD adherence. A PRISMA-compliant search identified peer-reviewed studies applying machine learning to EHR-based cardiovascular risk prediction. Risk of bias was assessed using PROBAST, and narrative synthesis was conducted due to heterogeneity. Twenty-nine studies were included. XGBoost, random forest, and neural networks were the most common models and generally outperformed logistic regression and traditional risk scores in discrimination. However, calibration was infrequently reported, and external validation was limited, often showing reduced performance. Machine learning models demonstrate improved predictive discrimination over conventional risk scores, but limited calibration assessment and weak external validation constrain clinical applicability. Stronger validation frameworks are needed for clinical translation.
Emergency department crowding is a persistent global healthcare challenge linked to longer wait times, increased patients leaving without being seen, worse clinical outcomes, and staff burnout. It also contributes to ambulance diversion and inefficient resource use, worsening hospital operational strain. This systematic review evaluates machine learning models for predicting ED crowding and optimizing patient flow, focusing on input features (e.g., arrival rates, acuity, bed availability) and reported operational outcomes such as waiting times and ambulance delays. A PRISMA-compliant review was conducted across PubMed, Embase, IEEE Xplore, and Scopus. Included studies applied machine learning to ED crowding or patient flow prediction and reported operational or crowding outcomes. Due to heterogeneity, a narrative synthesis was used, and risk of bias was assessed using an adapted tool. Thirty-two studies met inclusion criteria, using classification, regression, time-series, and deep learning models. Common predictors included arrival patterns, occupancy, and bed availability. While predictive performance was generally high, few studies evaluated real-world operational impacts, and most remained retrospective. Although machine learning models demonstrate strong predictive accuracy for ED crowding, evidence of real-world operational benefits remains limited. A clear gap exists between prediction and implementation into clinical workflow and decision-making. Future research should focus on translating predictions into measurable improvements in ED performance.
Transformer-based architectures have significantly advanced clinical natural language processing by improving the capture of contextual relationships in unstructured electronic health records compared to earlier recurrent and convolutional models, with domain-specific variants such as ClinicalBERT and BioBERT designed to better handle clinical terminology, abbreviations, and specialized language, thereby improving information extraction performance, although the relative impact of different pre-training strategies remains insufficiently synthesized and requires systematic evaluation of corpus selection and fine-tuning approaches; this systematic review mapped studies focusing on pre-training corpora, fine-tuning methods, and named entity recognition performance across entity types such as medications, diseases, procedures, laboratory tests, and social determinants of health, using PRISMA-guided methods and searches across PubMed, ACL Anthology, arXiv, and IEEE Xplore, identifying 32 eligible studies from 1,247 records; findings showed that ClinicalBERT, BioBERT, and PubMedBERT were the most frequently evaluated models, pre-trained on datasets such as MIMIC-III, PubMed abstracts, and mixed biomedical corpora, with consistent evidence that domain-specific pre-training outperforms general-domain BERT models on benchmarks like i2b2 and n2c2 despite variation across entity types and fine-tuning strategies, while clinical pre-training on large EHR corpora improves named entity recognition and optimized fine-tuning approaches such as lower learning rates and data augmentation further enhance performance, particularly for medications and diseases, underscoring the importance of domain adaptation and the need for more standardized evaluation protocols in clinical NLP research.
Patient no-shows in outpatient clinics (5%–30% across specialties) disrupt scheduling efficiency, increase wait times, and strain healthcare resources. To address this, healthcare systems are increasingly applying machine learning (ML) for predictive scheduling support. This systematic review synthesizes ML approaches for predicting outpatient no-shows, focusing on model types, feature usage, and reported operational deployment outcomes, with emphasis on translation into clinical scheduling practice. A PRISMA-compliant search of PubMed, Embase, IEEE Xplore, Scopus, and Web of Science identified studies using ML for no-show prediction in outpatient settings. Data on models, features, performance, and implementation were extracted. Risk of bias was assessed using an adapted PROBAST tool. Thirty-two studies were included. Logistic regression, random forest, and XGBoost were the most commonly used models. Historical attendance data was the dominant predictive feature. Fewer than 20% of studies reported real-world implementation, and reported intervention outcomes (e.g., overbooking, reminders) were inconsistent. While ML models show strong predictive performance, real-world deployment and evidence of operational impact remain limited. This gap highlights the need to prioritize implementation-focused research to translate predictive accuracy into measurable improvements in clinic efficiency and access.
Alzheimer’s disease (AD) is the leading cause of dementia, affecting over 50 million people worldwide, with prevalence expected to triple by 2050. Early detection is crucial for clinical trial enrollment and care planning, and multimodal data (MRI, PET, CSF biomarkers, and cognitive assessments) provides complementary information on neurodegeneration, metabolism, and protein aggregation. This systematic review synthesizes AI/ML approaches for early AD detection using multimodal data, focusing on fusion strategies and performance across disease stages. Following PRISMA guidelines, searches of PubMed, IEEE Xplore, Scopus, Web of Science, and arXiv (2017–2023) identified studies using ML/DL with at least two modalities and reporting diagnostic performance. From 1,247 records, 35 studies were included. MRI was the most used modality (>90%), followed by cognitive tests (70–80%), PET (40–50%), and CSF (20–30%). Early fusion was most common, with increasing use of intermediate fusion. Multimodal models achieved AUROC of 0.90–0.98 for AD vs controls, but lower performance (0.70–0.85) for predicting MCI conversion to AD. Overall, multimodal AI improves early AD detection, with strong performance for diagnosis but persistent challenges in forecasting MCI progression due to heterogeneity and limited longitudinal data.
The integration of artificial intelligence into clinical decision support systems offers improved diagnostic accuracy and efficiency, but the opacity of many machine learning models raises concerns about trust, accountability, and regulatory compliance. Explainable artificial intelligence (XAI) has been proposed to address this by making model predictions interpretable to clinicians; however, its true clinical value remains uncertain, and evaluation has not kept pace with methodological development. This systematic review aimed to identify XAI methods used in clinical decision support systems, assess how they are evaluated with clinicians, and determine whether explanations improve diagnostic accuracy, trust, mental models, and efficiency. Following PRISMA guidelines, we searched PubMed, Web of Science, IEEE Xplore, ACM Digital Library, and Scopus for studies published between 2017 and 2024. Eligible studies included original research evaluating XAI in clinical decision support systems with clinician participants and reporting quantitative or qualitative outcomes. Risk of bias was assessed using adapted QUADAS-2 and ROBIS tools, and findings were synthesized narratively with subgroup analyses. From 2,847 records, 68 studies were included. The most common XAI methods were SHAP-based feature attribution (38%), saliency or heatmap methods (29%), concept-based approaches such as TCAV (15%), and counterfactual or example-based explanations (12%). Radiology was the dominant field (54%), followed by dermatology (18%) and pathology (12%). Evaluation approaches were highly inconsistent, with few validated instruments and most studies relying on Likert-scale trust measures or qualitative feedback. Only 16% of studies showed improved diagnostic accuracy with explanations, 67% showed no significant effect, and 17% reported reduced accuracy due to over-reliance or misinterpretation. Although 82% of studies reported increased clinician trust, trust rarely correlated with actual diagnostic performance. Overall, while XAI methods are widely studied in clinical decision support, their evaluation is inconsistent and their benefits are limited. Explanations tend to increase clinician trust without reliably improving diagnostic accuracy, and may sometimes worsen performance, highlighting a trust–accuracy gap that poses important safety concerns for clinical deployment.
Postoperative complications including SSI (2–20%), VTE (1–5%), and respiratory failure (1–8%) significantly increase morbidity, mortality, length of stay, and readmissions. This systematic review assessed machine learning models predicting these outcomes, their performance, external validation, and clinical deployment. A PRISMA-based search (2017–2024) identified 32 eligible studies. Models such as random forest and XGBoost showed AUROC ranges of 0.70–0.85 for SSI, 0.75–0.90 for VTE (outperforming Caprini scores), and 0.75–0.88 for respiratory failure. However, fewer than 20% of studies included external validation and less than 5% reported clinical deployment. Overall, while machine learning models show strong retrospective performance, limited validation and minimal real-world implementation remain major barriers to clinical translation.
Suicidality and depression are major global health burdens, with over 700,000 suicide deaths annually and ~280 million people affected by major depressive disorder. Early risk prediction could support prevention, but traditional methods show limited accuracy. This PRISMA-compliant systematic review evaluated machine learning models for predicting suicidality and depression across electronic health records, social media, and wearable sensor data, focusing on performance, unimodal vs multimodal approaches, and ethical reporting. Searches of PubMed, PsycINFO, IEEE Xplore, arXiv, and ACM Digital Library identified eligible studies. EHR-based models showed AUROC 0.70–0.85 for suicide attempt prediction, social media models 0.70–0.80 for suicidal ideation, and wearable sensor models lower performance (0.65–0.75). Multimodal approaches improved performance by 5–10% over unimodal models. However, fewer than 20% of studies reported ethical considerations such as privacy, bias, or deployment safeguards. Overall, machine learning shows moderate-to-good predictive performance, with multimodal models performing best, but ethical reporting remains critically insufficient for clinical translation.
Large language models (LLMs) have rapidly advanced since the transformer architecture was introduced in 2017, with systems such as GPT-3, GPT-4, Med-PaLM, and Claude increasingly explored for applications in medical education, clinical documentation, decision support, and patient communication, raising both optimism and concerns regarding safety and reliability. This systematic review synthesizes evidence across studies retrieved from PubMed, arXiv, ACL Anthology, IEEE Xplore, and Google Scholar that empirically evaluated LLMs in clinical settings using quantitative performance metrics, with risk of bias assessed using an adapted PROBAST framework for machine learning research. Findings show that LLMs achieve 60–90% accuracy on USMLE-style examinations, with leading models such as GPT-4 and Med-PaLM 2 reaching or surpassing passing thresholds, while in clinical documentation tasks they can reduce physician workload by approximately 30–50% in generating outputs such as discharge summaries, though human review remains consistently required. Performance in clinical decision support is more variable and specialty-dependent, and hallucination rates ranging from 5–30% have been reported, alongside persistent issues of bias and overconfidence in incorrect outputs. Overall, while LLMs demonstrate strong capabilities in structured medical knowledge tasks and documentation support, current limitations including hallucinations, bias, and lack of prospective clinical validation prevent safe autonomous deployment, making clinician oversight and robust safety safeguards essential for any clinical use.
Oncology drug development is an expensive and high-failure process, with costs exceeding two billion dollars per approved drug and success rates below 10%. Deep learning has recently been explored as a strategy to improve efficiency across the drug discovery pipeline. This systematic review evaluates its application in target identification, compound screening and de novo drug design, and clinical trial optimization. Following PRISMA 2020 guidelines, multiple databases were searched and studies were screened using predefined inclusion criteria, with risk of bias assessed via established tools. The literature shows that graph neural networks and transformer-based models are the most widely used architectures, particularly in early-stage discovery tasks. Although many studies report strong in silico performance, often with AUC values above 0.80, only a small proportion demonstrate experimental or clinical validation. Overall, deep learning significantly advances computational drug discovery in oncology, but translation into clinically validated therapies remains limited, especially in trial optimization, highlighting the need for stronger prospective and experimental validation frameworks.
Sepsis continues to be a major contributor to morbidity and mortality among hospitalized patients globally, especially within intensive care and emergency departments, where rapid recognition is essential for improving survival through timely treatment. In recent years, machine learning approaches have gained attention for their ability to predict sepsis onset using routinely collected electronic health record data. This systematic review, conducted in accordance with PRISMA 2020 guidelines, synthesizes evidence from studies published between 2017 and 2025, focusing on model architectures, feature selection and engineering strategies, prediction time horizons, and validation methodologies. Searches across major biomedical and informatics databases identified 67 eligible studies. The included literature shows that logistic regression, ensemble tree-based algorithms, and deep learning models are most frequently applied for sepsis prediction tasks. However, the majority of studies rely on retrospective datasets with internal validation, while only a limited number incorporate prospective or real-world validation frameworks. Overall, although reported model performance is often strong in retrospective analyses, a consistent decline in accuracy is observed when models are evaluated in real clinical environments. These findings highlight that prospective validation and improved generalizability are still underdeveloped areas, underscoring the need for future research to emphasize real-time deployment and robust external validation before clinical integration.
Sleep disorders, including obstructive sleep apnea, insomnia, restless legs syndrome, narcolepsy, and central sleep apnea, represent a major public health burden. Polysomnography is the diagnostic gold standard but is resource-intensive, leading to increasing use of home sleep apnea testing and wearable devices to improve accessibility. This systematic review evaluates deep learning models in sleep medicine across polysomnography, home sleep apnea testing, and wearable data, focusing on architectures, signal types, validation approaches, diagnostic tasks, and clinical readiness. A PRISMA 2020–compliant search was conducted in PubMed, IEEE Xplore, Scopus, and Web of Science for studies published from 2017 to 2025, including those applying deep learning for sleep staging, apnea/hypopnea detection, or sleep disorder diagnosis using PSG, HSAT, or wearable-derived signals. Twenty-nine studies were included. Convolutional neural networks were the most widely used architecture, often combined with recurrent or hybrid models for temporal dependencies, while transformer-based models have recently emerged for long-sequence sleep analysis. Deep learning methods demonstrate strong performance in sleep staging and respiratory event detection, especially using polysomnography data. However, limited external validation, heterogeneous datasets, and a lack of prospective clinical deployment remain major barriers to clinical translation.
Rare diseases are challenging for AI development due to sparse patient populations, fragmented expertise, and strong inter-site variability, making federated learning a promising privacy-preserving solution for multi-institutional model training. This systematic review evaluates federated learning approaches for rare disease diagnosis and related data-scarce clinical settings, with emphasis on handling extreme data scarcity, class imbalance, heterogeneity, and privacy constraints. A PRISMA 2020-compliant search of PubMed, IEEE Xplore, Scopus, Web of Science, and arXiv (2017–2025) identified 2,015 records, with 56 studies included after screening. The most commonly used strategies included FedProx-based optimization, personalized federated learning, class-aware aggregation, generative data augmentation, and domain adaptation techniques. Overall, standard federated averaging is often insufficient under severe scarcity and distribution shift, while hybrid approaches combining personalization, augmentation, and domain adaptation show greater promise for improving performance in rare disease applications.
Public health emergencies reveal critical weaknesses in healthcare supply chains, especially when PPE demand outpaces procurement and distribution capacity, making predictive analytics an important tool for forecasting demand and improving allocation during crises. This systematic review evaluates predictive analytics models for PPE demand forecasting and distribution optimization during public health emergencies, focusing on model types, data sources, validation approaches, performance metrics, equity considerations, and implementation readiness. Following PRISMA 2020 guidelines, searches were conducted in PubMed, Web of Science, Scopus, IEEE Xplore, and Google Scholar for studies published between 2017 and 2025, yielding 2,847 records, of which 35 met inclusion criteria. Included studies comprised time series and statistical models (34%), machine learning and hybrid approaches (29%), optimization methods (26%), and simulation or digital twin frameworks (11%), with limited evidence of real-world deployment. Overall, findings indicate that predictive analytics can enhance PPE supply chain resilience by improving demand forecasting, allocation decisions, and scenario testing, but widespread adoption is limited by poor data interoperability, insufficient prospective validation, weak equity integration, and limited operational integration into healthcare decision systems.
Generative artificial intelligence (AI), including GANs, VAEs, and diffusion models, is increasingly used for synthesizing and enhancing medical images, helping address challenges such as limited data, expensive acquisition, and rare disease representation. This systematic review examines studies on generative AI methods for MRI, CT, X-ray, and pathology image synthesis from 2017 to 2026, focusing on synthesis tasks, evaluation strategies, and clinical utility. A PRISMA 2020-compliant search of PubMed, IEEE Xplore, Scopus, and Web of Science identified peer-reviewed research on generative models for medical image synthesis, augmentation, harmonization, or cross-modality translation. Findings show a shift from GAN-based methods to diffusion models post-2022, with MRI and CT studies emphasizing cross-modality translation, and X-ray and pathology studies focusing on augmentation and diagnostic utility. Despite GANs' continued dominance, diffusion models are gaining traction for improving image fidelity and diversity. However, evaluation practices remain inconsistent, with limited inclusion of clinically relevant assessments. This review follows PRISMA 2020 guidelines and provides a narrative synthesis of the evidence.
Clinical trial recruitment is hindered by slow, costly, and labor-intensive processes, particularly due to the complexity of eligibility criteria often written in free text. This systematic review examines the use of large language models (LLMs) for matching clinical trial eligibility criteria to electronic health records (EHR). It evaluates zero-shot, few-shot, and fine-tuned LLM approaches, comparing their strengths, limitations, and deployment readiness in supporting patient-trial matching. Thirty-three studies published from 2017 to 2026 were included, with findings showing that zero-shot prompting is most adaptable for simple criteria, few-shot prompting offers consistent reasoning for ambiguous criteria, and fine-tuned models excel in task-specific performance but require labeled data and are less portable. The review concludes that no single approach is optimal for all trial screening tasks, and hybrid workflows combining various methods with human verification are most suitable for clinical use.
This systematic review examines the use of edge artificial intelligence (AI) and wearable sensors for real-time patient monitoring in smart hospitals and home settings, focusing on detecting deterioration, falls, arrhythmias, and infection-related changes. The review synthesizes studies from 2017 to 2026 on edge AI architectures, wearable sensor fusion, and clinical alert systems, emphasizing latency, power constraints, alert performance, and integration into clinical workflows. A PRISMA 2020-compliant search identified 127 studies from 2,100 records, with findings showing that while edge AI execution grew post-2020, it still represented a minority of designs. Sensor fusion was often linked to broader event coverage but increased implementation complexity. The review concludes that edge AI can reduce latency and enhance privacy but introduces challenges related to power usage, model complexity, device reliability, and maintenance, with limited clinical validation of alert systems and few studies addressing alert fatigue or clinician response.