The integration of artificial intelligence (AI) into healthcare systems marks a fundamental shift from isolated predictive analytics tools to embedded, scalable architectures that support autonomous governance. This narrative review synthesizes 28 peer-reviewed publications from leading journals to examine AI’s role across healthcare infrastructure and clinical analytics. Early work established deep learning foundations for risk prediction, diagnostic support, and prognostic modelling using multimodal data. These capabilities rapidly evolved into system-level applications that enhance data ingestion, real-time inference, and operational optimisation across entire care ecosystems.By the early 2020s, attention turned to deployment realities, including clinician acceptance, cost-effectiveness, and integration into existing workflows. Frameworks for responsible implementation emerged alongside regulatory perspectives that emphasise safety, equity, and continuous oversight. Recent contributions highlight the transition toward closed-loop systems in which predictive outputs inform decisions, trigger interventions, and feed outcome data back for model recalibration. Governance architectures now address ethical challenges, explainability gaps, and the move from generalist to specialised medical AI.This review organises the literature through an original systems-level lens spanning four interconnected pillars—data foundations, analytic intelligence, deployment mechanisms, and governance layers—rather than replicating prior application-specific taxonomies. Cross-study analysis reveals consistent patterns: predictive analytics serve as the foundational engine, clinical decision support acts as the execution layer, closed-loop feedback enables adaptation, and governance ensures sustainable autonomy. The synthesis demonstrates that AI is no longer an adjunct technology but a core infrastructural element reshaping how healthcare systems ingest, process, act upon, and learn from data at scale.Trajectory as a coherent progression toward autonomous yet human-centred governance, the review provides clinicians, system architects, and policymakers with a unified understanding of current capabilities and the infrastructural requirements for responsible scaling.
Foundation models, characterized by their large-scale pretraining on diverse datasets, represent a transformative paradigm in artificial intelligence (AI) applications for healthcare systems and analytics. These models, often based on transformer architectures, enable generalist capabilities that extend beyond narrow task-specific AI, facilitating integration into complex healthcare infrastructures. This review synthesizes recent literature on the architectural integration of foundation models into healthcare systems, emphasizing their role in enhancing clinical analytics, decision support, and operational efficiency while addressing critical oversight considerations, including ethical, regulatory, and safety frameworks.In healthcare systems, foundation models are increasingly deployed to process multimodal data streams, including electronic health records (EHRs), medical imaging, and real-time patient monitoring. Architectural integration involves embedding these models within hospital information systems, enabling seamless data ingestion, inference, and feedback loops. For instance, models like those adapted from large language models (LLMs) support natural language processing for EHR mining, predictive analytics for disease progression, and generative tasks for synthetic data augmentation. Oversight considerations are paramount, encompassing regulatory compliance, bias mitigation, and human-AI collaboration protocols to ensure patient safety and equity.The synthesis highlights key architectural patterns: federated learning for privacy-preserving model training, hybrid human-AI workflows for clinical decision-making, and adaptive systems for continuous model recalibration. Analytics applications span precision medicine, where foundation models integrate genomic and clinical data for personalized interventions, to population health management, optimizing resource allocation through predictive modeling. Ethical oversight includes checklists for AI deployment in low- and middle-income countries (LMICs), emphasizing equitable access and cultural adaptability.Challenges in integration include data interoperability, model interpretability, and scalability in resource-constrained settings. Regulatory imperatives call for validation frameworks and safety standards to govern the rollout of generative AI. This review provides an original systems-level framing, structuring the discourse around data-to-decision pipelines, governance overlays, and evaluative metrics for sustainable adoption.Ultimately, foundation models hold promise for closed-loop healthcare systems, where AI-driven insights inform interventions and feedback refines models iteratively. However, rigorous oversight is essential to balance innovation with accountability, ensuring these technologies augment rather than disrupt clinical workflows. By synthesizing high-impact publications, this narrative review offers integrative insights for researchers, clinicians, and policymakers navigating AI-enabled healthcare transformation.
Large language models (LLMs) have rapidly advanced since the transformer architecture was introduced in 2017, with systems such as GPT-3, GPT-4, Med-PaLM, and Claude increasingly explored for applications in medical education, clinical documentation, decision support, and patient communication, raising both optimism and concerns regarding safety and reliability. This systematic review synthesizes evidence across studies retrieved from PubMed, arXiv, ACL Anthology, IEEE Xplore, and Google Scholar that empirically evaluated LLMs in clinical settings using quantitative performance metrics, with risk of bias assessed using an adapted PROBAST framework for machine learning research. Findings show that LLMs achieve 60–90% accuracy on USMLE-style examinations, with leading models such as GPT-4 and Med-PaLM 2 reaching or surpassing passing thresholds, while in clinical documentation tasks they can reduce physician workload by approximately 30–50% in generating outputs such as discharge summaries, though human review remains consistently required. Performance in clinical decision support is more variable and specialty-dependent, and hallucination rates ranging from 5–30% have been reported, alongside persistent issues of bias and overconfidence in incorrect outputs. Overall, while LLMs demonstrate strong capabilities in structured medical knowledge tasks and documentation support, current limitations including hallucinations, bias, and lack of prospective clinical validation prevent safe autonomous deployment, making clinician oversight and robust safety safeguards essential for any clinical use.
Oncology drug development is an expensive and high-failure process, with costs exceeding two billion dollars per approved drug and success rates below 10%. Deep learning has recently been explored as a strategy to improve efficiency across the drug discovery pipeline. This systematic review evaluates its application in target identification, compound screening and de novo drug design, and clinical trial optimization. Following PRISMA 2020 guidelines, multiple databases were searched and studies were screened using predefined inclusion criteria, with risk of bias assessed via established tools. The literature shows that graph neural networks and transformer-based models are the most widely used architectures, particularly in early-stage discovery tasks. Although many studies report strong in silico performance, often with AUC values above 0.80, only a small proportion demonstrate experimental or clinical validation. Overall, deep learning significantly advances computational drug discovery in oncology, but translation into clinically validated therapies remains limited, especially in trial optimization, highlighting the need for stronger prospective and experimental validation frameworks.
Sepsis continues to be a major contributor to morbidity and mortality among hospitalized patients globally, especially within intensive care and emergency departments, where rapid recognition is essential for improving survival through timely treatment. In recent years, machine learning approaches have gained attention for their ability to predict sepsis onset using routinely collected electronic health record data. This systematic review, conducted in accordance with PRISMA 2020 guidelines, synthesizes evidence from studies published between 2017 and 2025, focusing on model architectures, feature selection and engineering strategies, prediction time horizons, and validation methodologies. Searches across major biomedical and informatics databases identified 67 eligible studies. The included literature shows that logistic regression, ensemble tree-based algorithms, and deep learning models are most frequently applied for sepsis prediction tasks. However, the majority of studies rely on retrospective datasets with internal validation, while only a limited number incorporate prospective or real-world validation frameworks. Overall, although reported model performance is often strong in retrospective analyses, a consistent decline in accuracy is observed when models are evaluated in real clinical environments. These findings highlight that prospective validation and improved generalizability are still underdeveloped areas, underscoring the need for future research to emphasize real-time deployment and robust external validation before clinical integration.
Synthetic electronic health record (EHR) data generation has emerged as a potential solution to balancing clinical data accessibility with patient privacy, using generative artificial intelligence to simulate tabular, longitudinal, and textual health records without exposing identifiable patient information. This critical review, informed by PRISMA-ScR methodology, examines studies published between 2017 and 2025 focusing on generative models for synthetic EHR creation, with particular attention to privacy risks, data fidelity, downstream task utility, and ethical or regulatory considerations. A total of 67 studies were included after systematic screening, showing a dominance of GAN-based approaches alongside growing use of diffusion models and large language models in recent years, although privacy assessment and benchmarking practices remain inconsistent. Overall, the evidence suggests that while synthetic EHR data can facilitate data sharing, research, and model development, achieving a balance between realism, utility, and privacy remains challenging, as high statistical fidelity does not necessarily translate into clinical usefulness and strong downstream performance does not ensure adequate privacy protection.
Sleep disorders, including obstructive sleep apnea, insomnia, restless legs syndrome, narcolepsy, and central sleep apnea, represent a major public health burden. Polysomnography is the diagnostic gold standard but is resource-intensive, leading to increasing use of home sleep apnea testing and wearable devices to improve accessibility. This systematic review evaluates deep learning models in sleep medicine across polysomnography, home sleep apnea testing, and wearable data, focusing on architectures, signal types, validation approaches, diagnostic tasks, and clinical readiness. A PRISMA 2020–compliant search was conducted in PubMed, IEEE Xplore, Scopus, and Web of Science for studies published from 2017 to 2025, including those applying deep learning for sleep staging, apnea/hypopnea detection, or sleep disorder diagnosis using PSG, HSAT, or wearable-derived signals. Twenty-nine studies were included. Convolutional neural networks were the most widely used architecture, often combined with recurrent or hybrid models for temporal dependencies, while transformer-based models have recently emerged for long-sequence sleep analysis. Deep learning methods demonstrate strong performance in sleep staging and respiratory event detection, especially using polysomnography data. However, limited external validation, heterogeneous datasets, and a lack of prospective clinical deployment remain major barriers to clinical translation.
Rare diseases are challenging for AI development due to sparse patient populations, fragmented expertise, and strong inter-site variability, making federated learning a promising privacy-preserving solution for multi-institutional model training. This systematic review evaluates federated learning approaches for rare disease diagnosis and related data-scarce clinical settings, with emphasis on handling extreme data scarcity, class imbalance, heterogeneity, and privacy constraints. A PRISMA 2020-compliant search of PubMed, IEEE Xplore, Scopus, Web of Science, and arXiv (2017–2025) identified 2,015 records, with 56 studies included after screening. The most commonly used strategies included FedProx-based optimization, personalized federated learning, class-aware aggregation, generative data augmentation, and domain adaptation techniques. Overall, standard federated averaging is often insufficient under severe scarcity and distribution shift, while hybrid approaches combining personalization, augmentation, and domain adaptation show greater promise for improving performance in rare disease applications.
Public health emergencies reveal critical weaknesses in healthcare supply chains, especially when PPE demand outpaces procurement and distribution capacity, making predictive analytics an important tool for forecasting demand and improving allocation during crises. This systematic review evaluates predictive analytics models for PPE demand forecasting and distribution optimization during public health emergencies, focusing on model types, data sources, validation approaches, performance metrics, equity considerations, and implementation readiness. Following PRISMA 2020 guidelines, searches were conducted in PubMed, Web of Science, Scopus, IEEE Xplore, and Google Scholar for studies published between 2017 and 2025, yielding 2,847 records, of which 35 met inclusion criteria. Included studies comprised time series and statistical models (34%), machine learning and hybrid approaches (29%), optimization methods (26%), and simulation or digital twin frameworks (11%), with limited evidence of real-world deployment. Overall, findings indicate that predictive analytics can enhance PPE supply chain resilience by improving demand forecasting, allocation decisions, and scenario testing, but widespread adoption is limited by poor data interoperability, insufficient prospective validation, weak equity integration, and limited operational integration into healthcare decision systems.
Generative artificial intelligence (AI), including GANs, VAEs, and diffusion models, is increasingly used for synthesizing and enhancing medical images, helping address challenges such as limited data, expensive acquisition, and rare disease representation. This systematic review examines studies on generative AI methods for MRI, CT, X-ray, and pathology image synthesis from 2017 to 2026, focusing on synthesis tasks, evaluation strategies, and clinical utility. A PRISMA 2020-compliant search of PubMed, IEEE Xplore, Scopus, and Web of Science identified peer-reviewed research on generative models for medical image synthesis, augmentation, harmonization, or cross-modality translation. Findings show a shift from GAN-based methods to diffusion models post-2022, with MRI and CT studies emphasizing cross-modality translation, and X-ray and pathology studies focusing on augmentation and diagnostic utility. Despite GANs' continued dominance, diffusion models are gaining traction for improving image fidelity and diversity. However, evaluation practices remain inconsistent, with limited inclusion of clinically relevant assessments. This review follows PRISMA 2020 guidelines and provides a narrative synthesis of the evidence.
Clinical trial recruitment is hindered by slow, costly, and labor-intensive processes, particularly due to the complexity of eligibility criteria often written in free text. This systematic review examines the use of large language models (LLMs) for matching clinical trial eligibility criteria to electronic health records (EHR). It evaluates zero-shot, few-shot, and fine-tuned LLM approaches, comparing their strengths, limitations, and deployment readiness in supporting patient-trial matching. Thirty-three studies published from 2017 to 2026 were included, with findings showing that zero-shot prompting is most adaptable for simple criteria, few-shot prompting offers consistent reasoning for ambiguous criteria, and fine-tuned models excel in task-specific performance but require labeled data and are less portable. The review concludes that no single approach is optimal for all trial screening tasks, and hybrid workflows combining various methods with human verification are most suitable for clinical use.
Federated and decentralized machine learning offer the potential to extract valuable healthcare insights from siloed data without requiring the centralization of sensitive patient records, addressing long-standing privacy and governance challenges. This critical review assesses federated learning in healthcare through three lenses: privacy-preserving technologies, incentive mechanisms, and regulatory compliance frameworks. It examines whether the claims in existing literature are substantiated by real-world evidence from healthcare settings. The review reveals considerable enthusiasm for federated learning but identifies gaps, including incomplete implementation of privacy technologies, theoretical incentive mechanisms, and regulatory compliance often assumed but not validated. Additionally, real-world deployments are limited in scale and duration. The review concludes that the gap between federated learning's theoretical potential and clinical application remains significant, with overstated privacy claims and a lack of established frameworks for incentives and compliance.
This systematic review examines the use of edge artificial intelligence (AI) and wearable sensors for real-time patient monitoring in smart hospitals and home settings, focusing on detecting deterioration, falls, arrhythmias, and infection-related changes. The review synthesizes studies from 2017 to 2026 on edge AI architectures, wearable sensor fusion, and clinical alert systems, emphasizing latency, power constraints, alert performance, and integration into clinical workflows. A PRISMA 2020-compliant search identified 127 studies from 2,100 records, with findings showing that while edge AI execution grew post-2020, it still represented a minority of designs. Sensor fusion was often linked to broader event coverage but increased implementation complexity. The review concludes that edge AI can reduce latency and enhance privacy but introduces challenges related to power usage, model complexity, device reliability, and maintenance, with limited clinical validation of alert systems and few studies addressing alert fatigue or clinician response.