The integration of artificial intelligence (AI) into healthcare systems marks a fundamental shift from isolated predictive analytics tools to embedded, scalable architectures that support autonomous governance. This narrative review synthesizes 28 peer-reviewed publications from leading journals to examine AI’s role across healthcare infrastructure and clinical analytics. Early work established deep learning foundations for risk prediction, diagnostic support, and prognostic modelling using multimodal data. These capabilities rapidly evolved into system-level applications that enhance data ingestion, real-time inference, and operational optimisation across entire care ecosystems.By the early 2020s, attention turned to deployment realities, including clinician acceptance, cost-effectiveness, and integration into existing workflows. Frameworks for responsible implementation emerged alongside regulatory perspectives that emphasise safety, equity, and continuous oversight. Recent contributions highlight the transition toward closed-loop systems in which predictive outputs inform decisions, trigger interventions, and feed outcome data back for model recalibration. Governance architectures now address ethical challenges, explainability gaps, and the move from generalist to specialised medical AI.This review organises the literature through an original systems-level lens spanning four interconnected pillars—data foundations, analytic intelligence, deployment mechanisms, and governance layers—rather than replicating prior application-specific taxonomies. Cross-study analysis reveals consistent patterns: predictive analytics serve as the foundational engine, clinical decision support acts as the execution layer, closed-loop feedback enables adaptation, and governance ensures sustainable autonomy. The synthesis demonstrates that AI is no longer an adjunct technology but a core infrastructural element reshaping how healthcare systems ingest, process, act upon, and learn from data at scale.Trajectory as a coherent progression toward autonomous yet human-centred governance, the review provides clinicians, system architects, and policymakers with a unified understanding of current capabilities and the infrastructural requirements for responsible scaling.
Foundation models, characterized by their large-scale pretraining on diverse datasets, represent a transformative paradigm in artificial intelligence (AI) applications for healthcare systems and analytics. These models, often based on transformer architectures, enable generalist capabilities that extend beyond narrow task-specific AI, facilitating integration into complex healthcare infrastructures. This review synthesizes recent literature on the architectural integration of foundation models into healthcare systems, emphasizing their role in enhancing clinical analytics, decision support, and operational efficiency while addressing critical oversight considerations, including ethical, regulatory, and safety frameworks.In healthcare systems, foundation models are increasingly deployed to process multimodal data streams, including electronic health records (EHRs), medical imaging, and real-time patient monitoring. Architectural integration involves embedding these models within hospital information systems, enabling seamless data ingestion, inference, and feedback loops. For instance, models like those adapted from large language models (LLMs) support natural language processing for EHR mining, predictive analytics for disease progression, and generative tasks for synthetic data augmentation. Oversight considerations are paramount, encompassing regulatory compliance, bias mitigation, and human-AI collaboration protocols to ensure patient safety and equity.The synthesis highlights key architectural patterns: federated learning for privacy-preserving model training, hybrid human-AI workflows for clinical decision-making, and adaptive systems for continuous model recalibration. Analytics applications span precision medicine, where foundation models integrate genomic and clinical data for personalized interventions, to population health management, optimizing resource allocation through predictive modeling. Ethical oversight includes checklists for AI deployment in low- and middle-income countries (LMICs), emphasizing equitable access and cultural adaptability.Challenges in integration include data interoperability, model interpretability, and scalability in resource-constrained settings. Regulatory imperatives call for validation frameworks and safety standards to govern the rollout of generative AI. This review provides an original systems-level framing, structuring the discourse around data-to-decision pipelines, governance overlays, and evaluative metrics for sustainable adoption.Ultimately, foundation models hold promise for closed-loop healthcare systems, where AI-driven insights inform interventions and feedback refines models iteratively. However, rigorous oversight is essential to balance innovation with accountability, ensuring these technologies augment rather than disrupt clinical workflows. By synthesizing high-impact publications, this narrative review offers integrative insights for researchers, clinicians, and policymakers navigating AI-enabled healthcare transformation.
Multi-agent systems (MAS) represent a paradigm shift in artificial intelligence applications for healthcare operations, enabling distributed, autonomous entities to collaborate in complex environments characterized by uncertainty, heterogeneity, and real-time demands. This narrative review synthesizes recent advancements in MAS for healthcare systems and analytics, focusing on coordination theory, safety constraints, and implementation considerations. We examine how MAS facilitates intelligent coordination among agents—such as AI models, human clinicians, and IoT devices—to optimize operational workflows, enhance clinical decision-making, and ensure patient safety. Coordination theory in MAS underscores the mechanisms for agent interaction, including negotiation protocols, consensus algorithms, and hierarchical structures, which are critical for synchronizing tasks in healthcare settings like emergency response and chronic disease management. For instance, MAS enables adaptive resource allocation in hospitals by modeling agents as decision-makers that negotiate bed assignments or staff scheduling based on real-time data inputs. Safety constraints emerge as a pivotal concern, encompassing formal verification methods, fault-tolerant designs, and ethical safeguards to mitigate risks such as erroneous agent decisions leading to adverse patient outcomes. Implementation considerations address scalability, interoperability with legacy systems, and regulatory compliance, highlighting challenges in deploying MAS in fog-cloud architectures for remote monitoring. The review integrates a systems-level perspective, illustrating how MAS evolve from isolated AI tools to interconnected ecosystems that support closed-loop healthcare processes—from data acquisition to intervention feedback. We propose an original interpretive framework that structures MAS across layers: perceptual (data sensing), cognitive (analytics and decision fusion), coordinative (agent interaction), and governance (safety and oversight). This framework reveals cross-study insights, such as the role of large language models (LLMs) in augmenting agent rationality and the integration of digital twins for simulation-based safety testing. Comparative analysis shows that while MAS excel in dynamic environments like cardiology case retrieval or pain management, persistent gaps in standardization hinder widespread adoption. By synthesizing these elements, the review offers novel insights into MAS as enablers of resilient healthcare infrastructure, emphasizing the need for hybrid human-AI coordination to balance autonomy with oversight. Future implications include advancing MAS toward predictive analytics in personalized medicine, with recommendations for interdisciplinary research to address implementation barriers. Ultimately, this work advocates for MAS as foundational to next-generation healthcare analytics, promoting efficiency, equity, and safety in operational contexts.
The integration of social determinants of health (SDoH) into artificial intelligence (AI) systems for healthcare represents a pivotal advancement in addressing inequities within clinical analytics and decision-making frameworks. SDoH encompass socioeconomic, environmental, and behavioral factors that profoundly influence health outcomes, yet their incorporation into AI models has been inconsistent, often exacerbating biases rather than mitigating them. This narrative review synthesizes recent literature on strategies for embedding SDoH data into AI pipelines, elucidates mechanisms of bias propagation, and evaluates approaches to equity assessment in healthcare systems. Drawing from peer-reviewed publications, we highlight the evolution of AI applications in healthcare analytics, where machine learning algorithms increasingly process electronic health records (EHRs), wearable data, and population-level datasets to predict risks and optimize interventions. However, without deliberate integration of SDoH, these systems risk perpetuating disparities, as evidenced by models that underperform for underrepresented groups due to skewed training data. Integration strategies range from data augmentation techniques, such as linking EHRs with geospatial SDoH indices, to hybrid modeling approaches that fuse clinical variables with socioeconomic proxies. For instance, federated learning frameworks enable cross-institutional data sharing while preserving privacy, facilitating broader SDoH representation. Bias mechanisms are multifaceted, including selection bias from non-diverse datasets, algorithmic amplification of historical inequities, and deployment biases in real-world settings where AI outputs influence resource allocation. Studies demonstrate how unaddressed confounders, like zip code-based proxies for race or income, can lead to discriminatory predictions in areas such as readmission risk or treatment recommendations. Equity evaluation methodologies emphasize fairness metrics, such as demographic parity and equalized odds, adapted for healthcare contexts. Prospective audits, involving diverse stakeholder input, are recommended to assess model performance across SDoH strata. Consensus emerges on the need for governance structures that incorporate ethical AI principles, including transparency in SDoH feature engineering and continuous monitoring for drift. Challenges persist in standardizing SDoH data collection, with calls for interoperable ontologies to enhance AI generalizability. This review proposes a systems-level framework for SDoH-aware AI, advocating for closed-loop systems that integrate feedback from equity audits into model retraining cycles. Ultimately, advancing SDoH integration in healthcare AI requires interdisciplinary collaboration between clinicians, data scientists, and policymakers to foster equitable systems. By prioritizing bias mitigation and equity-centric design, AI can transition from a tool that mirrors societal inequities to one that actively reduces them, promoting health justice in analytics-driven care. Future directions include scalable implementations in low-resource settings and regulatory frameworks to enforce SDoH considerations. This synthesis underscores the transformative potential of SDoH-informed AI while cautioning against unchecked deployment that could widen health gaps.
The integration of artificial intelligence into healthcare systems has transitioned from isolated diagnostic tools to comprehensive workflow automation platforms that fundamentally reshape clinical operations, decision cycles, and accountability frameworks. This narrative review synthesizes studies that examine how AI-driven task modeling, human–AI interaction dynamics, and evolving accountability structures collectively enable scalable, safe, and ethically grounded automation across healthcare analytics and delivery infrastructures. Rather than cataloging isolated applications, the analysis adopts an original systems-level lens that organizes the literature into four interdependent layers—data orchestration, model orchestration, deployment orchestration, and governance orchestration—revealing recurring patterns of closed-loop intelligence that link real-time data ingestion to automated intervention and continuous recalibration. Task modeling emerges as the foundational mechanism through which heterogeneous clinical workflows are decomposed into machine-executable primitives while preserving human oversight at critical decision nodes. Multiple integrative reviews demonstrate that well-designed task ontologies reduce cognitive burden on clinicians by 30%–50% in high-volume settings such as nursing documentation, pathology slide triage, and echocardiographic measurement, yet success critically depends on explicit representation of human factors, including workload, trust calibration, and exception-handling protocols. Human factors literature further highlights the bidirectional influence between automation and clinician performance. While AI scribes and large language model-assisted note generation improve throughput, they simultaneously introduce new forms of automation bias and alert fatigue that must be mitigated through adaptive interface design and real-time transparency mechanisms. Accountability structures constitute the least mature yet most decisive layer of AI-enabled healthcare automation. Governance models that embed continuous human–AI shared liability, audit trails for every automated decision, and dynamic recalibration triggers are shown to be essential for regulatory acceptance and clinical adoption. Studies of real-world deployments in pathology foundation models and closed-loop infection prevention systems illustrate that accountability is not an afterthought but an infrastructural requirement: without traceable lineage from raw data through model inference to clinical action and feedback, organizations cannot fulfill medico-legal or ethical obligations. This review contributes an original integrative framework—the clinical intelligence loop—that formalizes the end-to-end automation architecture as a continuous cycle of data ingestion, task-modeled inference, human-augmented decision fusion, intervention execution, outcome monitoring, and governance-driven recalibration. Cross-study synthesis reveals that systems achieving sustained performance do so by maintaining tight coupling across all five stages rather than optimizing any single component in isolation. The analysis underscores that workflow automation in healthcare AI succeeds only when task modeling is human-centered, human factors are explicitly engineered into the loop, and accountability is infrastructural rather than retrofitted. These insights provide both theoretical scaffolding and practical guidance for health-system leaders, regulators, and technology developers seeking to scale responsible AI automation beyond pilot projects.
The rapid expansion of unstructured narrative data within patient safety event (PSE) reporting systems presents both a valuable source of safety intelligence and a major analytical challenge for healthcare organizations. Traditional manual review processes are labor-intensive, subjective, and incapable of scaling to the vast volumes of incident reports generated across modern health systems. Artificial intelligence techniques, particularly natural language processing and machine learning, provide scalable approaches for extracting meaningful insights from these narratives. This narrative review synthesizes advances in AI-enabled PSE analytics across three interconnected domains: automated narrative mining, data-driven taxonomy development, and integration within learning health systems that transform safety data into continuous improvement cycles. Evidence indicates that AI methods can improve event classification, accelerate detection of emerging safety signals, and reduce the analytical burden on safety teams. However, challenges remain regarding model generalisability, interpretability, and governance. AI-driven narrative analytics is emerging as a foundational component of next-generation safety intelligence infrastructures.
Wearable and mobile sensing technologies are transforming healthcare by enabling continuous monitoring, real-time analytics, and personalized interventions. This narrative review explores recent advances in artificial intelligence (AI)–driven healthcare analytics, focusing on validation frameworks, drift detection, and generalization challenges associated with wearable sensing systems. Modern wearable devices equipped with biosensors capture physiological signals such as heart rate, activity, and stress indicators, while AI algorithms analyze multimodal data to generate actionable clinical insights. Ensuring reliability requires robust validation strategies that address sensor accuracy, data integrity, and clinical relevance in real-world settings. Drift detection methods help maintain model performance despite environmental changes and user variability. At the same time, generalization techniques support reliable deployment across diverse populations and clinical contexts, advancing scalable and adaptive digital healthcare systems.
Large language models (LLMs) have rapidly advanced since the transformer architecture was introduced in 2017, with systems such as GPT-3, GPT-4, Med-PaLM, and Claude increasingly explored for applications in medical education, clinical documentation, decision support, and patient communication, raising both optimism and concerns regarding safety and reliability. This systematic review synthesizes evidence across studies retrieved from PubMed, arXiv, ACL Anthology, IEEE Xplore, and Google Scholar that empirically evaluated LLMs in clinical settings using quantitative performance metrics, with risk of bias assessed using an adapted PROBAST framework for machine learning research. Findings show that LLMs achieve 60–90% accuracy on USMLE-style examinations, with leading models such as GPT-4 and Med-PaLM 2 reaching or surpassing passing thresholds, while in clinical documentation tasks they can reduce physician workload by approximately 30–50% in generating outputs such as discharge summaries, though human review remains consistently required. Performance in clinical decision support is more variable and specialty-dependent, and hallucination rates ranging from 5–30% have been reported, alongside persistent issues of bias and overconfidence in incorrect outputs. Overall, while LLMs demonstrate strong capabilities in structured medical knowledge tasks and documentation support, current limitations including hallucinations, bias, and lack of prospective clinical validation prevent safe autonomous deployment, making clinician oversight and robust safety safeguards essential for any clinical use.
Oncology drug development is an expensive and high-failure process, with costs exceeding two billion dollars per approved drug and success rates below 10%. Deep learning has recently been explored as a strategy to improve efficiency across the drug discovery pipeline. This systematic review evaluates its application in target identification, compound screening and de novo drug design, and clinical trial optimization. Following PRISMA 2020 guidelines, multiple databases were searched and studies were screened using predefined inclusion criteria, with risk of bias assessed via established tools. The literature shows that graph neural networks and transformer-based models are the most widely used architectures, particularly in early-stage discovery tasks. Although many studies report strong in silico performance, often with AUC values above 0.80, only a small proportion demonstrate experimental or clinical validation. Overall, deep learning significantly advances computational drug discovery in oncology, but translation into clinically validated therapies remains limited, especially in trial optimization, highlighting the need for stronger prospective and experimental validation frameworks.
Sepsis continues to be a major contributor to morbidity and mortality among hospitalized patients globally, especially within intensive care and emergency departments, where rapid recognition is essential for improving survival through timely treatment. In recent years, machine learning approaches have gained attention for their ability to predict sepsis onset using routinely collected electronic health record data. This systematic review, conducted in accordance with PRISMA 2020 guidelines, synthesizes evidence from studies published between 2017 and 2025, focusing on model architectures, feature selection and engineering strategies, prediction time horizons, and validation methodologies. Searches across major biomedical and informatics databases identified 67 eligible studies. The included literature shows that logistic regression, ensemble tree-based algorithms, and deep learning models are most frequently applied for sepsis prediction tasks. However, the majority of studies rely on retrospective datasets with internal validation, while only a limited number incorporate prospective or real-world validation frameworks. Overall, although reported model performance is often strong in retrospective analyses, a consistent decline in accuracy is observed when models are evaluated in real clinical environments. These findings highlight that prospective validation and improved generalizability are still underdeveloped areas, underscoring the need for future research to emphasize real-time deployment and robust external validation before clinical integration.
Synthetic electronic health record (EHR) data generation has emerged as a potential solution to balancing clinical data accessibility with patient privacy, using generative artificial intelligence to simulate tabular, longitudinal, and textual health records without exposing identifiable patient information. This critical review, informed by PRISMA-ScR methodology, examines studies published between 2017 and 2025 focusing on generative models for synthetic EHR creation, with particular attention to privacy risks, data fidelity, downstream task utility, and ethical or regulatory considerations. A total of 67 studies were included after systematic screening, showing a dominance of GAN-based approaches alongside growing use of diffusion models and large language models in recent years, although privacy assessment and benchmarking practices remain inconsistent. Overall, the evidence suggests that while synthetic EHR data can facilitate data sharing, research, and model development, achieving a balance between realism, utility, and privacy remains challenging, as high statistical fidelity does not necessarily translate into clinical usefulness and strong downstream performance does not ensure adequate privacy protection.
Sleep disorders, including obstructive sleep apnea, insomnia, restless legs syndrome, narcolepsy, and central sleep apnea, represent a major public health burden. Polysomnography is the diagnostic gold standard but is resource-intensive, leading to increasing use of home sleep apnea testing and wearable devices to improve accessibility. This systematic review evaluates deep learning models in sleep medicine across polysomnography, home sleep apnea testing, and wearable data, focusing on architectures, signal types, validation approaches, diagnostic tasks, and clinical readiness. A PRISMA 2020–compliant search was conducted in PubMed, IEEE Xplore, Scopus, and Web of Science for studies published from 2017 to 2025, including those applying deep learning for sleep staging, apnea/hypopnea detection, or sleep disorder diagnosis using PSG, HSAT, or wearable-derived signals. Twenty-nine studies were included. Convolutional neural networks were the most widely used architecture, often combined with recurrent or hybrid models for temporal dependencies, while transformer-based models have recently emerged for long-sequence sleep analysis. Deep learning methods demonstrate strong performance in sleep staging and respiratory event detection, especially using polysomnography data. However, limited external validation, heterogeneous datasets, and a lack of prospective clinical deployment remain major barriers to clinical translation.
Rare diseases are challenging for AI development due to sparse patient populations, fragmented expertise, and strong inter-site variability, making federated learning a promising privacy-preserving solution for multi-institutional model training. This systematic review evaluates federated learning approaches for rare disease diagnosis and related data-scarce clinical settings, with emphasis on handling extreme data scarcity, class imbalance, heterogeneity, and privacy constraints. A PRISMA 2020-compliant search of PubMed, IEEE Xplore, Scopus, Web of Science, and arXiv (2017–2025) identified 2,015 records, with 56 studies included after screening. The most commonly used strategies included FedProx-based optimization, personalized federated learning, class-aware aggregation, generative data augmentation, and domain adaptation techniques. Overall, standard federated averaging is often insufficient under severe scarcity and distribution shift, while hybrid approaches combining personalization, augmentation, and domain adaptation show greater promise for improving performance in rare disease applications.
Public health emergencies reveal critical weaknesses in healthcare supply chains, especially when PPE demand outpaces procurement and distribution capacity, making predictive analytics an important tool for forecasting demand and improving allocation during crises. This systematic review evaluates predictive analytics models for PPE demand forecasting and distribution optimization during public health emergencies, focusing on model types, data sources, validation approaches, performance metrics, equity considerations, and implementation readiness. Following PRISMA 2020 guidelines, searches were conducted in PubMed, Web of Science, Scopus, IEEE Xplore, and Google Scholar for studies published between 2017 and 2025, yielding 2,847 records, of which 35 met inclusion criteria. Included studies comprised time series and statistical models (34%), machine learning and hybrid approaches (29%), optimization methods (26%), and simulation or digital twin frameworks (11%), with limited evidence of real-world deployment. Overall, findings indicate that predictive analytics can enhance PPE supply chain resilience by improving demand forecasting, allocation decisions, and scenario testing, but widespread adoption is limited by poor data interoperability, insufficient prospective validation, weak equity integration, and limited operational integration into healthcare decision systems.
Generative artificial intelligence (AI), including GANs, VAEs, and diffusion models, is increasingly used for synthesizing and enhancing medical images, helping address challenges such as limited data, expensive acquisition, and rare disease representation. This systematic review examines studies on generative AI methods for MRI, CT, X-ray, and pathology image synthesis from 2017 to 2026, focusing on synthesis tasks, evaluation strategies, and clinical utility. A PRISMA 2020-compliant search of PubMed, IEEE Xplore, Scopus, and Web of Science identified peer-reviewed research on generative models for medical image synthesis, augmentation, harmonization, or cross-modality translation. Findings show a shift from GAN-based methods to diffusion models post-2022, with MRI and CT studies emphasizing cross-modality translation, and X-ray and pathology studies focusing on augmentation and diagnostic utility. Despite GANs' continued dominance, diffusion models are gaining traction for improving image fidelity and diversity. However, evaluation practices remain inconsistent, with limited inclusion of clinically relevant assessments. This review follows PRISMA 2020 guidelines and provides a narrative synthesis of the evidence.
Clinical trial recruitment is hindered by slow, costly, and labor-intensive processes, particularly due to the complexity of eligibility criteria often written in free text. This systematic review examines the use of large language models (LLMs) for matching clinical trial eligibility criteria to electronic health records (EHR). It evaluates zero-shot, few-shot, and fine-tuned LLM approaches, comparing their strengths, limitations, and deployment readiness in supporting patient-trial matching. Thirty-three studies published from 2017 to 2026 were included, with findings showing that zero-shot prompting is most adaptable for simple criteria, few-shot prompting offers consistent reasoning for ambiguous criteria, and fine-tuned models excel in task-specific performance but require labeled data and are less portable. The review concludes that no single approach is optimal for all trial screening tasks, and hybrid workflows combining various methods with human verification are most suitable for clinical use.
Federated and decentralized machine learning offer the potential to extract valuable healthcare insights from siloed data without requiring the centralization of sensitive patient records, addressing long-standing privacy and governance challenges. This critical review assesses federated learning in healthcare through three lenses: privacy-preserving technologies, incentive mechanisms, and regulatory compliance frameworks. It examines whether the claims in existing literature are substantiated by real-world evidence from healthcare settings. The review reveals considerable enthusiasm for federated learning but identifies gaps, including incomplete implementation of privacy technologies, theoretical incentive mechanisms, and regulatory compliance often assumed but not validated. Additionally, real-world deployments are limited in scale and duration. The review concludes that the gap between federated learning's theoretical potential and clinical application remains significant, with overstated privacy claims and a lack of established frameworks for incentives and compliance.
This systematic review examines the use of edge artificial intelligence (AI) and wearable sensors for real-time patient monitoring in smart hospitals and home settings, focusing on detecting deterioration, falls, arrhythmias, and infection-related changes. The review synthesizes studies from 2017 to 2026 on edge AI architectures, wearable sensor fusion, and clinical alert systems, emphasizing latency, power constraints, alert performance, and integration into clinical workflows. A PRISMA 2020-compliant search identified 127 studies from 2,100 records, with findings showing that while edge AI execution grew post-2020, it still represented a minority of designs. Sensor fusion was often linked to broader event coverage but increased implementation complexity. The review concludes that edge AI can reduce latency and enhance privacy but introduces challenges related to power usage, model complexity, device reliability, and maintenance, with limited clinical validation of alert systems and few studies addressing alert fatigue or clinician response.