The integration of artificial intelligence (AI) into healthcare systems has revolutionized clinical analytics, enabling enhanced diagnostic accuracy, predictive modeling, and personalized treatment pathways. However, the opacity of many AI models poses significant challenges to their clinical adoption, necessitating advancements in explainable AI (XAI) to ensure interpretability and transparency. This narrative review synthesizes the literature on XAI within clinical systems, focusing on interpretability mechanisms, transparency frameworks, and deployment constraints in healthcare analytics. Drawing from high-impact studies, we examine how XAI addresses the “black box” nature of machine learning models in high-stakes medical decisions, particularly in contexts where performance has traditionally been prioritized over explainability. Key themes include the shift toward inherently interpretable models for critical applications, such as diagnostic imaging and predictive analytics, where post-hoc explanations often fall short. We explore the ethical imperatives for responsible AI deployment, including strategies for mitigating harm through transparent systems that align with clinical workflows. The review integrates perspectives on XAI in clinical diagnostics, emphasizing challenges in balancing model complexity with user trust. Transparency is framed not merely as a technical feature but as a systemic requirement, incorporating structured reporting practices for AI interventions and standardized modeling approaches. Deployment constraints are analyzed through the lens of real-world integration, including regulatory considerations, data privacy concerns, and human–AI interaction dynamics in healthcare infrastructures. We synthesize evidence from diverse applications, such as lung cancer diagnosis via explainable models and radiographic assessments, underscoring the need for multidisciplinary approaches to XAI. Furthermore, the review highlights biases in AI systems, particularly sex and gender disparities, and advocates for inclusive analytics to foster equitable healthcare. Clinical applications beyond the black box are discussed, with calls for standardized reporting to enhance reproducibility and trust. We position XAI as essential for closed-loop systems that incorporate feedback mechanisms, ensuring ongoing model recalibration in dynamic clinical environments. The synthesis reveals persistent gaps in current XAI deployments, such as overreliance on surrogate explanations that may mislead clinicians. Ultimately, this review proposes a systems-level framework for XAI in healthcare, integrating data ingestion, inference, decision support, and governance loops to overcome transparency barriers. This comprehensive overview informs the development of future AI-enabled healthcare infrastructures, emphasizing interpretability as a cornerstone for safe and effective clinical analytics.
The integration of multi-modal intelligence in healthcare represents a transformative paradigm, where artificial intelligence (AI) systems synthesize diverse clinical data streams—ranging from electronic health records (EHRs), imaging, genomics, and wearable sensor data—to enable more cohesive, predictive, and actionable insights. This narrative review synthesizes recent advancements in AI for healthcare systems and analytics, focusing on conceptual integration patterns that bridge disparate data modalities to enhance clinical decision-making and system-level efficiencies. We explore how multi-modal AI frameworks address the heterogeneity of healthcare data, fostering intelligent systems that support precision health, risk stratification, and closed-loop interventions. Key themes include the evolution of multi-modal machine learning techniques, such as fusion models that combine radiological imaging with clinical parameters for improved diagnostic accuracy, and the role of large language models (LLMs) in processing unstructured textual data alongside structured metrics. For instance, integrated frameworks leverage deep residual networks and transformers to handle multimodal inputs, enabling applications in areas like pulmonary hypertension prediction and Alzheimer’s disease progression forecasting. We highlight systems-level architectures that incorporate feedback loops for continuous model refinement, emphasizing the need for robust data modeling in federated learning environments to ensure privacy and interoperability across healthcare infrastructures. Challenges in data fusion, such as handling dataset shifts and ensuring equitable access to digital health tools, are contextualized within broader analytics pipelines. The review underscores original synthesis logic by framing integration patterns through a systems lens: data ingestion, intelligent inference, decision support, and governance. This approach reveals how multi-modal AI not only amplifies analytic capabilities but also redefines healthcare delivery models, from virtual biopsies using mammography data to comprehensive communication skills training for physicians via AI-driven video analysis. Ultimately, this synthesis positions multi-modal intelligence as a cornerstone for next-generation healthcare systems, promoting seamless interoperability and human-AI collaboration. By avoiding empirical benchmarks and focusing on conceptual patterns, we provide an interpretive framework that guides future deployments, ensuring AI enhances rather than disrupts clinical workflows.
The integration of artificial intelligence (AI) into healthcare systems has transformed population health analytics, enabling scalable infrastructures that process vast datasets to inform clinical decisions, resource allocation, and policy-making. This narrative review synthesizes recent literature on AI system architectures and governance models, focusing on how these elements underpin analytics-driven healthcare ecosystems. We examine the evolution of AI-enabled infrastructures, emphasizing federated learning, explainable models, and ethical frameworks to address data privacy, interoperability, and equity in population-level analytics. Key architectures include vertically integrated systems that streamline data ingestion, model deployment, and real-time inference, as seen in federated approaches that mitigate data silos while preserving patient confidentiality. Governance models are critical for ensuring trustworthy AI deployment, incorporating regulatory oversight, ethical principles adapted from military contexts to healthcare, and consensus-based guidelines for prediction models. We highlight the role of blockchain and data trusts in enhancing transparency and consent mechanisms, particularly in global health responses to pandemics and chronic disease management. The review structures the discourse around systems-level framing, integrating data flows, algorithmic decision support, and closed-loop feedback mechanisms that adapt to clinical outcomes. For instance, electronic health record (EHR)-based prediction models facilitate acute illness forecasting and outcome prediction in conditions like rheumatoid arthritis and oncology. We propose an original synthesis logic that conceptualizes AI infrastructures as adaptive networks, where governance acts as a regulatory layer overlaying architectural components to balance innovation with risk mitigation. Challenges such as bias in commercial datasets and the need for international cooperation are noted, but the emphasis remains on infrastructural resilience. Ultimately, this synthesis underscores the imperative for hybrid human-AI systems that prioritize population health equity, with governance models evolving to support sustainable analytics infrastructures. By positioning AI as a foundational tool for healthcare transformation, the review advocates for interdisciplinary collaboration to refine these systems, ensuring they deliver actionable insights while upholding ethical standards in diverse healthcare settings.
The integration of artificial intelligence (AI) into healthcare systems has revolutionized clinical analytics, enabling predictive modeling, diagnostic support, and personalized interventions. However, the post-deployment phase of these AI systems presents unique challenges, particularly in maintaining performance amid evolving clinical environments. This narrative review synthesizes recent literature on post-deployment monitoring strategies for clinical AI, focusing on drift detection, feedback governance, and update policies within healthcare systems and analytics frameworks. We examine how data shifts—arising from changes in patient demographics, clinical protocols, or external factors—can degrade AI model efficacy, leading to suboptimal outcomes in high-stakes settings like disease prediction and resource allocation. Drift detection emerges as a cornerstone, encompassing statistical methods to identify concept drift, covariate shift, and label drift in real-time healthcare data streams. Techniques such as nonparametric monitoring and ensemble-based approaches allow for proactive identification of performance decay, ensuring AI systems remain aligned with dynamic clinical realities. Feedback governance integrates human-in-the-loop mechanisms, where clinician inputs refine AI outputs, fostering trust and regulatory compliance in governance structures. Update policies, including retraining schedules and federated learning paradigms, to address the need for iterative model evolution without disrupting clinical workflows. We highlight systems-level perspectives, such as closed-loop architectures that link monitoring to automated updates, emphasizing interoperability across electronic health records (EHRs) and AI pipelines. Comparative analysis reveals gaps in current practices, including limited scalability in resource-constrained settings and ethical considerations in data privacy during monitoring. Through an original synthesis, we propose an integrative framework for AI lifecycle management in healthcare, underscoring the interplay between drift metrics, governance protocols, and policy-driven updates to enhance patient safety and system resilience. This review underscores the imperative for standardized monitoring protocols, informed by multidisciplinary insights, to bridge the translational gap from AI development to sustained clinical utility. By addressing these elements, healthcare AI can achieve robust, adaptive performance, ultimately improving analytics-driven decision-making and outcomes in diverse clinical contexts. Future directions include harmonizing international guidelines for AI monitoring, integrating explainable AI for better feedback loops, and leveraging emerging technologies like edge computing for real-time drift management. This synthesis provides a foundation for researchers and practitioners to advance post-deployment strategies, ensuring AI’s enduring impact on healthcare systems.
The integration of retrieval-augmented generation (RAG) into healthcare systems represents a transformative approach to enhancing the reliability, interpretability, and safety of artificial intelligence (AI)-driven clinical analytics. By combining large language models (LLMs) with external knowledge retrieval mechanisms, RAG mitigates hallucinations inherent in standalone generative models, ensuring outputs are grounded in verifiable evidence from electronic health records (EHRs), clinical guidelines, and peer-reviewed literature. This narrative review synthesizes recent advancements in RAG applications for healthcare, focusing on evidence-grounded strategies, tailored evaluation metrics, and robust safety controls to facilitate trustworthy deployment in high-stakes medical environments. Evidence grounded in RAG frameworks involves dynamic retrieval of contextually relevant information to inform generative responses, thereby improving factual accuracy in tasks such as clinical summarization, decision support, and patient education. Studies demonstrate that RAG-enhanced LLMs outperform traditional models in extracting key clinical insights from EHRs, with applications spanning orthopedic patient education, neurosurgical consultations, and precision oncology treatment matching. For instance, integrating vector databases with LLMs enables real-time querying of molecular data to align therapeutic recommendations with patient-specific profiles, reducing errors in evidence-based practice. However, the efficacy of grounding depends on the quality of retrieved sources, necessitating hybrid retrieval techniques that balance semantic similarity and domain-specific relevance. Evaluation metrics for RAG in healthcare extend beyond conventional natural language processing benchmarks to incorporate clinical validity, coherence with medical knowledge, and user-centric outcomes. Metrics such as faithfulness scores, which assess alignment between generated content and retrieved evidence, have been adapted for biomedical contexts, revealing improvements in accuracy for tasks like fitness assessments and diabetes education. Safety controls are paramount, encompassing bias mitigation through multi-agent conversational frameworks, privacy-preserving retrieval in federated systems, and hallucination detection via uncertainty quantification. Regulatory perspectives emphasize the need for standardized safety benchmarks to prevent misinformation in patient-facing tools. This review highlights systems-level insights, including closed-loop architectures where RAG facilitates iterative feedback between data ingestion, inference, and clinical intervention. Challenges in scalability, such as computational overhead in resource-constrained settings, are addressed through optimized retrieval pipelines. We propose an original interpretive framework for RAG deployment, emphasizing interoperability with existing healthcare infrastructures to enhance analytics workflows. Ultimately, RAG holds promise for democratizing AI in healthcare, provided rigorous evaluation and safety protocols are embedded from design to implementation, paving the way for equitable, evidence-driven clinical intelligence.
The integration of generative artificial intelligence (AI) into clinical workflows represents a transformative shift in healthcare systems and analytics, promising enhanced efficiency in documentation tasks while introducing novel challenges in reliability and governance. This narrative review synthesizes recent literature on the utility of generative AI models, such as large language models (LLMs), in automating clinical documentation, including patient notes, discharge summaries, and diagnostic reports, which traditionally consume significant clinician time. Studies highlight how these tools can streamline data ingestion from electronic health records (EHRs), generating coherent narratives that align with clinical standards, thereby reducing administrative burdens and allowing more focus on patient care. For instance, generative AI has demonstrated proficiency in summarizing complex medical dialogues and classifying clinical notes, often outperforming traditional methods in speed and accuracy, as evidenced by evaluations in German healthcare settings and emergency departments. However, the utility is tempered by inherent failure modes, including hallucinations—where models produce factually incorrect information—and biases amplified from training data, which can propagate errors in clinical decision-making. Oversight mechanisms are critical to mitigate these risks, encompassing human-in-the-loop verification, regulatory frameworks like the EU AI Act, and ethical guidelines for deployment in high-stakes environments. From a systems-level perspective, generative AI enables closed-loop analytics in healthcare infrastructure, where data flows from ingestion to inference, informing interventions and feeding back for model recalibration. This review examines how LLMs facilitate intelligent clinical decision support, such as in patient care document verification using EHRs and prompt engineering for medical education. Yet, failures such as catastrophic errors in multimodal AI applications underscore the need for robust oversight, including transparency in model training and post-deployment monitoring. Comparative analyses reveal that while generative AI excels in low-risk documentation tasks, its application in critical sectors demands interdisciplinary expertise to address trust deficits and ensure equitable outcomes. The review integrates cross-study insights, proposing an original framework for AI-enabled healthcare loops that emphasizes governance at each stage to balance innovation with safety. Emerging perspectives indicate that generative AI’s role in healthcare analytics extends to predictive modeling and administrative functions, with consensus statements advocating for standardized evaluation frameworks to assess real-world viability. Challenges in failure modes, such as over-reliance on AI outputs without verification, highlight the imperative for oversight mechanisms that incorporate legal and ethical considerations, ensuring compliance with therapeutic approvals and preventing misuse in controlled substance contexts. Ultimately, this synthesis underscores the dual-edged nature of generative AI in clinical workflows: its documentation utility can revolutionize healthcare delivery, but only through vigilant oversight to avert failures that compromise patient safety. By structuring the discourse around data-model-deployment-governance continua, this review offers a novel interpretive lens for future implementations, urging stakeholders to prioritize human oversight in AI-augmented systems.
Transformer-based architectures have significantly advanced clinical natural language processing by improving the capture of contextual relationships in unstructured electronic health records compared to earlier recurrent and convolutional models, with domain-specific variants such as ClinicalBERT and BioBERT designed to better handle clinical terminology, abbreviations, and specialized language, thereby improving information extraction performance, although the relative impact of different pre-training strategies remains insufficiently synthesized and requires systematic evaluation of corpus selection and fine-tuning approaches; this systematic review mapped studies focusing on pre-training corpora, fine-tuning methods, and named entity recognition performance across entity types such as medications, diseases, procedures, laboratory tests, and social determinants of health, using PRISMA-guided methods and searches across PubMed, ACL Anthology, arXiv, and IEEE Xplore, identifying 32 eligible studies from 1,247 records; findings showed that ClinicalBERT, BioBERT, and PubMedBERT were the most frequently evaluated models, pre-trained on datasets such as MIMIC-III, PubMed abstracts, and mixed biomedical corpora, with consistent evidence that domain-specific pre-training outperforms general-domain BERT models on benchmarks like i2b2 and n2c2 despite variation across entity types and fine-tuning strategies, while clinical pre-training on large EHR corpora improves named entity recognition and optimized fine-tuning approaches such as lower learning rates and data augmentation further enhance performance, particularly for medications and diseases, underscoring the importance of domain adaptation and the need for more standardized evaluation protocols in clinical NLP research.
Patient no-shows in outpatient clinics (5%–30% across specialties) disrupt scheduling efficiency, increase wait times, and strain healthcare resources. To address this, healthcare systems are increasingly applying machine learning (ML) for predictive scheduling support. This systematic review synthesizes ML approaches for predicting outpatient no-shows, focusing on model types, feature usage, and reported operational deployment outcomes, with emphasis on translation into clinical scheduling practice. A PRISMA-compliant search of PubMed, Embase, IEEE Xplore, Scopus, and Web of Science identified studies using ML for no-show prediction in outpatient settings. Data on models, features, performance, and implementation were extracted. Risk of bias was assessed using an adapted PROBAST tool. Thirty-two studies were included. Logistic regression, random forest, and XGBoost were the most commonly used models. Historical attendance data was the dominant predictive feature. Fewer than 20% of studies reported real-world implementation, and reported intervention outcomes (e.g., overbooking, reminders) were inconsistent. While ML models show strong predictive performance, real-world deployment and evidence of operational impact remain limited. This gap highlights the need to prioritize implementation-focused research to translate predictive accuracy into measurable improvements in clinic efficiency and access.
Alzheimer’s disease (AD) is the leading cause of dementia, affecting over 50 million people worldwide, with prevalence expected to triple by 2050. Early detection is crucial for clinical trial enrollment and care planning, and multimodal data (MRI, PET, CSF biomarkers, and cognitive assessments) provides complementary information on neurodegeneration, metabolism, and protein aggregation. This systematic review synthesizes AI/ML approaches for early AD detection using multimodal data, focusing on fusion strategies and performance across disease stages. Following PRISMA guidelines, searches of PubMed, IEEE Xplore, Scopus, Web of Science, and arXiv (2017–2023) identified studies using ML/DL with at least two modalities and reporting diagnostic performance. From 1,247 records, 35 studies were included. MRI was the most used modality (>90%), followed by cognitive tests (70–80%), PET (40–50%), and CSF (20–30%). Early fusion was most common, with increasing use of intermediate fusion. Multimodal models achieved AUROC of 0.90–0.98 for AD vs controls, but lower performance (0.70–0.85) for predicting MCI conversion to AD. Overall, multimodal AI improves early AD detection, with strong performance for diagnosis but persistent challenges in forecasting MCI progression due to heterogeneity and limited longitudinal data.