Clinical Intelligence Research Press Clinical Intelligence Research Press

Search

Search results:
Large Language Models in Clinical Medicine from 2017 to 2025: A Systematic Review of Performance on Medical Licensing Examinations, Clinical Documentation, Decision Support, and Safety Concerns
Large language models (LLMs) have rapidly advanced since the transformer architecture was introduced in 2017, with systems such as GPT-3, GPT-4, Med-PaLM, and Claude increasingly explored for applications in medical education, clinical documentation, decision support, and patient communication, raising both optimism and concerns regarding safety and reliability. This systematic review synthesizes evidence across studies retrieved from PubMed, arXiv, ACL Anthology, IEEE Xplore, and Google Scholar that empirically evaluated LLMs in clinical settings using quantitative performance metrics, with risk of bias assessed using an adapted PROBAST framework for machine learning research. Findings show that LLMs achieve 60–90% accuracy on USMLE-style examinations, with leading models such as GPT-4 and Med-PaLM 2 reaching or surpassing passing thresholds, while in clinical documentation tasks they can reduce physician workload by approximately 30–50% in generating outputs such as discharge summaries, though human review remains consistently required. Performance in clinical decision support is more variable and specialty-dependent, and hallucination rates ranging from 5–30% have been reported, alongside persistent issues of bias and overconfidence in incorrect outputs. Overall, while LLMs demonstrate strong capabilities in structured medical knowledge tasks and documentation support, current limitations including hallucinations, bias, and lack of prospective clinical validation prevent safe autonomous deployment, making clinician oversight and robust safety safeguards essential for any clinical use.
Journal of Artificial Intelligence for Healthcare Systems
Review | Open access | 20 January 2026 | Article: 121

Deep Learning for Oncology Drug Discovery: A Systematic Review of Target Identification, Compound Screening, and Clinical Trial Optimization
Oncology drug development is an expensive and high-failure process, with costs exceeding two billion dollars per approved drug and success rates below 10%. Deep learning has recently been explored as a strategy to improve efficiency across the drug discovery pipeline. This systematic review evaluates its application in target identification, compound screening and de novo drug design, and clinical trial optimization. Following PRISMA 2020 guidelines, multiple databases were searched and studies were screened using predefined inclusion criteria, with risk of bias assessed via established tools. The literature shows that graph neural networks and transformer-based models are the most widely used architectures, particularly in early-stage discovery tasks. Although many studies report strong in silico performance, often with AUC values above 0.80, only a small proportion demonstrate experimental or clinical validation. Overall, deep learning significantly advances computational drug discovery in oncology, but translation into clinically validated therapies remains limited, especially in trial optimization, highlighting the need for stronger prospective and experimental validation frameworks.
Journal of Artificial Intelligence for Healthcare Systems
Review | Open access | 20 January 2026 | Article: 122

Machine Learning for Predicting Sepsis in Hospitalized Patients: A Systematic Review of Model Types, Feature Engineering Approaches, Prediction Horizons, and Prospective Validation Studies
Sepsis continues to be a major contributor to morbidity and mortality among hospitalized patients globally, especially within intensive care and emergency departments, where rapid recognition is essential for improving survival through timely treatment. In recent years, machine learning approaches have gained attention for their ability to predict sepsis onset using routinely collected electronic health record data. This systematic review, conducted in accordance with PRISMA 2020 guidelines, synthesizes evidence from studies published between 2017 and 2025, focusing on model architectures, feature selection and engineering strategies, prediction time horizons, and validation methodologies. Searches across major biomedical and informatics databases identified 67 eligible studies. The included literature shows that logistic regression, ensemble tree-based algorithms, and deep learning models are most frequently applied for sepsis prediction tasks. However, the majority of studies rely on retrospective datasets with internal validation, while only a limited number incorporate prospective or real-world validation frameworks. Overall, although reported model performance is often strong in retrospective analyses, a consistent decline in accuracy is observed when models are evaluated in real clinical environments. These findings highlight that prospective validation and improved generalizability are still underdeveloped areas, underscoring the need for future research to emphasize real-time deployment and robust external validation before clinical integration.
Journal of Artificial Intelligence for Healthcare Systems
Review | Open access | 20 January 2026 | Article: 123

Artificial Intelligence for Sleep Medicine and Sleep Disorder Diagnosis: A Systematic Review of Deep Learning Models for Polysomnography, Home Sleep Apnea Testing, and Wearable Device Analysis
Sleep disorders, including obstructive sleep apnea, insomnia, restless legs syndrome, narcolepsy, and central sleep apnea, represent a major public health burden. Polysomnography is the diagnostic gold standard but is resource-intensive, leading to increasing use of home sleep apnea testing and wearable devices to improve accessibility. This systematic review evaluates deep learning models in sleep medicine across polysomnography, home sleep apnea testing, and wearable data, focusing on architectures, signal types, validation approaches, diagnostic tasks, and clinical readiness. A PRISMA 2020–compliant search was conducted in PubMed, IEEE Xplore, Scopus, and Web of Science for studies published from 2017 to 2025, including those applying deep learning for sleep staging, apnea/hypopnea detection, or sleep disorder diagnosis using PSG, HSAT, or wearable-derived signals. Twenty-nine studies were included. Convolutional neural networks were the most widely used architecture, often combined with recurrent or hybrid models for temporal dependencies, while transformer-based models have recently emerged for long-sequence sleep analysis. Deep learning methods demonstrate strong performance in sleep staging and respiratory event detection, especially using polysomnography data. However, limited external validation, heterogeneous datasets, and a lack of prospective clinical deployment remain major barriers to clinical translation.
Journal of Artificial Intelligence for Healthcare Systems
Review | Open access | 20 January 2026 | Article: 125

Federated Learning for Rare Disease Diagnosis and Research: A Systematic Review of Methods for Handling Extreme Data Scarcity, Class Imbalance, and Site Heterogeneity
Rare diseases are challenging for AI development due to sparse patient populations, fragmented expertise, and strong inter-site variability, making federated learning a promising privacy-preserving solution for multi-institutional model training. This systematic review evaluates federated learning approaches for rare disease diagnosis and related data-scarce clinical settings, with emphasis on handling extreme data scarcity, class imbalance, heterogeneity, and privacy constraints. A PRISMA 2020-compliant search of PubMed, IEEE Xplore, Scopus, Web of Science, and arXiv (2017–2025) identified 2,015 records, with 56 studies included after screening. The most commonly used strategies included FedProx-based optimization, personalized federated learning, class-aware aggregation, generative data augmentation, and domain adaptation techniques. Overall, standard federated averaging is often insufficient under severe scarcity and distribution shift, while hybrid approaches combining personalization, augmentation, and domain adaptation show greater promise for improving performance in rare disease applications.
Journal of Artificial Intelligence for Healthcare Systems
Review | Open access | 20 January 2026 | Article: 126

Predictive Analytics for Healthcare Supply Chain Resilience during Public Health Emergencies: A Systematic Review of Models for Personal Protective Equipment Demand Forecasting and Distribution Optimization
Public health emergencies reveal critical weaknesses in healthcare supply chains, especially when PPE demand outpaces procurement and distribution capacity, making predictive analytics an important tool for forecasting demand and improving allocation during crises. This systematic review evaluates predictive analytics models for PPE demand forecasting and distribution optimization during public health emergencies, focusing on model types, data sources, validation approaches, performance metrics, equity considerations, and implementation readiness. Following PRISMA 2020 guidelines, searches were conducted in PubMed, Web of Science, Scopus, IEEE Xplore, and Google Scholar for studies published between 2017 and 2025, yielding 2,847 records, of which 35 met inclusion criteria. Included studies comprised time series and statistical models (34%), machine learning and hybrid approaches (29%), optimization methods (26%), and simulation or digital twin frameworks (11%), with limited evidence of real-world deployment. Overall, findings indicate that predictive analytics can enhance PPE supply chain resilience by improving demand forecasting, allocation decisions, and scenario testing, but widespread adoption is limited by poor data interoperability, insufficient prospective validation, weak equity integration, and limited operational integration into healthcare decision systems.
Journal of Artificial Intelligence for Healthcare Systems
Review | Open access | 20 July 2026 | Article: 127

Generative Artificial Intelligence for Medical Imaging Synthesis and Augmentation from 2017 to 2026: A Systematic Review of Diffusion Models, GANs, and VAEs for MRI, CT, X-Ray, and Pathology
Generative artificial intelligence (AI), including GANs, VAEs, and diffusion models, is increasingly used for synthesizing and enhancing medical images, helping address challenges such as limited data, expensive acquisition, and rare disease representation. This systematic review examines studies on generative AI methods for MRI, CT, X-ray, and pathology image synthesis from 2017 to 2026, focusing on synthesis tasks, evaluation strategies, and clinical utility. A PRISMA 2020-compliant search of PubMed, IEEE Xplore, Scopus, and Web of Science identified peer-reviewed research on generative models for medical image synthesis, augmentation, harmonization, or cross-modality translation. Findings show a shift from GAN-based methods to diffusion models post-2022, with MRI and CT studies emphasizing cross-modality translation, and X-ray and pathology studies focusing on augmentation and diagnostic utility. Despite GANs' continued dominance, diffusion models are gaining traction for improving image fidelity and diversity. However, evaluation practices remain inconsistent, with limited inclusion of clinically relevant assessments. This review follows PRISMA 2020 guidelines and provides a narrative synthesis of the evidence.
Journal of Artificial Intelligence for Healthcare Systems
Review | Open access | 20 July 2026 | Article: 140

Large Language Models for Clinical Trial Patient Screening and Recruitment: A Systematic Review of Zero-Shot, Few-Shot, and Fine-Tuned Approaches for Matching Eligibility Criteria to Electronic Health Records
Clinical trial recruitment is hindered by slow, costly, and labor-intensive processes, particularly due to the complexity of eligibility criteria often written in free text. This systematic review examines the use of large language models (LLMs) for matching clinical trial eligibility criteria to electronic health records (EHR). It evaluates zero-shot, few-shot, and fine-tuned LLM approaches, comparing their strengths, limitations, and deployment readiness in supporting patient-trial matching. Thirty-three studies published from 2017 to 2026 were included, with findings showing that zero-shot prompting is most adaptable for simple criteria, few-shot prompting offers consistent reasoning for ambiguous criteria, and fine-tuned models excel in task-specific performance but require labeled data and are less portable. The review concludes that no single approach is optimal for all trial screening tasks, and hybrid workflows combining various methods with human verification are most suitable for clinical use.
Journal of Artificial Intelligence for Healthcare Systems
Review | Open access | 20 July 2026 | Article: 141

Artificial Intelligence for Real-Time Patient Monitoring in Smart Hospitals and Home Settings: A Systematic Review of Edge AI Architectures, Wearable Sensor Fusion, and Clinical Alert Systems
This systematic review examines the use of edge artificial intelligence (AI) and wearable sensors for real-time patient monitoring in smart hospitals and home settings, focusing on detecting deterioration, falls, arrhythmias, and infection-related changes. The review synthesizes studies from 2017 to 2026 on edge AI architectures, wearable sensor fusion, and clinical alert systems, emphasizing latency, power constraints, alert performance, and integration into clinical workflows. A PRISMA 2020-compliant search identified 127 studies from 2,100 records, with findings showing that while edge AI execution grew post-2020, it still represented a minority of designs. Sensor fusion was often linked to broader event coverage but increased implementation complexity. The review concludes that edge AI can reduce latency and enhance privacy but introduces challenges related to power usage, model complexity, device reliability, and maintenance, with limited clinical validation of alert systems and few studies addressing alert fatigue or clinician response.
Journal of Artificial Intelligence for Healthcare Systems
Review | Open access | 20 July 2026 | Article: 143
Filters
Clear All

Subject
AI-driven Diagnostics Artificial Intelligence in Health Informatics Artificial Intelligence in Healthcare Big Data in Healthcare Clinical Data Mining Clinical Decision Support Systems Clinical Informatics Computer Vision Connected Health Systems Deep Learning Digital Health Digital Healthcare Innovation Digital Transformation in Healthcare Electronic Health Records Ethical AI in Healthcare Explainable AI Health Data Analytics Health Data Privacy Health Informatics Health Information Management Health Information Systems Health System Optimization Health Technology Assessment Healthcare Data Science Healthcare Informatics Healthcare Information Security Healthcare Management Healthcare Management Information Systems Intelligent Medical Systems Internet of Medical Things (IoMT) Interoperability in Healthcare Systems Machine Learning Medical Data Analytics Medical Data Management Medical Imaging Mobile Health (mHealth) Natural Language Processing Precision Medicine Predictive Analytics Remote Patient Monitoring Smart Healthcare Systems Telemedicine Wearable Health Technologies e-Health




Access type