The integration of large language models (LLMs) into clinical healthcare systems represents a transformative shift in how data analytics, decision support, and operational infrastructure are conceptualized and deployed. This narrative review synthesizes recent advancements in LLMs within healthcare, focusing on their roles in enhancing clinical analytics, infrastructural frameworks, and oversight mechanisms while addressing inherent risk dynamics. Drawing from peer-reviewed literature, we examine how LLMs facilitate the processing of vast unstructured clinical data, such as electronic health records and patient narratives, to generate actionable insights that inform diagnostics, treatment planning, and resource allocation. Key infrastructural elements include scalable deployment pipelines that integrate LLMs with existing hospital information systems, enabling real-time analytics and predictive modeling without disrupting legacy workflows. Oversight is emphasized through regulatory frameworks that ensure ethical deployment, data privacy compliance, and bias mitigation, as LLMs amplify risks related to misinformation, algorithmic opacity, and equitable access in diverse clinical settings. Risk dynamics are explored in terms of model hallucinations, dependency on training data quality, and potential for exacerbating healthcare disparities if not properly governed. The review highlights systems-level analytics where LLMs contribute to closed-loop healthcare ecosystems, from data ingestion and inference to feedback-driven recalibration, fostering adaptive intelligence in clinical decision-making. For instance, LLMs have been adapted for tasks like text summarization, diagnostic reasoning, and patient communication, outperforming traditional methods in efficiency while requiring robust validation to maintain clinical fidelity. We underscore the need for interdisciplinary collaboration between clinicians, data scientists, and policymakers to harness LLMs' potential in optimizing healthcare delivery. By synthesizing cross-study evidence, this review proposes an original interpretive framework for LLM-enabled healthcare systems, structured around data-model-deployment-governance cycles, to guide future implementations. Ultimately, while LLMs promise enhanced analytics and infrastructural resilience, their clinical adoption demands vigilant oversight to balance innovation with patient safety and ethical integrity. This synthesis not only maps the current landscape but also identifies infrastructural gaps in scaling LLMs for equitable, high-stakes clinical environments, paving the way for more resilient healthcare analytics paradigms.
Large language models (LLMs) have rapidly advanced since the transformer architecture was introduced in 2017, with systems such as GPT-3, GPT-4, Med-PaLM, and Claude increasingly explored for applications in medical education, clinical documentation, decision support, and patient communication, raising both optimism and concerns regarding safety and reliability. This systematic review synthesizes evidence across studies retrieved from PubMed, arXiv, ACL Anthology, IEEE Xplore, and Google Scholar that empirically evaluated LLMs in clinical settings using quantitative performance metrics, with risk of bias assessed using an adapted PROBAST framework for machine learning research. Findings show that LLMs achieve 60–90% accuracy on USMLE-style examinations, with leading models such as GPT-4 and Med-PaLM 2 reaching or surpassing passing thresholds, while in clinical documentation tasks they can reduce physician workload by approximately 30–50% in generating outputs such as discharge summaries, though human review remains consistently required. Performance in clinical decision support is more variable and specialty-dependent, and hallucination rates ranging from 5–30% have been reported, alongside persistent issues of bias and overconfidence in incorrect outputs. Overall, while LLMs demonstrate strong capabilities in structured medical knowledge tasks and documentation support, current limitations including hallucinations, bias, and lack of prospective clinical validation prevent safe autonomous deployment, making clinician oversight and robust safety safeguards essential for any clinical use.
Clinical trial recruitment is hindered by slow, costly, and labor-intensive processes, particularly due to the complexity of eligibility criteria often written in free text. This systematic review examines the use of large language models (LLMs) for matching clinical trial eligibility criteria to electronic health records (EHR). It evaluates zero-shot, few-shot, and fine-tuned LLM approaches, comparing their strengths, limitations, and deployment readiness in supporting patient-trial matching. Thirty-three studies published from 2017 to 2026 were included, with findings showing that zero-shot prompting is most adaptable for simple criteria, few-shot prompting offers consistent reasoning for ambiguous criteria, and fine-tuned models excel in task-specific performance but require labeled data and are less portable. The review concludes that no single approach is optimal for all trial screening tasks, and hybrid workflows combining various methods with human verification are most suitable for clinical use.