Clinical Intelligence Research Press Clinical Intelligence Research Press

Search

Search results:
Large Language Models in Clinical Medicine from 2017 to 2025: A Systematic Review of Performance on Medical Licensing Examinations, Clinical Documentation, Decision Support, and Safety Concerns
Large language models (LLMs) have rapidly advanced since the transformer architecture was introduced in 2017, with systems such as GPT-3, GPT-4, Med-PaLM, and Claude increasingly explored for applications in medical education, clinical documentation, decision support, and patient communication, raising both optimism and concerns regarding safety and reliability. This systematic review synthesizes evidence across studies retrieved from PubMed, arXiv, ACL Anthology, IEEE Xplore, and Google Scholar that empirically evaluated LLMs in clinical settings using quantitative performance metrics, with risk of bias assessed using an adapted PROBAST framework for machine learning research. Findings show that LLMs achieve 60–90% accuracy on USMLE-style examinations, with leading models such as GPT-4 and Med-PaLM 2 reaching or surpassing passing thresholds, while in clinical documentation tasks they can reduce physician workload by approximately 30–50% in generating outputs such as discharge summaries, though human review remains consistently required. Performance in clinical decision support is more variable and specialty-dependent, and hallucination rates ranging from 5–30% have been reported, alongside persistent issues of bias and overconfidence in incorrect outputs. Overall, while LLMs demonstrate strong capabilities in structured medical knowledge tasks and documentation support, current limitations including hallucinations, bias, and lack of prospective clinical validation prevent safe autonomous deployment, making clinician oversight and robust safety safeguards essential for any clinical use.
Journal of Artificial Intelligence for Healthcare Systems
Review | Open access | 20 January 2026 | Article: 121

Large Language Model with Retrieval-Augmented Generation and Chain-of-Thought Reasoning for Differential Diagnosis Generation from Emergency Department Triage Notes and Vital Signs
This article proposes a conceptual framework for a diagnostic support system in emergency departments that leverages large language models, retrieval-augmented generation, and chain-of-thought reasoning. By combining triage notes and vital signs, the system generates a ranked differential diagnosis list to assist clinicians without replacing their judgment. The framework includes components like a triage note encoder, a vital sign encoder, a retrieval module, and a diagnosis ranker, using evidence from clinical guidelines, curated references, and de-identified prior cases. The approach grounds the model in authoritative knowledge while ensuring transparency and explainability in the diagnostic process. However, prospective validation, integration into workflows, and clinician oversight are crucial before implementation to ensure safety and effectiveness.
Journal of Artificial Intelligence for Healthcare Systems
Original Research | Open access | 20 July 2026 | Article: 129

Large Language Models for Clinical Trial Patient Screening and Recruitment: A Systematic Review of Zero-Shot, Few-Shot, and Fine-Tuned Approaches for Matching Eligibility Criteria to Electronic Health Records
Clinical trial recruitment is hindered by slow, costly, and labor-intensive processes, particularly due to the complexity of eligibility criteria often written in free text. This systematic review examines the use of large language models (LLMs) for matching clinical trial eligibility criteria to electronic health records (EHR). It evaluates zero-shot, few-shot, and fine-tuned LLM approaches, comparing their strengths, limitations, and deployment readiness in supporting patient-trial matching. Thirty-three studies published from 2017 to 2026 were included, with findings showing that zero-shot prompting is most adaptable for simple criteria, few-shot prompting offers consistent reasoning for ambiguous criteria, and fine-tuned models excel in task-specific performance but require labeled data and are less portable. The review concludes that no single approach is optimal for all trial screening tasks, and hybrid workflows combining various methods with human verification are most suitable for clinical use.
Journal of Artificial Intelligence for Healthcare Systems
Review | Open access | 20 July 2026 | Article: 141

Generative Artificial Intelligence Framework for Producing Draft Clinical Governance Dashboards from Quality Metrics, Audit Reports, Incident Logs, Compliance Indicators, and Executive Reporting Templates
Hospital boards and quality committees depend on monthly clinical governance dashboards to oversee safety, quality, compliance, and organizational risk. Yet producing these reports often requires repeated manual consolidation of metrics, audit findings, incident summaries, and executive commentary from fragmented operational systems. Manual dashboard assembly can delay insight, increase administrative burden, and introduce errors when data are copied across spreadsheets, documents, and presentation templates. These workflows also make it difficult to maintain consistent language, trace every claim to its source, and produce timely board-ready narratives. This article proposes a generative AI framework that ingests quality metrics, audit reports, incident logs, compliance indicators, and executive reporting templates to produce a complete draft clinical governance dashboard. The framework is conceptual and intended to support first-draft preparation rather than autonomous publication. The proposed architecture includes a multi-source data ingestion layer, a template-guided large language model, a factual verification module, and a collaborative human-review interface. Together, these components could transform scattered institutional data into structured narrative sections, exception summaries, and draft governance commentary. The framework could reduce report preparation time, improve formatting consistency, and allow quality specialists to focus on interpretation, escalation, and improvement planning rather than repetitive assembly. Its value would depend on source grounding, auditability, privacy protection, and a clear approval workflow. Automated draft governance reporting is an audacious but feasible application of generative AI in complex health systems. With careful design, human oversight, and rigorous evaluation, such systems could support higher-integrity reporting without replacing clinical accountability.
Journal of Health Informatics and Digital Systems
Original Research | Open access | 25 February 2026 | Article: 120

Generative Artificial Intelligence for Healthcare Administration: A Systematic Review of Applications in Documentation Support, Operational Reporting, Patient Communication, Governance Dashboards, and Workflow Automation
Healthcare administration is document-intensive, communication-heavy, and increasingly dependent on digital systems that require timely synthesis of clinical and operational information. Generative artificial intelligence has been proposed as a potential means of reducing administrative workload across documentation, reporting, patient communication, governance, and workflow coordination. This systematic review examined applications of generative artificial intelligence across five healthcare administrative domains: documentation support, operational reporting, patient communication, governance dashboards, and workflow automation. The review aimed to characterize reported use cases, model types, evaluation approaches, implementation barriers, and evidence maturity from 2017 to 2025. A PRISMA 2020-compliant search strategy was applied to PubMed, Scopus, IEEE Xplore, and Web of Science for studies published between January 1, 2017, and December 31, 2025. Screening was conducted in duplicate, and eligible studies were synthesized narratively by administrative domain. Documentation support was the most mature domain, particularly for clinical summarization, discharge summaries, patient-message drafting, and document classification. Operational reporting and governance dashboards were emerging areas, while patient communication and workflow automation showed diverse prototypes but limited prospective validation. Common challenges included hallucination, privacy, bias, regulatory uncertainty, integration burden, and limited evidence of real-world administrative impact. Generative artificial intelligence appears technically capable of supporting multiple healthcare administrative tasks, but evidence remains uneven across domains. Real-world impact, safety, scalability, and governance require stronger evaluation before these systems can be relied upon for high-stakes administrative decision-making.
Journal of Health Informatics and Digital Systems
Review | Open access | 25 February 2026 | Article: 121

Large Language Model Agent for Drafting Structured Responses to Non-Urgent Patient Portal Messages Using Local Clinical Protocols, Physician-Approved Templates, Patient History, and Message Intent Classification
Non-urgent patient portal messages consume a substantial fraction of clinician time and can shift attention away from higher-acuity care. As portal communication becomes a routine part of ambulatory medicine, inbox management increasingly functions as an additional clinical workload. Current inbox workflows often depend on manual review, interpretation, and free-text reply by clinicians. This creates delays, variability, and cognitive burden, especially for routine requests that could be answered through standardized guidance. This article proposes a large language model agent that classifies incoming message intent, identifies non-urgent queries suitable for templated responses, and drafts a structured reply. The draft is grounded in local clinical protocols, physician-approved templates, and patient-specific information from the electronic health record. The proposed agent includes a message intent classifier, a protocol-and-template retrieval module, a retrieval-augmented drafting model, a rule-based safety filter, and a human-verification interface. These components work together to keep generated text within a clinically approved and auditable workflow. By preparing protocol-adherent draft responses for clinician review, the agent could reduce routine documentation effort while preserving clinician judgment. Its value depends on safety boundaries, template quality, patient-data integration, and seamless fit within the existing portal workflow. The agent represents a pragmatic near-term use of large language models in clinical administration. It supports automation without removing human oversight from patient-facing communication.
Journal of Health Informatics and Digital Systems
Original Research | Open access | 25 February 2026 | Article: 130

Retrieval-Augmented Generation System for Supporting Hospital Accreditation Preparation Using Policy Documents, Audit Findings, Quality Indicators, Regulatory Standards, and Departmental Evidence Files
Hospital accreditation requires organizations to demonstrate that clinical, operational, safety, and governance processes are aligned with externally defined standards. This demonstration depends on extensive internal documentation, including current policies, procedure manuals, audit reports, quality dashboards, meeting records, training logs, and department-level evidence files. Accreditation preparation is often conducted through manual document searches across fragmented repositories. This process is slow, resource-intensive, and vulnerable to missed evidence, outdated documents, inconsistent interpretations, and duplication of staff effort. This article proposes a retrieval-augmented generation system designed to support hospital accreditation preparation without replacing human judgment. The system indexes internal accreditation-relevant documents and allows authorized users to ask natural-language questions that receive synthesized, citation-backed responses. The proposed system includes a multi-source ingestion pipeline, a metadata-rich vector database, a permissioned large language model, a citation-grounding layer, and a human review dashboard. Together, these components would support evidence retrieval, gap identification, traceability, and accreditation team collaboration. The system could assist accreditation teams by reducing time spent locating and cross-referencing evidence. It would also be expected to improve the completeness and consistency of preparation materials by grounding responses in authoritative internal sources. A specialized retrieval-augmented generation system could help hospitals move from episodic accreditation preparation toward continuous audit readiness. Its value would depend on document quality, governance safeguards, human review, and rigorous evaluation in real accreditation workflows.
Journal of Health Informatics and Digital Systems
Original Research | Open access | 20 July 2026 | Article: 135

Artificial Intelligence for Digital Patient Communication: A Systematic Review of Portal Message Triage, Chatbot Support, Care Navigation, Automated Response Drafting, and Patient Engagement Analytics
Asynchronous digital communication with patients has become a routine component of modern healthcare delivery. The rapid growth of patient portals, chatbots, and digital front-door tools has created opportunities for more responsive care, while also increasing communication workload for clinical teams. This systematic review examined artificial intelligence applications in digital patient communication from 2017 to 2026. The review focused on portal message triage, chatbot support, care navigation, automated response drafting, and patient engagement analytics. A PRISMA 2020-compliant review was conducted using structured searches of PubMed, Scopus, IEEE Xplore, and Web of Science. Records were screened by two reviewers, with data extracted on communication domain, AI approach, clinical setting, evaluation strategy, safety reporting, and implementation maturity. The literature was dominated by studies of chatbot support and portal message triage, with a growing body of work on large language model-enabled response drafting. Care navigation and patient engagement analytics were less frequently evaluated, and most studies emphasized technical performance, user satisfaction, or feasibility rather than health outcomes or workload reduction in real-world settings. AI for patient communication appears technically promising in isolated tasks, particularly message classification, chatbot interaction, and draft response generation. However, evidence remains limited regarding safe, equitable, and effective deployment across integrated communication workflows.
Journal of Health Informatics and Digital Systems
Review | Open access | 20 July 2026 | Article: 141
Filters
Clear All

Subject
AI-driven Diagnostics Artificial Intelligence in Health Informatics Artificial Intelligence in Healthcare Big Data in Healthcare Clinical Data Mining Clinical Decision Support Systems Clinical Informatics Computer Vision Connected Health Systems Deep Learning Digital Health Digital Healthcare Innovation Digital Transformation in Healthcare Electronic Health Records Ethical AI in Healthcare Explainable AI Health Data Analytics Health Data Privacy Health Informatics Health Information Management Health Information Systems Health System Optimization Health Technology Assessment Healthcare Data Science Healthcare Informatics Healthcare Information Security Healthcare Management Healthcare Management Information Systems Intelligent Medical Systems Internet of Medical Things (IoMT) Interoperability in Healthcare Systems Machine Learning Medical Data Analytics Medical Data Management Medical Imaging Mobile Health (mHealth) Natural Language Processing Precision Medicine Predictive Analytics Remote Patient Monitoring Smart Healthcare Systems Telemedicine Wearable Health Technologies e-Health




Access type