Clinical documentation is central to continuity of care, coding, billing, compliance, and quality measurement, yet it remains a major source of administrative burden for physicians. Artificial intelligence, especially natural language processing, has been proposed as a way to improve documentation quality, coding accuracy, billing support, and clinical workflow efficiency. This systematic review synthesised evidence from 2017 to 2023 on natural language processing methods applied to clinical documentation improvement. The review focused on automated clinical coding, note quality, billing support, computer-assisted physician documentation, and physician workflow efficiency. A PRISMA 2020-compliant search strategy was applied to PubMed, Scopus, IEEE Xplore, and Web of Science for publications from 1 January 2017 to 31 December 2023. Screening was performed in duplicate, and eligible studies were narratively synthesised by documentation domain, model type, deployment maturity, and evaluation approach. The evidence base expanded rapidly during the review period, especially in automated coding and clinical note generation. Several studies reported technically promising systems for ICD coding, documentation summarisation, and speech-derived notes, whereas billing outcomes and sustained workflow effects were less commonly evaluated. Natural language processing for clinical documentation improvement appears most mature for automated coding and note analysis. Evidence for billing support and physician workflow transformation remains less developed, particularly in prospective clinical environments.
Clinical documentation has become both a clinical asset and an administrative burden, shaping continuity of care, reimbursement, compliance, quality reporting, and medico-legal accountability. Studies using electronic health record logs and time-motion observations reported that physicians spend substantial time on documentation and desktop medicine, contributing to perceived workload and after-hours work [1, 2]. Documentation burden also has implications for billing accuracy and regulatory compliance, because incomplete, redundant, or poorly structured notes may weaken the link between clinical care, coded diagnoses, and medical necessity [3, 4]. These concerns created a strong rationale for clinical documentation improvement strategies that extend beyond conventional auditing and manual coding workflows [5].
Natural language processing has emerged as a major technical approach for clinical documentation improvement because much of the clinically and financially relevant evidence is embedded in free-text notes. Reviews of clinical natural language processing described a growing body of systems for extracting concepts, normalising terminology, identifying clinical events, and structuring narrative documentation [6, 7]. Since 2017, the field has increasingly shifted from rule-based extraction and classical machine learning toward neural architectures, attention mechanisms, and transformer-based approaches for coding and documentation-related tasks [8, 9]. This technical evolution has made it possible to address clinical note generation, automated coding, and documentation quality at a scale that would be difficult through manual review alone [10, 11].
The evidence base remains fragmented across several domains that are often evaluated separately. Automated coding studies commonly focus on ICD code prediction from discharge summaries or clinical notes, whereas note quality studies examine redundancy, copy-forward behaviour, documentation completeness, or speech recognition errors [12-15]. Workflow studies more often assess ambient speech, automated medical scribes, and physician time, while billing support is frequently inferred through coding alignment rather than directly measured through denials or revenue cycle outcomes [15, 16]. As a result, prior evidence has not always connected coding accuracy, note quality, billing compliance, and physician workflow into a single clinical documentation improvement framework [17, 18].
Figure 1 illustrates the integrated clinical documentation improvement framework linking natural language processing methods to coding accuracy, note quality, billing support, and physician workflow outcomes.

Figure 1. Integrated Clinical Documentation Improvement (CDI) Framework Linking NLP Capabilities to Clinical, Operational, and Financial Outcomes
The objective of this systematic review was to synthesise peer-reviewed evidence from 2017 to 2023 on artificial intelligence and natural language processing for clinical documentation improvement. The review followed PRISMA 2020 principles and examined four linked domains: automated clinical coding, note quality, billing support, and physician workflow efficiency. Because computer-assisted physician documentation and ambient documentation systems increasingly blend these domains, the review also considered human factors, deployment maturity, and evaluation methods. The emphasis was on systematic review-level synthesis rather than meta-analysis, because the included studies varied substantially in clinical setting, documentation type, model architecture, and outcome definition.
A structured search was conducted across PubMed, Scopus, IEEE Xplore, and Web of Science for studies published from 1 January 2017 to 31 December 2023. Search terms combined concepts for natural language processing, clinical documentation improvement, automated coding, note quality, billing support, computer-assisted physician documentation, ambient speech, physician workflow, ICD-10, CPT, SNOMED CT, and electronic health records [6, 7]. Example search concepts included natural language processing for clinical documentation improvement, automated coding from clinical notes, documentation completeness and accuracy, billing support using artificial intelligence, and ambient documentation for physician workflow.
Studies were eligible if they reported peer-reviewed research on artificial intelligence, machine learning, natural language processing, speech recognition, or automated summarisation applied to clinical documentation tasks. Eligible domains included automated coding, clinical note generation, note quality assessment, documentation burden, billing support, documentation-code alignment, and computer-assisted physician documentation. Studies were excluded if they were not focused on clinical documentation, did not involve clinical text or documentation workflow, were editorials or opinion pieces without empirical or systematic review content, or fell outside the publication window except where a prior review was used to contextualise the field. English-language studies were prioritised because the review focused on documentation systems and evaluation methods that could be consistently compared across included publications.
Records were screened in duplicate using title and abstract review followed by full-text assessment for studies that appeared relevant to clinical documentation improvement. The search identified 1,904 records, of which 1,512 remained after deduplication; 250 full-text articles were assessed, and 31 publications were selected in the final synthesis after exclusions for non-documentation focus, insufficient NLP detail, non-clinical text, or lack of relevance to CDI outcomes.
Figure 2 presents the PRISMA 2020 flow diagram detailing the identification, screening, eligibility, and inclusion of studies in this review.

Figure 2. PRISMA 2020 Flow Diagram of Study Selection for AI-Based Clinical Documentation Improvement Review
Data extraction captured bibliographic information, clinical setting, documentation type, NLP method, task definition, evaluation metrics, outcome domain, and deployment status. For automated coding studies, extraction focused on code system, model type, input note type, label granularity, explainability, and whether the system was evaluated retrospectively or in a clinical workflow [8-10]. For note quality, billing, and workflow studies, extraction captured documentation completeness, redundancy, speech-derived error patterns, physician time, after-hours charting, user experience, and evidence of revenue-cycle relevance [3, 4, 14]. Each study was assigned to one or more CDI domains because several systems addressed multiple downstream objectives, such as coding support and documentation completeness [18, 19].
Risk of bias was assessed narratively using domains adapted from PROBAST-AI and from methodological concerns specific to clinical NLP evaluation. Relevant domains included representativeness of training data, label quality, clinical setting, missingness, local coding practices, external validation, model transparency, and appropriateness of outcome measures [8, 10]. For workflow and documentation burden studies, additional attention was given to whether the design relied on retrospective logs, observation, self-report, or prospective deployment, because each approach captures a different aspect of physician work [1, 2, 4]. Studies of speech recognition and automated note generation were also reviewed for human correction requirements and potential documentation safety risks [14, 16].
Because the studies varied substantially in tasks, datasets, clinical settings, outcome measures, and deployment maturity, a meta-analysis was not performed. Instead, findings were narratively synthesised across four domains: automated coding, note quality, billing support, and physician workflow efficiency [7, 10]. Vote-counting by conceptual category was used to describe which model families and evaluation outcomes were most common, including convolutional neural networks, recurrent neural networks, attention models, transformer-based models, speech-recognition pipelines, and sequence-to-sequence summarisation [8, 9, 18]. The synthesis emphasised whether systems were technically validated, clinically deployed, evaluated by users, or linked to measurable financial and workflow outcomes [4, 15].
The final cited evidence set included 31 publications that collectively represented systematic reviews, automated coding models, clinical note quality studies, documentation burden analyses, speech-recognition evaluations, and automated note generation systems. The most common reasons for exclusion at full-text review were absence of a clinical documentation task, insufficient detail on natural language processing methods, focus on non-clinical text, or lack of relevance to coding, note quality, billing, or workflow outcomes [6, 7]. Several excluded studies addressed general EHR prediction or clinical decision support but did not evaluate documentation improvement as an explicit target. The included set therefore prioritised studies with direct implications for clinical documentation improvement rather than broader medical artificial intelligence [10, 15].
The evidence base showed a clear increase in publications after 2017, with early work focusing heavily on documentation burden, speech recognition, and neural coding models, followed by greater interest in transformers and automated note generation. Studies were conducted across inpatient, outpatient, emergency, and primary care settings, although many coding models relied on retrospective hospital datasets and benchmark corpora rather than prospective clinical deployments [8, 10, 12]. Geographically, the evidence included studies from North America, Europe, and Asia, with language and coding-system differences shaping the transferability of automated coding models [13, 20]. Clinical documentation improvement was therefore represented not as a single intervention type but as a collection of NLP-enabled tasks embedded in different health system contexts [6, 21].
Automated coding was the most technically developed domain, with studies applying neural and attention-based methods to assign ICD or related codes from free-text clinical documentation. Early deep learning approaches used convolutional, recurrent, and attention mechanisms, while later systems increasingly examined transformer-based or hybrid architectures for mapping long clinical notes to diagnosis codes [8, 9, 22]. Several studies reported that label-aware attention, residual convolutional networks, semantic matching, and knowledge integration could improve code assignment relative to simpler baselines, although evaluation conditions varied across datasets and code systems [13, 23, 24]. Systematic synthesis suggests that automated coding is the CDI domain with the strongest technical literature but not necessarily the strongest evidence of clinical implementation [10].
Despite extensive modelling work, prospective adoption of automated coding systems appeared limited in the included literature. Several studies evaluated models retrospectively on historical notes, which is valuable for technical benchmarking but does not fully capture coder trust, physician documentation behaviour, local coding variation, or payer-specific requirements [11, 12, 25]. Studies that used explainable approaches and label attention were relevant to implementation because they attempted to make code predictions interpretable to coders and clinicians [8, 26]. However, the review found that real-world adoption remains constrained by differences between benchmark performance and the operational requirements of coding departments [10].
Note quality studies addressed a different aspect of clinical documentation improvement by examining whether notes are complete, accurate, concise, and clinically meaningful. Evidence on note length and redundancy suggested that outpatient progress notes may become longer and more repetitive over time, creating risks of information overload and copy-forward bloat [3]. Studies of documentation concordance and speech-recognition-assisted notes also showed that clinical documents may diverge from observed physician behaviour or contain errors introduced during dictation and transcription workflows [14, 27]. These findings support the relevance of NLP tools that detect missing elements, reduce redundant text, and identify potential documentation inconsistencies [28].
Evidence linking NLP-driven note quality improvement to clinician satisfaction was less developed than evidence describing documentation burden itself. EHR log studies and workload assessments showed that documentation consumes substantial physician and nursing time, but fewer studies evaluated whether NLP interventions directly improved perceived note usefulness, satisfaction, or burnout-related outcomes [1, 4]. Automated summarisation and SOAP-note generation systems suggested a potential route to more concise and structured documentation, although many evaluations remained technical rather than user-centred [18, 19]. Overall, the literature indicates that reducing note length alone should not be treated as equivalent to improving clinical documentation quality unless completeness, accuracy, and clinician trust are also assessed [3, 28].
Billing support appeared in the evidence base primarily through automated coding, documentation-code alignment, and extraction of clinical evidence that could justify diagnoses or procedures. Automated ICD coding studies implicitly address revenue-cycle needs because diagnosis codes influence reimbursement, risk adjustment, quality reporting, and audit exposure [10, 13]. Explainable coding models were particularly relevant to billing support because they attempted to connect predicted codes with textual evidence in the note, which may assist coders in verifying medical necessity and documentation sufficiency [8, 26]. However, the literature rarely evaluated denial prevention, payer audit outcomes, or charge capture as primary endpoints [10].
Financial outcomes were among the least directly measured endpoints in the included literature. Although automated coding studies frequently evaluated accuracy-oriented metrics, few studies connected model outputs to actual reimbursement, denial reduction, coder productivity, or audit outcomes in deployed revenue-cycle workflows [9, 23, 29]. This gap may reflect commercial sensitivity, institutional variation in billing rules, and the difficulty of attributing financial outcomes to documentation tools alone. Consequently, the review found that billing support is a high-potential CDI domain but one with weaker independent evidence than automated coding or documentation burden measurement [4, 10].
Physician workflow studies focused on reducing documentation burden through speech recognition, automated medical scribes, and generation of structured clinical notes from patient-physician conversations. Automated scribe and medical conversation summarisation studies reported pipelines that transform dialogue into documentation, including systems for generating reports or SOAP-style notes [17-19]. These studies are important for CDI because generated notes can affect coding specificity, note completeness, billing defensibility, and physician time. However, evaluation methods varied substantially, with many systems assessed through technical or simulated documentation outputs rather than sustained clinical deployment [15].
Documentation burden studies using EHR logs showed that physicians divide work between patient-facing care and desktop medicine, with after-hours documentation contributing to workload concerns [1, 2]. Clinical documentation burden measurement using event log metadata provided a foundation for evaluating whether AI documentation tools reduce time in notes, orders, messages, and other EHR activities [4]. Although some documentation automation systems aimed to reduce cognitive load by converting conversations into structured notes, fewer studies measured task switching, after-hours charting, or cognitive burden as primary outcomes [15,19]. The evidence therefore supports workflow efficiency as a key promise of AI-CDI but not yet as a consistently demonstrated outcome [5].
The technical trajectory moved from rule-based and classical NLP methods toward deep learning, attention, and transformer-based architectures. Automated coding studies included convolutional neural networks, recurrent convolutional models, label attention models, semantic matching methods, knowledge integration, and transformer-based representations [8, 9, 11, 22]. Speech and note generation studies used automated speech recognition, sequence-to-sequence modelling, modular summarisation, and conversation-to-note pipelines [16, 18, 19]. Across domains, the strongest methodological trend was the increasing use of models that encode long narrative text while attempting to preserve interpretability for clinical or coding users [23, 26].
Table 1 provides a cross-domain analytical mapping of NLP capabilities to clinical documentation improvement outcomes and their relative evidence maturity.
Table 1. Cross-Domain Analytical Mapping of NLP Capabilities to CDI Outcomes and Evidence Maturity
NLP Capability | Primary CDI Domain | Secondary Effects Across Domains | Typical Evaluation Metrics | Evidence Maturity | Key Implementation Gap |
Automated ICD Coding | Coding Accuracy | Billing compliance, audit defensibility | F1 score, precision/recall, top-k accuracy | High (retrospective) | Limited prospective validation and coder trust integration |
Clinical Note Summarisation | Note Quality | Coding specificity, workflow efficiency | ROUGE, BLEU, human evaluation | Moderate | Risk of information omission affecting billing and care continuity |
Speech Recognition & Ambient Scribes | Workflow Efficiency | Note quality, coding downstream effects | Word error rate, time saved | Moderate | Error propagation and need for human correction |
Documentation Completeness Detection | Note Quality | Billing justification, audit readiness | Recall of missing elements, rule-based checks | Low–Moderate | Lack of linkage to clinical and financial endpoints |
Documentation-Code Alignment Models | Billing Support | Coding accuracy, compliance | Alignment accuracy, explainability scores | Low | Rare evaluation against real payer audit outcomes |
Explainable AI for Coding | Coding + Billing | Trust, workflow integration | Attention visualization, interpretability metrics | Emerging | Limited validation in real-world coder workflows |
Conversation-to-Note Generation | Workflow + Note Quality | Coding accuracy, clinician satisfaction | Human review, structural completeness | Emerging | Insufficient longitudinal workflow impact data |
Evaluation metrics varied by CDI domain and often reflected technical convenience more than operational relevance. Automated coding studies commonly relied on accuracy-oriented and classification metrics, while note generation studies used text-generation and human-review methods, and documentation burden studies used EHR logs, time-motion observations, or workload proxies [1, 8, 18]. Note quality studies examined length, redundancy, concordance, and error types, but these measures were not always connected to patient safety, billing defensibility, or clinician satisfaction [3, 14, 27]. Billing support evaluations were particularly limited because claim-level accuracy, denial prevention, and revenue-cycle return were rarely reported in a prospective and independently replicated way [10].
Deployment challenges were repeatedly linked to EHR integration, clinician trust, workflow fit, transparency, local documentation norms, and the risk of introducing new errors. Automated coding models may perform well on retrospective datasets but still require integration with coder workflows, local code policies, and audit processes [10, 13]. Speech and automated note systems must also manage transcription errors, summarisation omissions, inappropriate phrasing, and physician accountability for AI-generated documentation [14, 15]. Human factors are therefore central to AI-CDI because clinicians and coders must decide when to accept, edit, override, or ignore machine-generated suggestions [8, 28].
Automated clinical coding was the most mature area in the reviewed literature, supported by a dense technical literature using neural architectures, attention mechanisms, and transformer-based representations. Several studies reported that modern models can assign diagnosis codes from narrative clinical notes under retrospective evaluation conditions, but these results do not automatically translate into operational coding reliability [8, 11, 22]. Explainability and evidence highlighting are important because coders must understand why a code was suggested and whether the documentation supports it [8, 26]. The central gap is therefore not whether models can predict codes in principle, but whether they can be safely embedded into real-world coding and documentation workflows [10].
The note quality literature showed that documentation can become long, redundant, discordant, or vulnerable to speech-recognition errors. Studies on note length, observed behaviour, and dictated documents indicated that documentation quality cannot be reduced to the mere presence of more text [3, 14, 27]. NLP systems that reduce redundancy or generate summaries may improve usability, but they must preserve clinically relevant detail and avoid omitting information needed for billing, audit, or care coordination [18, 19]. Future studies should therefore treat conciseness, completeness, accuracy, and clinical usefulness as linked but distinct outcomes [28].
Billing support was frequently implied through automated coding but less often evaluated as a primary clinical documentation improvement outcome. Coding studies are relevant to reimbursement because they connect clinical text to billable diagnoses, but most evaluations stopped at code prediction rather than measuring denial prevention, audit outcomes, or documentation sufficiency [13, 23, 29]. Explainable models could support billing workflows by showing the textual evidence behind suggested codes, yet this capability still requires validation against payer rules and medical necessity criteria [8, 26]. The review therefore suggests that billing support is an important but underdeveloped bridge between NLP research and revenue-cycle practice [10].
Physician workflow improvement is a major motivation for AI documentation systems, but the evidence remains uneven. EHR log studies documented the scale of documentation burden, while automated scribe and note-generation studies demonstrated technical feasibility for converting conversations into structured documentation [1, 4, 16]. However, few studies provided long-term independent evidence that these tools reduce after-hours charting, improve clinician satisfaction, preserve note quality, and maintain billing integrity in routine care [15, 19]. This limits confidence in broad claims that AI documentation systems have already transformed physician workflow at scale [5].
Many included studies evaluated standalone models rather than tools embedded in native EHR documentation, coding, and billing workflows. This distinction matters because an NLP model that performs well offline may fail if it interrupts clinicians, generates excessive suggestions, requires duplicative review, or does not align with local templates and coding rules [10, 15]. Documentation burden studies show that workflow context is essential when assessing time and efficiency, because EHR work is distributed across notes, orders, inboxes, and after-hours tasks [1, 4]. AI-CDI should therefore be evaluated as a workflow intervention rather than merely as a predictive model [5].
Table 2 presents a conceptual evaluation framework for assessing AI-driven clinical documentation improvement across clinical, operational, and financial dimensions.
Table 2. Conceptual Framework for Evaluating AI-Driven Clinical Documentation Improvement Across Clinical, Operational, and Financial Dimensions
Evaluation Dimension | Subcomponents | Current Measurement Approach | Limitation Identified in Review | Recommended Future Metric |
Clinical Documentation Quality | Completeness, accuracy, conciseness | Note length, redundancy, concordance | Weak linkage to clinical outcomes | Composite clinical documentation quality index |
Coding Reliability | Code accuracy, specificity, consistency | Classification metrics (F1, precision) | Lack of real-world coding validation | Agreement with expert coders + audit outcomes |
Billing Integrity | Medical necessity, compliance, denial rates | Proxy via coding accuracy | Rare direct financial evaluation | Claim acceptance rate, denial reduction, revenue delta |
Workflow Efficiency | Time in EHR, after-hours work, task switching | EHR logs, time-motion studies | Limited longitudinal evidence | Sustained time savings and cognitive load indices |
Human-AI Interaction | Acceptance, override, editing burden | Rarely measured | Critical gap in usability evidence | Acceptance rate, override taxonomy, editing time |
Deployment Maturity | Retrospective vs prospective use | Dataset-based validation | Lack of real-world implementation studies | Multi-site prospective trials |
Equity & Bias | Performance across populations | Rare subgroup analysis | Underdeveloped domain | Stratified performance and fairness metrics |
Human-AI collaboration remains under-studied in clinical documentation improvement. Explainable coding models and automated note systems can provide suggestions, but the practical value depends on whether physicians and coders accept, revise, or override those outputs [8, 26]. Studies of documentation variation and concordance suggest that clinicians already differ in how they document care, which may influence how they respond to CDI prompts or generated text [27, 28]. Future evaluation should therefore examine acceptance rates, override reasons, editing burden, and the effect of suggestions on clinical reasoning and documentation ownership [15].
AI-generated clinical documentation raises ethical and regulatory questions related to accountability, documentation integrity, privacy, bias, and auditability. Speech recognition and automated note generation can introduce errors or omissions that may affect patient safety, coding, billing, and medico-legal responsibility [14, 16]. Automated coding and documentation-code alignment systems may also reflect historical documentation and coding practices, including local biases and incomplete labels [10, 11]. These issues suggest that CDI systems require transparent audit trails, human oversight, monitoring for drift, and governance processes that extend beyond model development [6, 7].
This review was limited by the heterogeneity of clinical documentation tasks, model architectures, datasets, clinical settings, and outcome measures. Because automated coding, note quality, billing support, and workflow efficiency were evaluated using different endpoints, a quantitative meta-analysis was not appropriate [7, 10]. The English-language focus may have underrepresented non-English documentation systems, although several studies demonstrated that language and coding-system context can materially influence model transferability [20]. The review also relied on published literature, which may overrepresent technically successful systems and underrepresent failed deployments or commercially sensitive billing outcomes [15].
The evidence base was strongest for retrospective automated coding and weaker for prospective workflow, billing, and implementation outcomes. Many studies used historical datasets or controlled evaluations rather than real-world deployment, limiting conclusions about clinician adoption, coder trust, revenue-cycle impact, and long-term sustainability [8-10]. Documentation burden and note quality studies provided important context but did not always test AI-CDI interventions directly [1, 3, 4]. As a result, the current literature supports cautious optimism about AI for CDI while also showing the need for independent replication, prospective evaluation, and multidimensional outcome assessment [5, 28].
Prior reviews helped establish the foundations of clinical natural language processing, but most addressed either broad clinical text processing or a narrower technical area rather than clinical documentation improvement as an integrated field. Reviews of clinical information extraction, chronic disease NLP, and clinical text machine learning described major progress in extracting structured information from notes, but they did not consistently connect extraction performance to coding, billing, documentation quality, and physician workflow outcomes [6, 7, 17, 21]. Similarly, reviews and methodological papers on automated coding emphasised the growth of ICD prediction models but gave less attention to note bloat, ambient documentation, or revenue-cycle implementation [10]. This review therefore extends prior work by organising the evidence around CDI functions rather than around NLP task type alone.
Some prior evidence focused on speech recognition, automated scribes, and medical conversation summarisation as solutions to physician documentation burden. These studies were important because they moved beyond retrospective note classification and examined documentation as a workflow problem involving conversation capture, report generation, and clinician review [15, 16, 18, 19]. However, this literature often remained separate from automated coding research, even though generated notes may directly influence coding specificity, medical necessity, billing compliance, and audit defensibility [13, 14]. By bringing these areas together, the present review highlights that documentation generation and coding support should be evaluated as connected parts of the same CDI ecosystem.
The novel contribution of this review is the emphasis on implementation gaps and multidimensional outcome evaluation. Automated coding studies often reported technical metrics, while documentation burden studies reported time or workload proxies, and note quality studies reported length, redundancy, concordance, or error patterns [1, 3, 8, 27]. Fewer studies connected these outcomes in a way that would allow health systems to determine whether an AI-CDI tool improves documentation quality, physician efficiency, coding reliability, and billing performance simultaneously [4, 26]. This synthesis therefore suggests that future CDI evidence should move from isolated model validation toward integrated clinical, operational, and financial evaluation [5, 29].
A major gap is the absence of long-term studies showing whether AI-CDI tools improve patient outcomes, documentation reliability, physician well-being, or sustained revenue-cycle performance. Automated coding systems have been extensively evaluated on historical notes, but few studies followed their use through coder review, claim submission, payer response, audit processes, and institutional financial outcomes [13, 23, 29]. Workflow and note-generation studies similarly require longer follow-up to determine whether initial time savings persist after novelty effects, training periods, and workflow adaptations [15, 19]. Future studies should therefore track clinical, financial, and operational outcomes together rather than treating them as separate endpoints [1, 4].
The literature provides limited evidence on how physicians and coders interact with AI-CDI suggestions in routine practice. Explainable coding models can highlight relevant note evidence, but the field still lacks detailed evaluation of when users accept, ignore, modify, or distrust suggestions [8, 24, 26]. This matters because poorly timed or poorly justified prompts may increase cognitive load, while well-designed suggestions could support more complete and specific documentation [27, 28]. Future research should examine trust calibration, override behaviour, editing burden, and the effect of AI suggestions on clinical reasoning and documentation habits [5, 15].
Equity and bias remain underdeveloped topics in AI-CDI research. Models trained on historical notes and codes may reproduce existing differences in documentation intensity, access to care, diagnostic labelling, or coding practices across patient populations [10, 11]. Automated note generation may also amplify biases if speech recognition, summarisation, or template selection performs differently across accents, languages, specialties, or clinical contexts [14, 16, 18]. Future evaluations should therefore stratify performance across demographic groups, clinical settings, and language contexts while also examining whether AI-CDI changes documentation quality or billing outcomes unevenly [21, 20].
The main implication for research practice is the need to shift from isolated accuracy studies to implementation-oriented evaluation across the four CDI domains. Automated coding models, note generation systems, note quality tools, and workflow interventions should be studied together because each can affect the others in clinical practice [8, 10, 18]. For example, a generated note may reduce physician typing time but still create coding ambiguity, billing risk, or downstream review burden if it lacks specificity or contains summarisation errors [14, 19]. Implementation science frameworks would help researchers examine adoption, fidelity, user adaptation, and sustainability in addition to model performance [4, 15].
For clinical practice, AI-CDI tools should be treated as assistive systems rather than autonomous replacements for physicians, coders, or CDI specialists. The included literature supports the technical feasibility of code prediction, speech-derived documentation, and clinical text summarisation, but it also shows persistent challenges in error detection, interpretability, integration, and user trust [13, 16, 26]. Human oversight remains critical because documentation is both a clinical communication tool and a legal and billing record [14, 27]. Health systems adopting these tools should therefore define review responsibilities, escalation processes, and monitoring procedures before broad deployment [5, 10].
Natural language processing for clinical documentation improvement progressed rapidly between 2017 and 2023, especially in automated clinical coding and note analysis. The reviewed literature shows that AI methods can extract, classify, summarise, and generate documentation-relevant information across a variety of clinical contexts.
Billing support and physician workflow efficiency remain less mature as evidence domains. Although many systems have plausible value for revenue-cycle support and documentation burden reduction, fewer studies have evaluated real-world billing outcomes, sustained physician time savings, or long-term implementation effects.
A coordinated research agenda is needed to standardise evaluation, improve external validation, measure implementation outcomes, and assess equity. Future studies should connect coding accuracy, note quality, billing integrity, and workflow impact rather than evaluating each domain in isolation.
The future of clinical documentation improvement lies not in replacing physicians, coders, or CDI specialists, but in augmenting their work with trustworthy, context-aware, auditable AI. The most valuable systems will be those that improve documentation while preserving clinical judgment, professional accountability, and patient-centred care.
None
None
None
None
Open Access The author(s) retain copyright. This article is licensed under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License. It may be shared and adapted for non-commercial purposes with appropriate attribution, an indication of changes, and distribution of adaptations under the same license. Third-party material may be subject to separate terms identified in its credit line. View the license at https://creativecommons.org/licenses/by-nc-sa/4.0/.