Type 2 diabetes affects over 400 million people worldwide and requires lifelong management through continuous monitoring of laboratory values, medications, and comorbidities, yet the use of longitudinal electronic health records for research is restricted by privacy regulations such as HIPAA and GDPR, making synthetic data generation an important alternative for preserving utility while protecting confidentiality. However, existing synthetic data models often fail to accurately capture temporal treatment effects and the gradual development of comorbidities, limiting their usefulness for downstream clinical and machine learning applications. To address this, a time-series generative adversarial network is proposed for longitudinal diabetes data, incorporating a temporal encoder for irregular sampling, a treatment-conditioned generator, and dual discriminators that evaluate both static patient characteristics and dynamic clinical trajectories to ensure consistency between interventions and outcomes. By explicitly modeling temporal dependencies and comorbidity structures, the framework produces more realistic synthetic patient records that better reflect disease progression and medication-response relationships, thereby enabling privacy-preserving data sharing while supporting robust secondary analyses and future applications in chronic disease modeling.
Synthetic electronic health record (EHR) data generation has emerged as a potential solution to balancing clinical data accessibility with patient privacy, using generative artificial intelligence to simulate tabular, longitudinal, and textual health records without exposing identifiable patient information. This critical review, informed by PRISMA-ScR methodology, examines studies published between 2017 and 2025 focusing on generative models for synthetic EHR creation, with particular attention to privacy risks, data fidelity, downstream task utility, and ethical or regulatory considerations. A total of 67 studies were included after systematic screening, showing a dominance of GAN-based approaches alongside growing use of diffusion models and large language models in recent years, although privacy assessment and benchmarking practices remain inconsistent. Overall, the evidence suggests that while synthetic EHR data can facilitate data sharing, research, and model development, achieving a balance between realism, utility, and privacy remains challenging, as high statistical fidelity does not necessarily translate into clinical usefulness and strong downstream performance does not ensure adequate privacy protection.