Adverse drug reactions (ADRs) are a major global health issue, contributing to significant morbidity, mortality, and healthcare costs. Many ADRs are detected only after widespread drug use, reflecting limitations of pre-market trials in capturing real-world patient variability. Although electronic health records (EHRs) collected between 2017 and 2023 provide rich data for post-market surveillance, they remain underused for systematic ADR detection. Current pharmacovigilance methods rely heavily on spontaneous reporting systems, which suffer from underreporting and bias, while supervised machine learning approaches require labeled ADR data that are often unavailable for rare or novel events. This paper proposes a variational autoencoder (VAE)-based unsupervised framework to detect ADR signals from multimodal EHR data, including clinical notes and laboratory results. The model learns normal patient data distributions and identifies deviations as potential safety signals without requiring labeled ADR examples. A multimodal architecture combines natural language processing of clinical notes with structured laboratory encoders, forming a shared latent space for anomaly detection based on reconstruction error. The framework enables detection of unknown ADRs by flagging abnormal patterns in patient records across large datasets from 2017 to 2023. Its unsupervised nature makes it suitable for identifying rare or previously unrecognized drug safety issues. Overall, this approach offers a scalable, proactive pharmacovigilance strategy that shifts drug safety monitoring from reactive reporting to predictive detection using routine EHR data.
Hypertension affects about 1.4 billion adults globally and is a major modifiable risk factor for cardiovascular disease. Although several first-line antihypertensive drug classes exist, randomized controlled trials typically report only average treatment effects (ATEs), which mask important variability in individual patient responses. As a result, clinical guidelines often assume a homogeneous patient population, leading to trial-and-error prescribing, delayed blood pressure control, and avoidable adverse effects. I argue that causal forest models combined with double machine learning (DML) enable reliable estimation of heterogeneous treatment effects (HTEs) from observational electronic health record data. These methods can approximate randomized trial validity while capturing clinically meaningful variation in treatment response across patients. Compared with traditional approaches, they are computationally feasible and better suited for individualized treatment assessment. Therefore, comparative effectiveness research in hypertension should move beyond ATE-focused analyses toward routine HTE estimation using causal machine learning. This shift would support more precise, data-driven prescribing and improve patient outcomes.
Telemedicine expanded rapidly in the United States during the COVID-19 public health emergency as Medicare and state Medicaid programs relaxed coverage restrictions. Diabetes affects about 37 million Americans, and key outcomes such as HbA1c, blood pressure, and LDL cholesterol are routinely tracked in electronic health records. However, the causal impact of telemedicine expansion on these outcomes remains uncertain, as simple pre–post comparisons are confounded by concurrent trends such as the pandemic and seasonal variation. Randomized policy experiments are impractical, leaving a gap in high-quality causal evidence. We argue that Bayesian structural time series (BSTS) applied to state-level EHR aggregates provides a strong alternative. BSTS constructs a synthetic counterfactual from similar states, modeling trend and seasonality to estimate what outcomes would have been without telemedicine expansion. This allows clearer separation of policy effects from underlying time dynamics and produces interpretable estimates with uncertainty bounds. Unlike difference-in-differences, BSTS does not rely on parallel trends assumptions that may be violated in this context. It offers a transparent framework for causal inference using routinely available aggregated data. Policymakers should prioritize such causal methods when evaluating whether telemedicine expansions should become permanent rather than relying on descriptive before–after analyses.
Telemedicine expanded rapidly during COVID-19 as Medicare and states relaxed long-standing restrictions. Diabetes outcomes (HbA1c, blood pressure, LDL cholesterol) are routinely tracked in EHRs, yet causal evidence that telemedicine improves these outcomes remains limited. Pre-post analyses cannot separate telemedicine effects from confounding time trends such as seasonality and pandemic-related changes, while randomized state-level policy trials are infeasible. This leaves a key evidence gap for policy decisions. Bayesian structural time series (BSTS) using state-level EHR aggregates is the most suitable approach for causal inference, constructing counterfactual outcomes from similar donor states. BSTS accounts for trends, seasonality, and autocorrelation, and uses synthetic control principles to reduce confounding. It also provides uncertainty estimates and works with routinely collected aggregate data. Policymakers should require BSTS-based evidence before making telemedicine coverage permanent. Researchers should apply these methods to existing policy variation and share data and code. States should build routine EHR-based monitoring systems for diabetes outcomes. Causal evaluation of telemedicine is feasible now using existing data and methods. Relying on pre-post studies or waiting for randomized trials delays actionable evidence needed for policy decisions.