The Berlin definition of ARDS provides standardized diagnostic criteria based on acute onset within one week of a known insult, bilateral chest imaging opacities not explained by other causes, respiratory failure not due to cardiac issues or fluid overload, and impaired oxygenation measured by the PaO₂/FiO₂ ratio, enabling consistent identification in intensive care; however, its clinical use is limited by variability in imaging interpretation and the need for rapid decision-making, often causing delays and inconsistent diagnoses. Current practice relies heavily on subjective assessment of chest X-rays and limited integration of clinical notes and laboratory trends, resulting in moderate inter-observer agreement and reduced diagnostic reliability. To overcome these challenges, a multimodal transformer framework is proposed that integrates chest X-rays, clinical notes, and laboratory data using vision transformers, BERT-based text encoders, and temporally aware lab embeddings, with cross-modal attention enabling interaction across data types and a fusion module producing final ARDS probability estimates. This integrated approach improves diagnostic accuracy by combining complementary information, enhances interpretability through attention mechanisms, and offers a more objective and timely method for ARDS detection, with potential to support earlier intervention and better outcomes in critically ill patients.
Neovascular age-related macular degeneration (wet AMD) is the most severe form of AMD, driven by choroidal neovascularization that can cause rapid, irreversible central vision loss. Early anti-VEGF treatment preserves vision, making timely identification of progression from intermediate AMD critically important. However, current surveillance methods are insufficient for accurately predicting which patients will convert to neovascular disease. Existing prediction models rely mainly on a single imaging modality such as fundus photography or optical coherence tomography (OCT), limiting their ability to capture the full spectrum of disease features. Important clinical factors—age, genetics, and lifestyle—are also often underused. This lack of integrated multimodal modeling limits accurate risk stratification. We propose a multimodal transformer framework that integrates fundus images, OCT volumes, and clinical variables to predict progression from intermediate to neovascular AMD. Modality-specific encoders convert each data type into unified token representations, which are then fused using a cross-modal transformer to generate a calibrated progression risk score. The system includes a vision transformer-based fundus encoder, a 3D OCT volume encoder, a clinical variable MLP encoder, a cross-modal attention module for information fusion, and a classifier that outputs time-to-neovascular conversion risk. The framework learns shared representations across modalities, enabling interaction between imaging biomarkers and clinical risk factors. Cross-modal attention helps uncover complex patterns that may precede neovascularization and are not visible in single-modality models. This framework enables integrated, multimodal risk prediction for AMD progression, offering a foundation for personalized monitoring and earlier intervention through improved risk stratification.