Neovascular age-related macular degeneration (wet AMD) is the most severe form of AMD, driven by choroidal neovascularization that can cause rapid, irreversible central vision loss. Early anti-VEGF treatment preserves vision, making timely identification of progression from intermediate AMD critically important. However, current surveillance methods are insufficient for accurately predicting which patients will convert to neovascular disease. Existing prediction models rely mainly on a single imaging modality such as fundus photography or optical coherence tomography (OCT), limiting their ability to capture the full spectrum of disease features. Important clinical factors—age, genetics, and lifestyle—are also often underused. This lack of integrated multimodal modeling limits accurate risk stratification. We propose a multimodal transformer framework that integrates fundus images, OCT volumes, and clinical variables to predict progression from intermediate to neovascular AMD. Modality-specific encoders convert each data type into unified token representations, which are then fused using a cross-modal transformer to generate a calibrated progression risk score. The system includes a vision transformer-based fundus encoder, a 3D OCT volume encoder, a clinical variable MLP encoder, a cross-modal attention module for information fusion, and a classifier that outputs time-to-neovascular conversion risk. The framework learns shared representations across modalities, enabling interaction between imaging biomarkers and clinical risk factors. Cross-modal attention helps uncover complex patterns that may precede neovascularization and are not visible in single-modality models. This framework enables integrated, multimodal risk prediction for AMD progression, offering a foundation for personalized monitoring and earlier intervention through improved risk stratification.