Emergency department crowding is a persistent global healthcare challenge linked to longer wait times, increased patients leaving without being seen, worse clinical outcomes, and staff burnout. It also contributes to ambulance diversion and inefficient resource use, worsening hospital operational strain. This systematic review evaluates machine learning models for predicting ED crowding and optimizing patient flow, focusing on input features (e.g., arrival rates, acuity, bed availability) and reported operational outcomes such as waiting times and ambulance delays. A PRISMA-compliant review was conducted across PubMed, Embase, IEEE Xplore, and Scopus. Included studies applied machine learning to ED crowding or patient flow prediction and reported operational or crowding outcomes. Due to heterogeneity, a narrative synthesis was used, and risk of bias was assessed using an adapted tool. Thirty-two studies met inclusion criteria, using classification, regression, time-series, and deep learning models. Common predictors included arrival patterns, occupancy, and bed availability. While predictive performance was generally high, few studies evaluated real-world operational impacts, and most remained retrospective. Although machine learning models demonstrate strong predictive accuracy for ED crowding, evidence of real-world operational benefits remains limited. A clear gap exists between prediction and implementation into clinical workflow and decision-making. Future research should focus on translating predictions into measurable improvements in ED performance.