Breast cancer remains a leading cause of cancer-related mortality among women worldwide, underscoring the importance of effective screening strategies for early detection and improved survival. Although conventional modalities such as mammography reduce mortality, they are limited by false positives, false negatives, and overdiagnosis, particularly in dense breast tissue and diverse populations. Deep learning, especially convolutional neural networks (CNNs), has shown promise in improving diagnostic accuracy and reducing inter-reader variability; however, its translation into routine clinical practice requires critical evaluation beyond reported performance metrics. This critical review evaluates CNN-based deep learning applications for breast cancer detection across mammography, ultrasound, and MRI, with emphasis on training strategies and barriers to clinical deployment. A targeted literature search identified peer-reviewed studies focusing on CNN architectures, transfer learning, and implementation challenges. Findings indicate that models such as ResNet, DenseNet, and EfficientNet perform well in controlled settings, supported by transfer learning and data augmentation approaches. However, these results often fail to translate into consistent clinical performance, particularly across imaging modalities and real-world workflows. Limitations including demographic bias, insufficient external validation, and weak evidence of outcome or cost-effectiveness highlight a substantial gap between experimental success and clinical readiness. The review concludes that while deep learning in breast imaging is promising, its adoption should remain cautious and evidence-driven until robust clinical benefit is clearly demonstrated.
Federated learning (FL) is promoted as a privacy-preserving method for training machine learning models across healthcare institutions without sharing patient data, with growing use in medical imaging, electronic health records, and rare disease research. This critical review examines FL studies from 2017–2024, focusing on privacy guarantees, statistical heterogeneity, communication efficiency, and real-world clinical deployment. A structured search of PubMed, IEEE Xplore, arXiv, and Google Scholar was conducted using relevant FL and healthcare terms, including studies addressing privacy, heterogeneity, communication, or deployment. Reported privacy guarantees are often overstated, with most studies relying on FedAvg without differential privacy. Statistical heterogeneity in non-IID settings remains largely unresolved. Fewer than 5% of studies report real-world deployment, typically at very small scale. A significant gap exists between FL research and clinical application. Current methods fall short of healthcare-grade privacy and real-world constraints, limiting readiness for high-stakes clinical use.
Postoperative complications including SSI (2–20%), VTE (1–5%), and respiratory failure (1–8%) significantly increase morbidity, mortality, length of stay, and readmissions. This systematic review assessed machine learning models predicting these outcomes, their performance, external validation, and clinical deployment. A PRISMA-based search (2017–2024) identified 32 eligible studies. Models such as random forest and XGBoost showed AUROC ranges of 0.70–0.85 for SSI, 0.75–0.90 for VTE (outperforming Caprini scores), and 0.75–0.88 for respiratory failure. However, fewer than 20% of studies included external validation and less than 5% reported clinical deployment. Overall, while machine learning models show strong retrospective performance, limited validation and minimal real-world implementation remain major barriers to clinical translation.