A Hybrid Deep Learning Framework for Breast Cancer Morphology Classification: Integrating DCGAN-based Augmentation, Attention Mechanisms, and Self-Supervised Transfer Learning
نویسندگان
1 Ph.D. Student, Electronics Engineering, Faculty of Engineering, Lorestan University, Lorestan, Iran
2 Associate Professor, Department of Electronics Engineering, Faculty of Engineering, Lorestan University, Lorestan, Iran
3 Associate Professor, Department of Electronics Engineering, Faculty of Engineering, Mohaghegh Ardabili University, Ardabil, Iran
doi
چکیده
Accurate breast cancer morphology classification in mammography faces challenges from scarce annotated datasets and class imbalance. This study introduces a novel deep learning framework to overcome these limitations, integrating Deep Convolutional Generative Adversarial Networks (DCGAN) for data augmentation, a hybrid YOLOv8s architecture with attention mechanisms, and a self-supervised transfer learning strategy. A core component is expanding the MIAS dataset. Using DCGAN, we generated 5,000 synthetic mammograms, mitigating class imbalance, ensuring balanced representation, and reducing overfitting risk. The DCGAN architecture showed stable training and effectively captured high-frequency lesion details. To enhance discriminative power, we integrated Convolutional Block Attention Modules (CBAM) into YOLOv8s. CBAM applies channel and spatial attention, critical for noise suppression and sharpening lesion boundaries. A "Dynamic Head Wrapper" seamlessly integrates CBAM layers, adapting dynamically to feature dimensions. Our framework also includes a self-supervised learning (SSL) phase using SimCLR for backbone pre-training on 20,000 unlabeled in-domain VinDr-Mammo images. This pre-training enabled the model to learn robust, domain-invariant features and intrinsic tissue representations, providing an effective "warm start." For clinical validity, evaluation used a hybrid test set of real MIAS and external CBIS-DDSM images, without synthetic data exposure. The proposed method achieved a robust mean Average Precision (mAP)@0.5 of 0.969, significantly outperforming standard YOLOv8s and cross-domain transfer learning. These results confirm that synergistic integration of in-domain self-supervision and attention mechanisms improves diagnostic accuracy and clinical reliability.