Text-Enhanced Semantic Segmentation via Contrastive Language-Image Pretraining Guided Multi-Modal Feature Fusion with Feature Refinement Approach

نویسندگان

1 Faculty of Computer Engineering, Shahrood University of Technology, Shahrood, Iran

2 Faculty of Computer Engineering, Shahrood University of Technology, Shahrood, Iran

3 School of Computer Engineering, Iran University of Science and Technology (IUST), Tehran, Iran

doi
10.5829/ije.2026.39.06c.11
چکیده

Image semantic segmentation is a fundamental task in computer vision. It is the process of pixel labeling and segmenting distinct parts of an image. It has wide applications in various scientific, medical, and industrial fields. Despite significant advancements, detailed and accurate segmentation remains challenging. In this paper, we propose a multi-modal semantic segmentation method to enrich feature maps. By taking both visual and textual features as multi-modal inputs, the model increases feature richness and achieves a more informative and detailed feature representation. In particular, ResNet-101 serves as the baseline model for extracting visual features. The Contrastive Language-Image Pretraining (CLIP) text encoder extracts textual features, making the representation multi-modal. Additionally, mid-level features from ResNet-101 are used to reduce information loss occurring in deep layers, thereby enhancing the image reconstruction process. The proposed model was tested on the COCO dataset, achieving a mean Intersection over Union (mIoU) of 64.44% on four categories: "person," "chair," "dining table," and "background." The effectiveness of the model is reflected in the proposed method. The proposed method outperforms the baseline DeepLabV3+ model, achieving a 7.87% improvement over its mIoU of 56.57%. These results underscore the potential of combining multi-modal image-text data and advanced attention mechanisms to enhance semantic segmentation performance.