Improving Visual Question Answering by Integrating Question-Specific Labels into Image Processing
نویسندگان
1 Faculty of Computer Engineering, Shahrood University of Technology, Shahrood, Iran
2 Faculty of Computer Engineering, Shahrood University of Technology, Shahrood, Iran.
3 Faculty of Computer Engineering, Shahrood University of Technology, Shahrood, Iran
doi
10.5829/ije.2026.39.12c.02چکیده
Visual Question Answering (VQA) combines computer vision and natural language processing to enable context-aware responses to image-based questions. We proposed a question-aware architecture that enhances visual feature extraction by generating semantically-driven, question-specific labels to guide image processing. Our model integrates ResNet-50 for visual feature extraction, BERT for textual processing, and an attention mechanism to fuse multimodal inputs. A dynamic label generation module aligns semantic attributes with visual regions, improving reasoning accuracy. This design further improves the model’s performance by carefully aligning visual and textual information to better understand and respond to a wide range of complex image-based questions. Evaluated on VQA v2, TextVQA, and VizWiz, our model achieves 73.3% accuracy on VQA v2, 0.57 ANLS on TextVQA, and 61.5% accuracy on VizWiz, surpassing baselines like LoRRA by up to 4% on VQA v2. Compared to state-of-the-art models, our model is competitive but requires improved OCR for text-heavy tasks and shows robustness on additional text-heavy benchmarks such as ST-VQA, ChartQA, and OCR-VQA. This architecture advances assistive technologies and medical diagnostics by enabling precise, context-aware visual reasoning.