Urban environmental visibility factors from street-view imagery using fine-tuned deep learning
نویسندگان
1 دانشگاه آزاد اسلامی
2 دانشگاه آزاد اسلامی
3 دانشگاه آزاد اسلامی
4 دانشگاه زابل
5 دانشگاه آزاد اسلامی
doi
10.22034/gjesm.2026.04.05چکیده
BACKGROUND AND OBJECTIVES: Street-level urban view factors are important indicators for assessing urban form. Environmental conditions and greenery exposure ultimately dictate residents’ quality of life; however, robust evaluation frameworks for assessing these determinants are still limited. This limitation is primarily due to the challenges in collecting extensive field data across large urban areas. To address this limitation, the objectives of this study were to leverage existing street-view imagery. This approach can be a cost-effective and widely available data source for urban environmental assessment. This study introduces an evaluative framework for quantifying building, tree, and sky view factors through the application of advanced deep learning models. METHODS: A DeepLabV3+ model, pre-trained on the Cityscapes dataset, was selected for this study. To identify critical urban landscape features (buildings, trees, and sky), the model underwent localized fine-tuning aimed at enhancing predictive performance. In the fine-tuning process, manually annotated street-view images collected directly from the study area were used. The historic city center of Chiang Mai, Thailand, was selected as the research site. Street-view images were collected using the Google Street View Static application programming interface. To enhance segmentation performance under local urban conditions, several data augmentation strategies were examined and benchmarked against a baseline model. FINDINGS: The site-specific fine-tuning significantly enhanced model performance. The results demonstrated that combining various data augmentation strategies provides greater robustness than single-method approaches. The multi-augmentation model achieved an optimal mean intersection over union of 0.908. This metric reflected robust performance in the segmentation of buildings, trees, and sky features. Statistical analysis revealed a strong negative correlation between building view factor and tree view factor (ρ= -0.71, p < 0.001). This inverse relation reflects both urban spatial arrangements and intrinsic compositional class limitations. CONCLUSION: This study demonstrates the effectiveness of combining localized fine-tuning and multiple data-augmentation techniques to adapt DeepLabV3+ to the complex urban environment of Chiang Mai. The model improves urban view factor extraction accuracy. This study systematically evaluates localized fine-tuning and data-augmentation strategies and demonstrates that a relatively small, locally labelled dataset can effectively improve the performance of a pretrained model in a site-specific urban context.