Logo image
Seeing Roads Through Words: A Language-Guided Framework for RGB-T Driving Scene Segmentation
Conference paper   Peer reviewed

Seeing Roads Through Words: A Language-Guided Framework for RGB-T Driving Scene Segmentation

Raturaj Reddy, Hrishav Bakul Barua, Junn Yong Loo, Thanh Thi Nguyen and Ganesh Krishnasamy
Proceedings of the 2026 IEEE International Conference on Image Processing (ICIP), pp.1-6
IEEE International Conference on Image Processing (ICIP), 2026 (Tampere, Finland, 13-Aug-2026–17-Aug-2026)
Institute of Electrical and Electronics Engineers
2026

Abstract

Lighting Printing Modeling Color Decoding Equations Semantic segmentation Conferences Labeling Pixel
Robust semantic segmentation of road scenes under adverse illumination, lighting, and shadow conditions remain a core challenge for autonomous driving applications. RGB-Thermal fusion is a standard approach, yet existing methods apply static fusion strategies uniformly across all conditions, allowing modality-specific noise to propagate throughout the network. Hence, we propose CLARITY that dynamically adapts its fusion strategy to the detected scene condition. Guided by vision-language model (VLM) priors, the network learns to modulate each modality’s contribution based on the illumination state while leveraging object embeddings for segmentation, rather than applying a fixed fusion policy. We further introduce two mechanisms - one which preserves valid dark-object semantics that prior noise-suppression methods incorrectly discard, and a hierarchical decoder that enforces structural consistency across scales to sharpen boundaries on thin objects. Experiments on the MFNet dataset demonstrate that CLARITY establishes a new state-of-the-art (SOTA), achieving 62.3% mIoU and 77.5% mAcc.

Details

Metrics

1 Record Views
Logo image