Hierarchical Domain Adaptation with Multi-Level Attention for RGB-Thermal Object Detection

Authors

  • Yunan Liu College of Computer Science and Software Engineering, Shenzhen University, Shenzhen 518060, China Author
  • Mingrong Gong College of Computing and Data Science, Nanyang Technological University, Singapore 639798, Singapore Author
  • Chaoqi Chen College of Computer Science and Software Engineering, Shenzhen University, Shenzhen 518060, China Author

DOI:

https://doi.org/10.64509/jicn.23.127

Keywords:

Domain Adaptation, RGB-Thermal Object Detection, Hierarchical Attention, Adversarial Learning, Faster R-CNN

Abstract

Cross-modal domain adaptation for RGB-to-Thermal object detection (DAOD) remains a challenging task due to the large modality gap, feature distribution mismatch, and limited supervision on the thermal domain. Conventional domain adaptation detectors often fail to capture consistent semantic representations across modalities, resulting in unstable adversarial learning and suboptimal detection accuracy. To address these challenges, we propose HDAM (Hierarchical Domain Adaptation with Multi-Level Attention), an enhanced domain-adaptive detection framework built upon Faster R-CNN. HDAM progressively aligns cross-domain features from low- to high-level representations through a hierarchical attention mechanism. Specifically, it comprises three key components: (1) Entropy-Guided Cross-level Attention (ECA), which leverages discriminator-derived entropy responses to reweight pixel- and mid-level features for more stable cross-modal encoding; (2) Multi-branch Global Semantics and Discriminator (MGSD), which performs global adversarial alignment and integrates multi-scale spatial information through attention-enhanced branches, while introducing an auxiliary image-level semantic consistency loss; and (3) Instance-Context Alignment (ICA), which fuses ROI-level and contextual embeddings and applies dual instance-level adversarial losses for precise fine-grained adaptation. Extensive experiments on benchmark RGB-Thermal datasets demonstrate that HDAM effectively bridges the modality gap and achieves improved cross-domain detection performance compared to recent DAOD methods.

Downloads

Download data is not yet available.

References

[1] Torralba, A., Efros, A.A.: Unbiased look at dataset bias. In Conference on Computer Vision and Pattern Recognition (CVPR) 2011, pp. 1521–1528 (2011). https://doi.org/10.1109/CVPR.2011.5995347

[2] Dai, Y., Wu, Y., Zhou, F., Barnard, K.: Attentional local contrast networks for infrared small target detection. IEEE transactions on geoscience and remote sensing 59(11), 9813–9824 (2021). https://doi.org/10.1109/TGRS.2020.3044958

[3] Shi, C., Zheng, Y., Chen, Z.: Domain adaptive thermal object detection with unbiased granularity alignment. ACM Transactions on Multimedia Computing, Communications and Applications 20(9), 1–23 (2024). https://doi.org/10.1145/3665892

[4] Berjawi, J., Dupas, Y., C´erin, C.: Towards a Generalizable Fusion Architecture for Multimodal Object Detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 2192–2200 (2025). https://doi.org/10.1109/ICCVW69036.2025.00231

[5] Pan, S.J., Yang, Q.: A Survey on Transfer Learning. IEEE Transactions on Knowledge and Data Engineering 22(10), 1345–1359 (2010). https://doi.org/10.1109/TKDE.2009.191

[6] Oza, P., Sindagi, V.A., VS, V., Patel, V.M.: Unsupervised domain adaptation of object detectors: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence 46(6), 4018–4040 (2023). https://doi.org/10.1109/TPAMI.2022.3217046

[7] Wilson, G., Cook, D.J.: A survey of unsupervised deep domain adaptation. ACM Transactions on Intelligent Systems and Technology (TIST) 11(5), 1–46 (2020). https://doi.org/10.1145/3400066

[8] Ganin, Y., Lempitsky, V.: Unsupervised domain adaptation by backpropagation. In Proceedings of the 32nd International Conference on Machine Learning, pp. 1180–1189 (2015)

[9] Long, M., Cao, Y., Wang, J., Jordan, M.: Learning transferable features with deep adaptation networks. In Proceedings of the 32nd International Conference on Machine Learning, pp. 97–105 (2015)

[10] Sun, B., Saenko, K.: Deep coral: Correlation alignment for deep domain adaptation. In European Conference on Computer Vision, pp. 443–450 (2016). https://doi.org/10.1007/978-3-319-49409-8 35

[11] Chen, Y., Li, W., Sakaridis, C., Dai, D., Van Gool, L.: Domain adaptive faster r-cnn for object detection in the wild. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 3339–3348 (2018). https://doi.org/10.1109/CVPR.2018.00352

[12] Saito, K., Ushiku, Y., Harada, T., Saenko, K.: Strongweak distribution alignment for adaptive object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 6956–6965 (2019). https://doi.org/10.1109/CVPR.2019.00712

[13] Kim, Y., Cho, D., Han, K., Panda, P., Hong, S.: Domain adaptation without source data. IEEE Transactions on Artificial Intelligence 2(6), 508–518 (2021). https://doi.org/10.1109/TAI.2021.3110179

[14] Peng, X., Bai, Q., Xia, X., Huang, Z., Saenko, K., Wang, B.: Moment matching for multi-source domain adaptation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 1406–1415 (2019). https://doi.org/10.1109/ICCV.2019.00149

[15] Do, D.P., Kim, T., Na, J., Kim, J., Lee, K., Cho, K., Hwang, W.: D3t: Distinctive dual-domain teacher zigzagging across rgb-thermal gap for domain-adaptive object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 23313–23322 (2024). https://doi.org/10.1109/CVPR52733.2024.02200

[16] Li, H., Zhang, R., Yao, H., Zhang, X., Hao, Y., Song, X., Peng, S., Zhao, Y., Zhao, C., Wu, Y., et al.: SEEN-DA: SEmantic ENtropy guided Domain-aware Attention for Domain Adaptive Object Detection. In Proceedings of the Computer Vision and Pattern Recognition Conference, pp. 25465–25475 (2025). https://doi.org/10.1109/CVPR52734.2025.02371

[17] Zhou, Y., Zhang, H., Sun, Y., Liu, Y., Kung, S.-Y.: Illumination Guided Domain Adaptation Object Detection in Thermal Imagery. Neurocomputing 653, 131237 (2025). https://doi.org/10.1016/j.neucom.2025.131237

[18] Yu, H., Deng, J., Li, W., Duan, L.: Towards unsupervised model selection for domain adaptive object detection. In Proceedings of the 38th International Conference on Neural Information Processing Systems, pp. 58423–58444 (2024)

[19] Rivadeneira, R.E., Sappa, A.D., Wang, C., Jiang, J., Zhong, Z., Chen, P., Wang, S.: Thermal Image SuperResolution Challenge Results - PBVS 2024. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pp. 3113–3122 (2024). https://doi.org/10.1109/CVPRW63382.2024.00317

[20] Xu, M., Wang, H., Ni, B., Tian, Q., Zhang, W.: Cross-domain detection via graph-induced prototype alignment. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 12355–12364 (2020). https://doi.org/10.1109/CVPR42600.2020.01174

[21] Medeiros, H.R., Belal, A., Muralidharan, S., Granger, E., Pedersoli, M.: Visual Modality Prompt for Adapting Vision-Language Object Detectors. In 2025 IEEE/CVF International Conference on Computer Vision (ICCV), pp. 2172–2182 (2025). https://doi.org/10.1109/ICCV51701.2025.00210

[22] Xu, C.-D., Zhao, X.-R., Jin, X., Wei, X.-S.: Exploring Categorical Regularization for Domain Adaptive Object Detection. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 11721–11730 (2020). https://doi.org/10.1109/CVPR42600.2020.01174

[23] Zellinger, W., Grubinger, T., Lughofer, E., Natschl¨ager, T., Saminger-Platz, S.: Central moment discrepancy (CMD) for domain-invariant representation learning. arXiv preprint arXiv:1702.08811 (2017). https://doi.org/10.48550/arXiv.1702.08811

[24] Gretton, A., Borgwardt, K.M., Rasch, M.J., Sch¨olkopf, B., Smola, A.: A kernel two-sample test. The journal of machine learning research 13(1), 723–773 (2012)

[25] Na, J., Jung, H., Chang, H.J., Hwang, W.: FixBi: Bridging Domain Spaces for Unsupervised Domain Adaptation. In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 1094–1103 (2021). https://doi.org/10.1109/CVPR46437.2021.00115

[26] Ren, S., He, K., Girshick, R., Sun, J.: Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks. IEEE Transactions on Pattern Analysis and Machine Intelligence 39(6), 1137–1149 (2017). https://doi.org/10.1109/TPAMI.2016.2577031

[27] He, Z., Zhang, L.: Multi-Adversarial Faster-RCNN for Unrestricted Object Detection. In 2019 IEEE/CVF International Conference on Computer Vision (ICCV), pp. 6667–6676 (2019). https://doi.org/10.1109/ICCV.2019.00677

[28] Jiao, L., Wei, H., Pan, Q.: Region and Sample Level Domain Adaptation for Unsupervised Infrared Target Detection in Aerial Remote Sensing Images. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing 18, 11289–11306 (2025). https://doi.org/10.1109/JSTARS.2025.3561737

[29] Neuwirth–Trapp, M., Bieshaar, M., Paudel, D.P., Van Gool, L.: Incremental Object Detection with Prompt-Based Methods. In 2025 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), pp. 5224–5232 (2025). https://doi.org/10.1109/ICCVW69036.2025.00544

[30] Li, T., Ye, M., Wu, T., Li, N., Li, S., Tang, S., Ji, L.: Pseudo Visible Feature Fine-Grained Fusion for Thermal Object Detection. In 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 6710–6719 (2025). https://doi.org/10.1109/CVPR52734.2025.00629

[31] Peng, H., Hu, Y., Yu, B., Zhang, Z.: TCAINet an RGB T salient object detection model with cross modal fusion and adaptive decoding. Scientific Reports 15(1), 14266 (2025). https://doi.org/10.1038/s41598-025-98423-z

[32] Roy, S., Krivosheev, E., Zhong, Z., Sebe, N., Ricci, E.: Curriculum Graph Co-Teaching for Multi-Target Domain Adaptation. In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 5347–5356 (2021). https://doi.org/10.1109/CVPR46437.2021.00531

[33] Chen, Y., Wang, H., Li, W., Sakaridis, C., Dai, D., Van Gool, L.: Scale-Aware Domain Adaptive Faster R-CNN. International Journal of Computer Vision 129(7), 2223–2243 (2021). https://doi.org/10.1007/s11263-021-01447-x

[34] Chen, H.-Y., Chao, W.-L.: Gradual domain adaptation without indexed intermediate domains. In Proceedings of the 35th International Conference on Neural Information Processing Systems, pp. 8201–8214 (2021)

[35] Wu, Z., Wang, X., Gonzalez, J., Goldstein, T., Davis, L.: ACE: Adapting to Changing Environments for Semantic Segmentation. In 2019 IEEE/CVF International Conference on Computer Vision (ICCV), pp. 2121–2130 (2019). https://doi.org/10.1109/ICCV.2019.00221

[36] Guan, D., Huang, J., Xiao, A., Lu, S., Cao, Y.: Uncertainty-Aware Unsupervised Domain Adaptation in Object Detection. IEEE Transactions on Multimedia 24, 2502–2514 (2022). https://doi.org/10.1109/TMM.2021.3082687

[37] Chen, C., Zheng, Z., Ding, X., Huang, Y., Dou, Q.: Harmonizing transferability and discriminability for adapting object detectors. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 8869–8878 (2020). https://doi.org/10.1109/CVPR42600.2020.00889

[38] Zhang, Y., Wang, Z., Li, J., Zhuang, J., Lin, Z.: Towards effective instance discrimination contrastive loss for unsupervised domain adaptation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 11388–11399 (2023). https://doi.org/10.1109/ICCV51070.2023.01046

[39] Tzeng, E., Hoffman, J., Saenko, K., Darrell, T.: Adversarial discriminative domain adaptation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 7167–7176 (2017). https://doi.org/10.1109/CVPR.2017.316

[40] Sun, T., Lu, C., Ling, H.: Local context-aware active domain adaptation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 18634–18643 (2023). https://doi.org/10.1109/ICCV51070.2023.01708

[41] Lin, T.-Y., Goyal, P., Girshick, R., He, K., Doll´ar, P.: Focal Loss for Dense Object Detection. In 2017 IEEE International Conference on Computer Vision (ICCV), pp. 2999–3007 (2017). https://doi.org/10.1109/ICCV.2017.324

[42] Zhang, H., Fromont, E., Lefevre, S., Avignon, B.: Multispectral Fusion for Object Detection with Cyclic Fuseand-Refine Blocks. In 2020 IEEE International Conference on Image Processing (ICIP), pp. 276–280 (2020). https://doi.org/10.1109/ICIP40778.2020.9191080

[43] Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R., Franke, U., Roth, S., Schiele, B.: The Cityscapes Dataset for Semantic Urban Scene Understanding. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 3213–3223 (2016). https://doi.org/10.1109/CVPR.2016.350

[44] Everingham, M., Van Gool, L., Williams, C.K., Winn, J., Zisserman, A.: The pascal visual object classes (voc) challenge. International journal of computer vision 88(2), 303–338 (2010). https://doi.org/10.1007/s11263-009-0275-4

[45] He, K., Zhang, X., Ren, S., Sun, J.: Deep Residual Learning for Image Recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 770–778 (2016). https://doi.org/10.1109/CVPR.2016.90

[46] Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., Fei-Fei, L.: Imagenet: A large-scale hierarchical image database. In 2009 IEEE Conference on Computer Vision and Pattern Recognition, pp. 248–255 (2009). https://doi.org/10.1109/CVPR.2009.5206848

[47] He, K., Gkioxari, G., Doll´ar, P., Girshick, R.: Mask r-cnn. In Proceedings of the IEEE International Conference on Computer Vision, pp. 2961–2969 (2017). https://doi.org/10.1109/TPAMI.2018.2844175

[48] Shen, Z., Maheshwari, H., Yao, W., Savvides, M.: Scl: Towards accurate domain adaptive object detection via gradient detach based stacked complementary losses. arXiv preprint arXiv:1911.02559 (2019). https://doi.org/10.48550/arXiv.1911.02559

[49] Wu, A., Liu, R., Han, Y., Zhu, L., Yang, Y.: VectorDecomposed Disentanglement for Domain-Invariant Object Detection. In 2021 IEEE/CVF International Conference on Computer Vision (ICCV), pp. 9322–9331 (2021). https://doi.org/10.1109/ICCV48922.2021.00921

[50] Marnissi, M.A., Fradi, H., Sahbani, A., Essoukri Ben Amara, N.: Feature distribution alignments for object detection in the thermal domain. The Visual Computer 39(3), 1081–1093 (2023). https://doi.org/10.1007/s00371-021-02386-x

[51] Zhao, L., Wang, L.: Task-specific inconsistency alignment for domain adaptive object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 14217–14226 (2022). https://doi.org/10.1109/CVPR52688.2022.01382

jicn127

Downloads

Published

2026-07-30

Issue

Section

Articles

How to Cite

Liu, Y., Gong, M., & Chen, C. (2026). Hierarchical Domain Adaptation with Multi-Level Attention for RGB-Thermal Object Detection. Journal of Intelligent Computing and Networking, 2(3), 1-13. https://doi.org/10.64509/jicn.23.127

Similar Articles

1-10 of 19

You may also start an advanced similarity search for this article.