Garment Image Segmentation Method Based on Wavelet Multi-Scale Features and Improved U-Net
DOI:
https://doi.org/10.64509/jdi.13.103Keywords:
Clothing Image Segmentation, U-Net, Deeply Separable Convolution, Attention Mechanism, Haar Wavelet Down Sampling, Residual NetworkAbstract
Against the background of the gradual integration of fashion design and AIGC, obtaining accurate clothing masks through image segmentation and further guiding diffusion models for pattern generation and reconstruction serves as a key technique for intelligent design. Nevertheless, practical scenarios are plagued by issues such as complex background interference and occlusion between human bodies and garments, which easily lead to blurred segmentation boundaries and loss of details, thus impairing the subsequent design generation performance. To address these problems, this paper proposes a clothing image segmentation method based on an improved U-Net architecture. By fusing wavelet sampling and partial convolution in the downsampling stage, the feature learning of the encoder is optimized and the influence of occluded regions is reduced. Furthermore, the Efficient Local Attention-T (ELA-T) mechanism is embedded into the skip connections to enhance the model's sensitivity to subtle boundary features. A novel residual multi-scale depthwise separable convolution module (R-MSDW) is designed to strengthen the model's capability in capturing multi-scale features and hierarchical semantic information. Finally, experiments are conducted on the DeepFashion2 dataset and a self-constructed dataset. The proposed method outperforms mainstream models such as U-Net, SegNet, and U2-Net, achieving an mIoU of 86.1% and an overall accuracy of 95.1% on the DeepFashion2 dataset, which surpasses vanilla U-Net by 8.78% in mIoU and 2.7% in overall accuracy respectively, and an mIoU of 85.4% and an overall accuracy of 94.8% on the self-constructed dataset. The proposed model helps moderately enhance the visual quality of generated fashion design images.
Downloads
References
[1] Chen, B., Zhang, Y., Yu, B., Liu, X.: Two-stage adjustable perceptual distillation network for virtual tryon. Journal of Graphics 43(2), 316-323 (2022)
[2] Xu, Y.: Garment image segmentation technology and application based on convolutional neural network. Master's Thesis, Donghua University (2021)
[3] Zhang, Q., Liu, L., Fu, X., Liu, L., Huang, Q.: Clothing Image Retrieval by Label Optimization and Semantic Segmentation. Journal of Computer-Aided Design & Computer Graphics 32(9), 1450-1465 (2020). https://doi.org/10.3724/SP.J.1089.2020.18122
[4] Xu, H., Bai, M., Wan, T., Xue, T., Tang, W.: Image semantic analysis and retrieval recommendation for clothing based on deep learning. Fangzhi Gaoxiao Jichukexue Xuebao 33(3), 64-72 (2020). https://doi.org/10.13338/j.issn.1006-8341.2020.03.011
[5] Zhang, L., Rao, A., Agrawala, M.: Adding Conditional Control to Text-to-Image Diffusion Models. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 3836-3847 (2023)
[6] Sezgin, M., Sankur, B.: Survey over image thresholding techniques and quantitative performance evaluation. Journal of Electronic Imaging 13(1), 146-165 (2004). https://doi.org/10.1117/1.1631315
[7] Adams, R., Bischof, L.: Seeded region growing. IEEE Transactions on Pattern Analysis and Machine Intelligence 16(6), 641-647 (1994). https://doi.org/10.1109/34.295913
[8] Canny, J.: A Computational Approach to Edge Detection. IEEE Transactions on Pattern Analysis and Machine Intelligence PAMI-8(6), 679-698 (1986). https://doi.org/10.1109/TPAMI.1986.4767851
[9] Huang, P., Zheng, Q., Liang, C.: Overview of Image Segmentation Methods. Journal of Wuhan University (Natural Science Edition) 66(6), 519-531 (2020). https://doi.org/10.14188/j.1671-8836.2019.0002
[10] Long, J., Shelhamer, E., Darrell, T.: Fully Convolutional Networks for Semantic Segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 3431-3440 (2015)
[11] Ronneberger, O., Fischer, P., Brox, T.: U-Net: Convolutional Networks for Biomedical Image Segmentation. In Medical Image Computing and Computer-Assisted Intervention - MICCAI 2015, pp. 234-241 (2015). https://doi.org/10.1007/978-3-319-24574-4_28
[12] Gu, X., Liu, X., R, Z.: Road Scene Semantic Segmentation Algorithm Based on Multi Feature Fusion. Science Technology and Engineering 21(33), 14251-14257 (2021). https://doi.org/10.3969/j.issn.1671-1815.2021.33.030
[13] Hua, W., Gu, M., Li, L., Cui, L.: Clothing image segmentation algorithm based on improved SOLOv2. Basic Sciences Journal of Textile Universities 34(4), 74-81 (2021). https://doi.org/10.13338/j.issn.1006-8341.2021.04.011
[14] Badrinarayanan, V., Kendall, A., Cipolla, R.: SegNet: A Deep Convolutional Encoder-Decoder Architecture for Image Segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence 39(12), 2481-2495 (2017). https://doi.org/10.1109/TPAMI.2016.2644615
[15] Lin, G., Milan, A., Shen, C., Reid, I.: RefineNet: Multi-path Refinement Networks for High-Resolution Semantic Segmentation. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 5168-5177 (2017). https://doi.org/10.1109/CVPR.2017.549
[16] Chen, L.-C., Zhu, Y., Papandreou, G., Schroff, F., Adam, H.: Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation. In Computer Vision - ECCV 2018, pp. 833-851 (2018). https://doi.org/10.1007/978-3-030-01234-2_49
[17] Martinsson, J., Mogren, O.: Semantic Segmentation of Fashion Images Using Feature Pyramid Networks. In 2019 IEEE/CVF International Conference on Computer Vision Workshop (ICCVW), pp. 3133-3136 (2019). https://doi.org/10.1109/ICCVW.2019.00382
[18] Qin, X., Zhang, Z., Huang, C., Dehghan, M., Zaiane, O.R., Jagersand, M.: U2-Net: Going deeper with nested U-structure for salient object detection. Pattern Recognition 106, 107404 (2020). https://doi.org/10.1016/j.patcog.2020.107404
[19] Guo, M.-H., Xu, T.-X., Liu, J.-J., Liu, Z.-N., Jiang, P.-T., Mu, T.-J., Zhang, S.-H., Martin, R.R., Cheng, M.-M., Hu, S.-M.: Attention mechanisms in computer vision: A survey. Computational Visual Media 8, 331-368 (2022). https://doi.org/10.1007/s41095-022-0271-y
[20] Fu, J., Liu, J., Tian, H., Li, Y., Bao, Y., Fang, Z., Lu, H.: Dual Attention Network for Scene Segmentation. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 3141-3149 (2019). https://doi.org/10.1109/CVPR.2019.00326
[21] Hu, J., Shen, L., Sun, G.: Squeeze-and-Excitation Networks. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 7132-7141 (2018). https://doi.org/10.1109/CVPR.2018.00745
[22] Woo, S., Park, J., Lee, J.-Y., Kweon, I.S.: CBAM: Convolutional Block Attention Module. In Computer Vision - ECCV 2018, pp. 3-19 (2018). https://doi.org/10.1007/978-3-030-01234-2_1
[23] Zhong, H., Zhang, Z., Peng, T., He, R., Hu, X., Zhang, J.: FMNet: feature alignment based multi-directional attention mechanism clothing image segmentation network. Zhongguo keji lunwen 18(03), 275-282 (2023)
[24] Chen, R., Chen, Y., Yu, K., Xu, Z.: A segmentation method for virtual clothing effect images. Industria Textila 76(2), 230-236 (2025). https://doi.org/10.35530/IT.076.02.2024111
[25] Guo, T., Mousavi, H.S., Vu, T.H., Monga, V.: Deep Wavelet Prediction for Image Super-Resolution. In 2017 IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pp. 1100-1109 (2017). https://doi.org/10.1109/CVPRW.2017.148
[26] Bae, W., Yoo, J., Ye, J.C.: Beyond Deep Residual Learning for Image Restoration: Persistent Homology-Guided Manifold Simplification. In 2017 IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pp. 1141-1149 (2017). https://doi.org/10.1109/CVPRW.2017.152
[27] Li, Q., Shen, L., Guo, S., Lai, Z.: Wavelet Integrated CNNs for Noise-Robust Image Classification. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 7243-7252 (2020). https://doi.org/10.1109/CVPR42600.2020.00727
[28] Bugár, G., Bánoci, V., Broda, M., Levický, D., Mikó, E.: Blind steganography based on 2D Haar transform. In Proceedings ELMAR-2013, pp. 31-35 (2013)
[29] Liu, Y., Wu, Y., Zhang, Q., Yan, F., Chen, S.: Road Crack Detection Based on Separable Convolution and Wave Transform Fusion. Computer Science 51(11A), 240100141 (2024). https://doi.org/10.11896/sjkk.240100141
[30] Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., Rabinovich, A.: Going deeper with convolutions. In 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 1-9 (2015). https://doi.org/10.1109/CVPR.2015.7298594
[31] Ge, Y., Zhang, R., Wang, X., Tang, X., Luo, P.: DeepFashion2: A Versatile Benchmark for Detection, Pose Estimation, Segmentation and Re-Identification of Clothing Images. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 5332-5340 (2019). https://doi.org/10.1109/CVPR.2019.00548
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Authors

This work is licensed under a Creative Commons Attribution 4.0 International License.
