Garment Image Segmentation Method Based on Wavelet Multi-Scale Features and Improved U-Net

Authors

  • Jiangtao Wei School of Textiles and Fashion, Shanghai University of Engineering Science, Shanghai 201620, China Author
  • Yanmei Li School of Textiles and Fashion, Shanghai University of Engineering Science, Shanghai 201620, China Author
  • Yu Chen School of Textiles and Fashion, Shanghai University of Engineering Science, Shanghai 201620, China Author

DOI:

https://doi.org/10.64509/jdi.13.103

Keywords:

Clothing Image Segmentation, U-Net, Deeply Separable Convolution, Attention Mechanism, Haar Wavelet Down Sampling, Residual Network

Abstract

Against the background of the gradual integration of fashion design and AIGC, obtaining accurate clothing masks through image segmentation and further guiding diffusion models for pattern generation and reconstruction serves as a key technique for intelligent design. Nevertheless, practical scenarios are plagued by issues such as complex background interference and occlusion between human bodies and garments, which easily lead to blurred segmentation boundaries and loss of details, thus impairing the subsequent design generation performance. To address these problems, this paper proposes a clothing image segmentation method based on an improved U-Net architecture. By fusing wavelet sampling and partial convolution in the downsampling stage, the feature learning of the encoder is optimized and the influence of occluded regions is reduced. Furthermore, the Efficient Local Attention-T (ELA-T) mechanism is embedded into the skip connections to enhance the model's sensitivity to subtle boundary features. A novel residual multi-scale depthwise separable convolution module (R-MSDW) is designed to strengthen the model's capability in capturing multi-scale features and hierarchical semantic information. Finally, experiments are conducted on the DeepFashion2 dataset and a self-constructed dataset. The proposed method outperforms mainstream models such as U-Net, SegNet, and U2-Net, achieving an mIoU of 86.1% and an overall accuracy of 95.1% on the DeepFashion2 dataset, which surpasses vanilla U-Net by 8.78% in mIoU and 2.7% in overall accuracy respectively, and an mIoU of 85.4% and an overall accuracy of 94.8% on the self-constructed dataset. The proposed model helps moderately enhance the visual quality of generated fashion design images.

Downloads

Download data is not yet available.

Author Biographies

  • Yanmei Li, School of Textiles and Fashion, Shanghai University of Engineering Science, Shanghai 201620, China

    I obtained a master's degree from Jiangnan University in 2001, a doctoral degree from Donghua University in 2009, a visiting scholar at the School of Computer Science, Fudan University in 2011, and a senior visiting scholar at Cornell University in the United States in 2015. Received the title of Outstanding Young Teacher in Shanghai in 2009; Received the title of Red Flag Bearer twice in 2014 and 2016.Senior Professor at the School of Textile and Apparel, Shanghai University of Engineering Science

  • Yu Chen, School of Textiles and Fashion, Shanghai University of Engineering Science, Shanghai 201620, China

    I graduated from the Department of Electronic Engineering at Shanghai Jiao Tong University with a bachelor's degree in 1999. In 2006, I obtained a doctoral degree in Computer and Automation from the University of Lille in France. In 2007, I went to Hong Kong through the Hong Kong Talents Program to work as the Technical Director of 3D Human Body Database at TPC Company. In 2017, I joined the University of Engineering as a teacher through the Talent Program. Senior Professor at the School of Textile and Apparel, Shanghai University of Engineering Science.

References

[1] Chen, B., Zhang, Y., Yu, B., Liu, X.: Two-stage adjustable perceptual distillation network for virtual tryon. Journal of Graphics 43(2), 316-323 (2022)

[2] Xu, Y.: Garment image segmentation technology and application based on convolutional neural network. Master's Thesis, Donghua University (2021)

[3] Zhang, Q., Liu, L., Fu, X., Liu, L., Huang, Q.: Clothing Image Retrieval by Label Optimization and Semantic Segmentation. Journal of Computer-Aided Design & Computer Graphics 32(9), 1450-1465 (2020). https://doi.org/10.3724/SP.J.1089.2020.18122

[4] Xu, H., Bai, M., Wan, T., Xue, T., Tang, W.: Image semantic analysis and retrieval recommendation for clothing based on deep learning. Fangzhi Gaoxiao Jichukexue Xuebao 33(3), 64-72 (2020). https://doi.org/10.13338/j.issn.1006-8341.2020.03.011

[5] Zhang, L., Rao, A., Agrawala, M.: Adding Conditional Control to Text-to-Image Diffusion Models. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 3836-3847 (2023)

[6] Sezgin, M., Sankur, B.: Survey over image thresholding techniques and quantitative performance evaluation. Journal of Electronic Imaging 13(1), 146-165 (2004). https://doi.org/10.1117/1.1631315

[7] Adams, R., Bischof, L.: Seeded region growing. IEEE Transactions on Pattern Analysis and Machine Intelligence 16(6), 641-647 (1994). https://doi.org/10.1109/34.295913

[8] Canny, J.: A Computational Approach to Edge Detection. IEEE Transactions on Pattern Analysis and Machine Intelligence PAMI-8(6), 679-698 (1986). https://doi.org/10.1109/TPAMI.1986.4767851

[9] Huang, P., Zheng, Q., Liang, C.: Overview of Image Segmentation Methods. Journal of Wuhan University (Natural Science Edition) 66(6), 519-531 (2020). https://doi.org/10.14188/j.1671-8836.2019.0002

[10] Long, J., Shelhamer, E., Darrell, T.: Fully Convolutional Networks for Semantic Segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 3431-3440 (2015)

[11] Ronneberger, O., Fischer, P., Brox, T.: U-Net: Convolutional Networks for Biomedical Image Segmentation. In Medical Image Computing and Computer-Assisted Intervention - MICCAI 2015, pp. 234-241 (2015). https://doi.org/10.1007/978-3-319-24574-4_28

[12] Gu, X., Liu, X., R, Z.: Road Scene Semantic Segmentation Algorithm Based on Multi Feature Fusion. Science Technology and Engineering 21(33), 14251-14257 (2021). https://doi.org/10.3969/j.issn.1671-1815.2021.33.030

[13] Hua, W., Gu, M., Li, L., Cui, L.: Clothing image segmentation algorithm based on improved SOLOv2. Basic Sciences Journal of Textile Universities 34(4), 74-81 (2021). https://doi.org/10.13338/j.issn.1006-8341.2021.04.011

[14] Badrinarayanan, V., Kendall, A., Cipolla, R.: SegNet: A Deep Convolutional Encoder-Decoder Architecture for Image Segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence 39(12), 2481-2495 (2017). https://doi.org/10.1109/TPAMI.2016.2644615

[15] Lin, G., Milan, A., Shen, C., Reid, I.: RefineNet: Multi-path Refinement Networks for High-Resolution Semantic Segmentation. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 5168-5177 (2017). https://doi.org/10.1109/CVPR.2017.549

[16] Chen, L.-C., Zhu, Y., Papandreou, G., Schroff, F., Adam, H.: Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation. In Computer Vision - ECCV 2018, pp. 833-851 (2018). https://doi.org/10.1007/978-3-030-01234-2_49

[17] Martinsson, J., Mogren, O.: Semantic Segmentation of Fashion Images Using Feature Pyramid Networks. In 2019 IEEE/CVF International Conference on Computer Vision Workshop (ICCVW), pp. 3133-3136 (2019). https://doi.org/10.1109/ICCVW.2019.00382

[18] Qin, X., Zhang, Z., Huang, C., Dehghan, M., Zaiane, O.R., Jagersand, M.: U2-Net: Going deeper with nested U-structure for salient object detection. Pattern Recognition 106, 107404 (2020). https://doi.org/10.1016/j.patcog.2020.107404

[19] Guo, M.-H., Xu, T.-X., Liu, J.-J., Liu, Z.-N., Jiang, P.-T., Mu, T.-J., Zhang, S.-H., Martin, R.R., Cheng, M.-M., Hu, S.-M.: Attention mechanisms in computer vision: A survey. Computational Visual Media 8, 331-368 (2022). https://doi.org/10.1007/s41095-022-0271-y

[20] Fu, J., Liu, J., Tian, H., Li, Y., Bao, Y., Fang, Z., Lu, H.: Dual Attention Network for Scene Segmentation. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 3141-3149 (2019). https://doi.org/10.1109/CVPR.2019.00326

[21] Hu, J., Shen, L., Sun, G.: Squeeze-and-Excitation Networks. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 7132-7141 (2018). https://doi.org/10.1109/CVPR.2018.00745

[22] Woo, S., Park, J., Lee, J.-Y., Kweon, I.S.: CBAM: Convolutional Block Attention Module. In Computer Vision - ECCV 2018, pp. 3-19 (2018). https://doi.org/10.1007/978-3-030-01234-2_1

[23] Zhong, H., Zhang, Z., Peng, T., He, R., Hu, X., Zhang, J.: FMNet: feature alignment based multi-directional attention mechanism clothing image segmentation network. Zhongguo keji lunwen 18(03), 275-282 (2023)

[24] Chen, R., Chen, Y., Yu, K., Xu, Z.: A segmentation method for virtual clothing effect images. Industria Textila 76(2), 230-236 (2025). https://doi.org/10.35530/IT.076.02.2024111

[25] Guo, T., Mousavi, H.S., Vu, T.H., Monga, V.: Deep Wavelet Prediction for Image Super-Resolution. In 2017 IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pp. 1100-1109 (2017). https://doi.org/10.1109/CVPRW.2017.148

[26] Bae, W., Yoo, J., Ye, J.C.: Beyond Deep Residual Learning for Image Restoration: Persistent Homology-Guided Manifold Simplification. In 2017 IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pp. 1141-1149 (2017). https://doi.org/10.1109/CVPRW.2017.152

[27] Li, Q., Shen, L., Guo, S., Lai, Z.: Wavelet Integrated CNNs for Noise-Robust Image Classification. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 7243-7252 (2020). https://doi.org/10.1109/CVPR42600.2020.00727

[28] Bugár, G., Bánoci, V., Broda, M., Levický, D., Mikó, E.: Blind steganography based on 2D Haar transform. In Proceedings ELMAR-2013, pp. 31-35 (2013)

[29] Liu, Y., Wu, Y., Zhang, Q., Yan, F., Chen, S.: Road Crack Detection Based on Separable Convolution and Wave Transform Fusion. Computer Science 51(11A), 240100141 (2024). https://doi.org/10.11896/sjkk.240100141

[30] Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., Rabinovich, A.: Going deeper with convolutions. In 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 1-9 (2015). https://doi.org/10.1109/CVPR.2015.7298594

[31] Ge, Y., Zhang, R., Wang, X., Tang, X., Luo, P.: DeepFashion2: A Versatile Benchmark for Detection, Pose Estimation, Segmentation and Re-Identification of Clothing Images. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 5332-5340 (2019). https://doi.org/10.1109/CVPR.2019.00548

JDI103

Downloads

Published

2026-07-03

Issue

Section

Articles

How to Cite

Wei, J., Li, Y., & Chen, Y. (2026). Garment Image Segmentation Method Based on Wavelet Multi-Scale Features and Improved U-Net. Journal of Design Intelligence , 1(3), 25-38. https://doi.org/10.64509/jdi.13.103

Similar Articles

You may also start an advanced similarity search for this article.