A VGG19-based Image Style Transfer Method That Integrates Improved UNet Semantic Segmentation With a Dual-Encoder Architecture
DOI:
https://doi.org/10.64509/jdi.13.102Keywords:
Image Style Transfer, Dual-Encoder Architecture, Improved UNet Semantic Mask, Attention Module, Hybrid-Domain Loss, Wavelet TransformAbstract
To address the common issues of image artifacts, background contamination, and loss of fine texture details in clothing design image translation using the existing VGG19 software, this paper proposes an improved VGG19 network method that incorporates a dual-encoder architecture, a hybrid-domain loss function, and a high-precision semantic mask generated by an improved UNet. In this method, power mean pooling and the SimAM attention mechanism are introduced to the VGG19 network, and two independent encoders are employed to extract and couple the structural features of the clothing and the style features of the image, respectively. The improvements include: 1) using a hybrid-domain loss function with wavelet-domain adjustments to constrain clothing structure and image style; 2) applying a high-precision spatial-domain mask—generated by an improved UNet semantic segmentation network—to completely purify the image background; and 3) employing a total variation loss function to achieve smooth transitions between image contents. Experimental results demonstrate that, compared to the traditional VGG19 network method, the proposed method improves the PSNR and SSIM metrics by 19.5% and 30.9%, respectively, while achieving a remarkable 36.8% reduction in the LPIPS metric, thereby facilitating the application of the VGG19 network method in clothing image style transfer.
Downloads
References
[1] Yu, F., Hong, Y., Liu, C.: Development of a hat style recognition system based on image processing and machine learning. Industria Textila 73(2), 204-212 (2022). https://doi.org/10.35530/IT.073.02.202050
[2] Chen, R., Chen, Y., Yu, K., Xu, Z.: A segmentation method for virtual clothing effect images. Industria Textila 76(2), 230-236 (2025). https://doi.org/10.35530/IT.076.02.2024111
[3] Wang, Z., Zhao, L., Chen, H., Zuo, Z., Li, A., Xing, W., Lu, D.: Evaluate and improve the quality of neural style transfer. Computer Vision and Image Understanding 207, 103203 (2021). https://doi.org/10.1016/j.cviu.2021.103203
[4] Li, M., Yang, C., Shi, B.: Research on Image Style Transfer Technology Based on Semantic Segmentation. Journal of Computer Engineering & Applications 56(24), 207-213 (2020). https://doi.org/10.3778/j.issn.1002-8331.1910-0238
[5] Li, C., Wand, M.: Combining markov random fields and convolutional neural networks for image synthesis. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2479-2486 (2016). https://doi.org/10.1109/CVPR.2016.272
[6] Shen, J., Shi, J., Gu, J., Qian, Q., Shu, X., Pang, L., Zhang, Z.: SADST: Style-aware dynamic style transfer for domain generalized semantic segmentation. Neural Networks 196, 108336 (2025). https://doi.org/10.1016/j.neunet.2025.108336
[7] Radenović, F., Tolias, G., Chum, O.: Fine-tuning CNN image retrieval with no human annotation. IEEE transactions on pattern analysis and machine intelligence 41(7), 1655-1668 (2018). https://doi.org/10.1109/TPAMI.2018.2846566
[8] Yang, L., Zhang, R.Y., Li, L., Xie, X.: SimAM: A Simple, Parameter-Free Attention Module for Convolutional Neural Networks. In Proceedings of the 38th International Conference on Machine Learning, pp. 11863-11874 (2021).
[9] Hu, J., Shen, L., Sun, G.: Squeeze-and-excitation networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 7132-7141 (2018). https://doi.org/10.1109/CVPR.2018.00745
[10] Woo, S., Park, J., Lee, J. Y., Kweon, I.S.: CBAM: Convolutional Block Attention Module. In Proceedings of the European conference on computer vision (ECCV), pp. 3-19 (2018). https://doi.org/10.1007/978-3-030-01234-2_1
[11] Zhao, L., Zhang, Z.: An improved pooling method for convolutional neural networks. Scientific Reports 14(1), 1589 (2024). https://doi.org/10.1038/s41598-024-51258-6
[12] Boussaad, L.: Binary pooling: A novel approach for local feature extraction in convolutional neural networks. Neurocomputing 654, 131201 (2025). https://doi.org/10.1016/j.neucom.2025.131201
[13] AL-Mekhlafi, H., Liu, S.: Structure-Aware Style Transfer Based on Multi-Scale and Edge Texture. In International Conference on Advanced Data Mining and Applications, pp.170-184 (2024). https://doi.org/10.1007/978-981-96-0847-8_12
[14] Shadoul, I., Al-Hmouz, R., Hossen, A., Mesbah, M., Deveci, M.: The effect of pooling parameters on the performance of convolution neural network. Artificial Intelligence Review 58(9), 271 (2025). https://doi.org/10.1007/s10462-025-11273-z
[15] Chen, D.Y., Tennent, H., Hsu, C.W.: Artadapter: Text-to-image style transfer using multi-level style encoder and explicit adaptation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 8619-8628 (2024). https://doi.org/10.1109/CVPR52733.2024.00823
[16] Chen, L., Fu, Y., Gu, L., Yan, C., Harada, T., Huang, G.: Frequency-aware feature fusion for dense image prediction. IEEE transactions on pattern analysis and machine intelligence, 46(12), 10763-10780 (2024). https://doi.org/10.1109/TPAMI.2024.3449959
[17] Huang, X., Belongie, S.: Arbitrary style transfer in real-time with adaptive instance normalization. In Proceedings of the IEEE international conference on computer vision, pp. 1510-1519 (2017). https://doi.org/10.1109/ICCV.2017.167
[18] Wang, H., Xing, P., Huang, R., Ai, H., Wang, Q., Bai, X.: Instantstyle-plus: Style transfer with content-preserving in text-to-image generation. arXiv preprint arXiv:2407.00788 (2024). https://doi.org/10.48550/arXiv.2407.00788
[19] Mishra, D., Hadar, O.: Accelerating neural style-transfer using contrastive learning for unsupervised satellite image super-resolution. IEEE Transactions on Geoscience and Remote Sensing 61, 1-14 (2023). https://doi.org/10.1109/TGRS.2023.3314283
[20] Gatys, L.A., Ecker, A.S., Bethge, M.: Image style transfer using convolutional neural networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2414-2423 (2016). https://doi.org/10.1109/CVPR.2016.265
[21] Esan, D.O., Owolawi, P.A., Tu, C.: Artistic Image Generation Using Deep Convolutional Generative Adversarial Networks. In International Conference on Computer and Communication Engineering, pp. 3-16 (2024). https://doi.org/10.1007/978-3-031-71079-7_1
[22] Simonyan, K., Zisserman, A.: Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 (2014). https://doi.org/10.48550/arXiv.1409.1556
[23] Wang, Z., Bovik, A.C.: Mean squared error: Love it or leave it? A new look at signal fidelity measures. IEEE signal processing magazine 26(1), 98-117 (2009). https://doi.org/10.1109/MSP.2008.930649
[24] Wang, Z., Bovik, A.C., Sheikh, H.R., Simoncelli, E.P.: Image quality assessment: from error visibility to structural similarity. IEEE transactions on image processing 13(4), 600-612 (2004). https://doi.org/10.1109/TIP.2003.819861
[25] Zhang, R., Isola, P., Efros, A.A., Shechtman, E., Wang, O.: The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 586-595 (2018). https://doi.org/10.1109/CVPR.2018.00068
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Authors

This work is licensed under a Creative Commons Attribution 4.0 International License.
