Fashion Style Quantification and Clustering Network Based on Unsupervised Latent Space Representation Learning

Authors

  • Zhiyuan Yu College of Textile and Clothing Engineering, Soochow University, Suzhou 215021, China Author
  • Ziyi Guo College of Textile and Clothing Engineering, Soochow University, Suzhou 215021, China Author
  • Yan Hong College of Textile and Clothing Engineering, Soochow University, Suzhou 215021, China Author

DOI:

https://doi.org/10.64509/jdi.13.119

Keywords:

Latent Space Representation, Fashion Aesthetics, Deep Clustering, Data Visualization, Multimodal Fusion

Abstract

 Amidst the rapid evolution of fashion e-commerce, efficiently matching massive inventories with personalized demands requires the objective quantification of aesthetic styles. However, traditional manual annotations are subjective and costly, while low-level visual features struggle with the "semantic gap" against high-level user preferences. To address these challenges, this study proposes LUCID (Latent Unsupervised Clustering with Interactive Display), an unsupervised latent space representation learning and clustering visualization network driven by vision-language priors. LUCID first utilizes a pre-trained Contrastive Language-Image Pre-training (CLIP) image encoder to extract high-dimensional semantic visual features. To overcome the "curse of dimensionality," it employs an auto-encoder integrated with a contrastive learning mechanism to compress these features into denoised, low-dimensional latent vectors. Subsequently, the K-means algorithm automatically clusters these representations, and Uniform Manifold Approximation and Projection (UMAP) generates an intuitive two-dimensional fashion style distribution map. This end-to-end workflow completely circumvents the reliance on manual labels, achieving effective feature dimensionality reduction and semantic refinement. Extensive experiments demonstrate that LUCID's unsupervised clustering aligns highly with human consensus, achieving ACC, NMI, and ARI scores of 84%, 75%, and 69% on the DeepFashion dataset, with strong generalization on Fashion-MNIST (68%, 69%, 56%). Ultimately, this research effectively bridges the aesthetic semantic gap, providing robust technical support for intelligent fashion analysis systems.

Downloads

Download data is not yet available.

References

[1] Hong, Y., Guo, S., Zeng, X., Zhang, J.: Human cognition modeling for the metaverse-oriented design system. IEEE Network 38(6), 243-251 (2004). https://doi.org/10.1109/MNET.2024.3377909

[2] Li, M., Zhang, J., Hong, Y., Xie, X., Zhang, M., Guo, S.: Advanced product personalization in blockchain-enabled metaverse: A diffusion model for automatic style generation. IEEE Internet Things Journal 12(8), 10304-10315 (2025). https://doi.org/10.1109/JIOT.2024.3511667

[3] Chakraborty, S., Hoque, M.S., Jeem, N.R., Biswas, M.C., Bardhan, D., Lobaton, E.: Fashion Recommendation Systems, Models and Methods: A Review. Informatics 8(3), 49 (2021). https://doi.org/10.3390/informatics8030049

[4] Hu, Z.-H., Li, X., Wei, C., Zhou, H.-L.: Examining Collaborative Filtering Algorithms for Clothing Recommendation in E-Commerce. Textile Research Journal 89(14), 2821-2835 (2019). https://doi.org/10.1177/0040517518801200

[5] McAuley, J., Targett, C., Shi, Q., Hengel, A.: Image-Based Recommendations on Styles and Substitutes. In Proceedings of the 38th International ACM SIGIR Conference on Research and Development in Information Retrieval, pp. 43-52 (2015). https://doi.org/10.1145/2766462.2767755

[6] Jeon, Y., Jin, S., Han, K.: FANCY: Human-Centered, Deep Learning-Based Framework for Fashion Style Analysis. In Proceedings of the Web Conference 2021, pp. 2367-2378 (2021). https://doi.org/10.1145/3442381.3449833

[7] Veit, A., Kovacs, B., Bell, S., McAuley, J., Bala, K., Belongie, S.: Learning Visual Clothing Style with Heterogeneous Dyadic Co-Occurrences. In 2015 IEEE International Conference on Computer Vision (ICCV), pp. 4642-4650 (2015). https://doi.org/10.1109/ICCV.2015.527

[8] Xie, J., Girshick, R., Farhadi, A.: Unsupervised Deep Embedding for Clustering Analysis. In Proceedings of the 33rd International Conference on International Conference on Machine Learning, pp. 478-487 (2016)

[9] Guo, X., Gao, L., Liu, X., Yin, J.: Improved Deep Embedded Clustering with Local Structure Preservation. In Proceedings of the International Joint Conference on Artificial Intelligence, pp. 1753-1759 (2017)

[10] Jiao, Y., Xie, N., Gao, Y., Wang, C.-C., Sun, Y.: Fine-Grained Fashion Representation Learning by Online Deep Clustering. In European Conference on Computer Vision, pp. 19-35 (2022). https://doi.org/10.1007/978-3-031-19812-0_2

[11] Yin, S., Fu, C., Zhao, S., Li, K., Sun, X., Xu, T., Chen, E.: A Survey on Multimodal Large Language Models. National Science Review 11(12), 403 (2024). https://doi.org/10.1093/nsr/nwae403

[12] Zhao, F., Zhang, C., Geng, B.: Deep Multimodal Data Fusion. ACM Computing Surveys 56(9), 1-36 (2024). https://doi.org/10.1145/3649447

[13] Ma, J., Sun, H., Yang, D., Zhang, H.: Personalized Fashion Recommendations for Diverse Body Shapes with Contrastive Multimodal Cross-Attention Network. ACM Transactions on Intelligent Systems and Technology 15(4), 1-21 (2024). https://doi.org/10.1145/3637217

[14] He, R., McAuley, J.: VBPR: Visual Bayesian Personalized Ranking from Implicit Feedback. In Proceedings of the Thirtieth AAAI Conference on Artificial Intelligence, pp. 144-150 (2015)

[15] Maaten, L., Hinton, G.: Visualizing Data Using t-SNE. Journal of Machine Learning Research 9(86), 2579-2605 (2008)

[16] Zhao, H., Chen, H., Yang, F., Liu, N., Deng, H., Cai, H., Wang, S., Yin, D., Du, M.: Explainability for Large Language Models: A Survey. ACM Transactions on Intelligent Systems and Technology 15(2), 1-38 (2024). https://doi.org/10.1145/3639372

[17] Krizhevsky, A., Sutskever, I., Hinton, G.E.: ImageNet Classification with Deep Convolutional Neural Networks. Communications of the ACM 60, 84-90 (2017). https://doi.org/10.1145/3065386

[18] Hsiao, W.-L., Grauman, K.: Creating Capsule Wardrobes from Fashion Images. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 7161-7170 (2018). https://doi.org/10.1109/CVPR.2018.00748

[19] Liu, Z., Luo, P., Qiu, S., Wang, X., Tang, X.: DeepFashion: Powering Robust Clothes Recognition and Retrieval with Rich Annotations. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 1096-1104 (2016). https://doi.org/10.1109/CVPR.2016.124

[20] Xiao, H., Rasul, K., Vollgraf, R.: Fashion-MNIST: A Novel Image Dataset for Benchmarking Machine Learning Algorithms. arXiv preprint arXiv:1708.07747 (2017). https://doi.org/10.48550/arXiv.1708.07747

[21] Hinton, G.E., Salakhutdinov, R.R.: Reducing the Dimensionality of Data with Neural Networks. Science 313(5786), 504-507 (2006). https://doi.org/10.1126/science.1127647

[22] Krizhevsky, A., Hinton, G.E.: Using Very Deep Autoencoders for Content-Based Image Retrieval. In Proceedings of the European Symposium on Artificial Neural Networks (2011)

[23] Olshausen, B.A., Field, D.J.: Sparse Coding with an Overcomplete Basis Set: A Strategy Employed by V1? Vision Research 37(23), 3311-3325 (1997). https://doi.org/10.1016/S0042-6989(97)00169-7

[24] Bell, A.J., Sejnowski, T.J.: The Independent Components of Natural Scenes Are Edge Filters. Vision Research 37(23), 3327-3338 (1997). https://doi.org/10.1016/S0042-6989(97)00121-1

[25] Vincent, P., Larochelle, H., Bengio, Y., Manzagol, P.-A.: Extracting and Composing Robust Features with Denoising Autoencoders. In Proceedings of the 25th International Conference on Machine Learning, pp. 1096-1103 (2008). https://doi.org/10.1145/1390156.1390294

[26] Kingma, D.P., Welling, M.: Auto-Encoding Variational Bayes. In International Conference on Learning Representations (2014)

[27] Makhzani, A., Shlens, J., Jaitly, N., Goodfellow, I., Frey, B.: Adversarial Autoencoders. arXiv preprint arXiv:1511.05644 (2015). https://doi.org/10.48550/arXiv.1511.05644

[28] Yang, B., Fu, X., Sidiropoulos, N.D., Hong, M.: Towards K-Means-Friendly Spaces: Simultaneous Deep Learning and Clustering. In Proceedings of the 34th International Conference on Machine Learning, pp. 3861-3870 (2017)

[29] Yang, J., Parikh, D., Batra, D.: Joint Unsupervised Learning of Deep Representations and Image Clusters. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 5147-5156 (2016). https://doi.org/10.1109/CVPR.2016.556

[30] Jiang, Z., Zheng, Y., Tan, H., Tang, B., Zhou, H.: VaDE: Variational Deep Embedding: An Unsupervised and Generative Approach to Clustering. In Proceedings of the Twenty-Sixth International Joint Conference on Artificial Intelligence, pp. 1965-1972 (2017). https://doi.org/10.24963/ijcai.2017/273

[31] MacQueen, J.: Some Methods for Classification and Analysis of Multivariate Observations. In Proceedings of the Fifth Berkeley Symposium on Mathematical Statistics and Probability, vol. 1, pp. 281-297 (1967)

[32] Aggarwal, C.C., Hinneburg, A., Keim, D.A.: On the Surprising Behavior of Distance Metrics in High Dimensional Space. In Proceedings of the 8th International Conference on Database Theory, pp. 420-434 (2001)

[33] Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.: Learning Transferable Visual Models From Natural Language Supervision. In Proceedings of the 38th International Conference on Machine Learning, pp. 8748-8763 (2021)

[34] Han, Y., Zhang, L., Chen, Q., Chen, Z., Li, Z., Yang, J., Cao, Z.: FashionSAP: Symbols and Attributes Prompt for Fine-Grained Fashion Vision-Language Pre-Training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 15028-15038 (2023). https://doi.org/10.1109/CVPR52729.2023.01443

[35] Hadsell, R., Chopra, S., LeCun, Y.: Dimensionality Reduction by Learning an Invariant Mapping. In Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, pp. 1735-1742 (2006). https://doi.org/10.1109/CVPR.2006.100

[36] Oord, A., Li, Y., Vinyals, O.: Representation Learning with Contrastive Predictive Coding. arXiv preprint arXiv:1807.03748 (2018). https://doi.org/10.48550/arXiv.1807.03748

[37] Chen, T., Kornblith, S., Norouzi, M., Hinton, G.: A Simple Framework for Contrastive Learning of Visual Representations. In Proceedings of the 37th International Conference on Machine Learning, vol. 119, pp. 1597-1607 (2020)

[38] McInnes, L., Healy, J., Saul, N., Grossberger, L.: UMAP: Uniform Manifold Approximation and Projection. Journal of Open Source Software 3(29), 861 (2018). https://doi.org/10.21105/joss.00861

[39] Cai, D., He, X., Han, J.: Locally Consistent Concept Factorization for Document Clustering. IEEE Transactions on Knowledge and Data Engineering 23(6), 902-913 (2010). https://doi.org/10.1109/TKDE.2010.165

[40] Santos, J.M., Embrechts, M.: On the Use of the Adjusted Rand Index as a Metric for Evaluating Supervised Classification. In International Conference on Artificial Neural Networks, pp. 175-184 (2009). https://doi.org/10.1007/978-3-642-04277-5_18

JDI 119

Downloads

Published

2026-07-10

Issue

Section

Articles

How to Cite

Yu, Z., Guo, Z., & Hong, Y. (2026). Fashion Style Quantification and Clustering Network Based on Unsupervised Latent Space Representation Learning. Journal of Design Intelligence , 1(3), 39-50. https://doi.org/10.64509/jdi.13.119