Перейти до основного вмісту
Hyperspherical Mixed Prototypes Networks
Lukashov Dmytro 1
1 Kharkiv National University of Radio Electronics, Kharkiv, Kharkiv, 61166, Ukraine
Keywords: Artificial intelligence, Neural networks, Deep learning, Computer vision
Abstract

This paper introduces hyperspherical mixed prototype networks. The key difference compared to hyperspherical prototype networks is mixing class prototypes and refined optimization objectives. This work proposes mixing data samples and the corresponding class prototypes in their respective spaces. In this case, the objective is to maximize the cosine similarity between a mixed sample and its corresponding mixed prototype. The other proposed objective is to minimize the cross-entropy between the dot product of a mixed sample with the original prototypes and the corresponding vector representing the proportion of each prototype in the mixture. The experiments show a performance improvement in the visual classification task compared to the baseline hyperspherical prototype networks for both optimization objectives. In addition, the results presented in this work outperform those given in the original hyperspherical prototype networks paper. Another key finding is that, when mixing is used, maximizing similarities as an optimization objective results in better accuracy and much better intraclass clustering of embeddings around their respective prototypes than cross-entropy.

References

[1] Sangdoo Yun et al. “CutMix: Regularization Strategy to Train Strong Classifiers With Localizable Features”. В: 2019 IEEE/CVF International Conference on Computer Vision (ICCV) (2019), с. 6022—6031. url: https://api.semanticscholar.org/CorpusID:152282661.
[2] Hongyi Zhang et al. “mixup: Beyond Empirical Risk Minimization”. В: International Conference on Learning Representations. 2018. url: https://openreview.net/forum?id=r1Ddp1-Rb.
[3] Jin-Ha Lee et al. “SmoothMix: a Simple Yet Effective Data Augmentation to Train Robust Classifiers”. В: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). 2020, с. 3264—3274. doi: 10.1109/CVPRW50498.2020.00386.
[4] Runji Liu et al. “Attentive Mix: An Efficient Data Augmentation Method for Object Detection”. В: 2021 7th International Conference on Computer and Communications (ICCC). 2021, с. 770—774. doi: 10.1109/ICCC54389.2021.9674718.
[5] Jang-Hyun Kim, Wonho Choo and Hyun Oh Song. “Puzzle Mix: Exploiting Saliency and Local Statistics for Optimal Mixup”. В: Proceedings of the 37th International Conference on Machine Learning. Ed. by Hal Daum´e III and Aarti Singh. Vol. 119. Proceedings of Machine Learning Research. PMLR, 13–18 Jul 2020, с. 5275—5285. url: https://proceedings.mlr.press/v119/kim20b.html.
[6] A. F. M. Shahab Uddin et al. “SaliencyMix: A Saliency Guided Data Augmentation Strategy for Better Regularization”. В: International Conference on Learning Representations. 2021. url: https://openreview.net/forum?id=-M0QkvBGTTq.
[7] Shaoli Huang, Xinchao Wang and Dacheng Tao. “SnapMix: Semantically Proportional Mixing for Augmenting Fine-grained Data”. В: AAAI Conference on Artificial Intelligence. 2020. url: https://api.semanticscholar.org/CorpusID:228063878.
[8] JangHyun Kim et al. “Co-Mixup: Saliency Guided Joint Mixup with Supermodular Diversity”. В: International Conference on Learning Representations. 2021. url: https://openreview.net/forum?id=gvxJzw8kW4b.
[9] Pascal Mettes, Elise van der Pol and Cees Snoek. “Hyperspherical Prototype Networks”. В: Advances in Neural Information Processing Systems. Ed. by H. Wallach et al. Vol. 32. Curran Associates, Inc., 2019. url: https://proceedings.neurips.cc/paper_files/paper/2019/file/02a32ad2669e6fe298e607fe7cc0e1a0-Paper.pdf.
[10] Tejaswi Kasarla et al. “Maximum class separation as inductive bias in one matrix”. В: Proceedings of the 36th International Conference on Neural Information Processing Systems. NIPS ’22. New Orleans, LA, USA: Curran Associates Inc., 2024. isbn: 9781713871088.
[11] Mina Ghadimi Atigh, Martin Keller-Ressel and Pascal Mettes. “Hyperbolic Busemann Learning with Ideal Prototypes”. В: Advances in Neural Information Processing Systems. Ed. by A. Beygelzimer et al. 2021. url: https://openreview.net/forum?id=c_XcmuxwAY.
[12] Federico Pernici et al. “Regular Polytope Networks”. В: IEEE Transactions on Neural Networks and Learning Systems 33.9 (2022), с. 4373—4387. doi: 10.1109/TNNLS.2021.3056762.
[13] Federico Pernici et al. “Maximally Compact and Separated Features with Regular Polytope Networks”. В: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops. Черв. 2019.
[14] Jiankang Deng et al. “ArcFace: Additive Angular Margin Loss for Deep Face Recognition”. В: 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2019, с. 4685—4694. doi: 10.1109/CVPR.2019.00482.
[15] Y. Huang et al. “CurricularFace: Adaptive Curriculum Learning Loss for Deep Face Recognition”. В: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2020), с. 5900—5909. url: https://api . semanticscholar . org /CorpusID:209050760.
[16] Xinlong Yang et al. “Prototypical Mixing and Retrieval-based Refinement for Label Noise-resistant Image Retrieval”. В: 2023 IEEE/CVF International Conference on Computer Vision (ICCV). 2023, с. 11205—11215. doi: 10.1109/ICCV51070.2023.01032.
[17] Jake Snell, Kevin Swersky and Richard S. Zemel. “Prototypical Networks for Fewshot Learning”. В: Neural Information Processing Systems. 2017. url: https://api.semanticscholar.org/CorpusID:309759.
[18] Hong-Ming Yang et al. “Convolutional Prototype Network for Open Set Recognition”. В: IEEE Transactions on Pattern Analysis and Machine Intelligence 44.5 (2022), с. 2358— 2370. doi: 10.1109/TPAMI.2020.3045079.
[19] Mengye Ren et al. “Meta-Learning for Semi-Supervised Few-Shot Classification”. В: International Conference on Learning Representations. 2018. url: https://openreview.net/forum?id=HJcSzz-CZ.
[20] Zhixiong Zeng et al. “PAN: Prototype-based Adaptive Network for Robust Crossmodal Retrieval”. В: Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval. SIGIR ’21. Virtual Event, Canada: Association for Computing Machinery, 2021, с. 1125—1134. isbn: 9781450380379. doi: 10.1145/3404835.3462867. url: https://doi.org/10.1145/3404835.3462867.
[21] J.J. Thomson. “XXIV. On the structure of the atom: an investigation of the stability and periods of oscillation of a number of corpuscles arranged at equal intervals around the circumference of a circle; with application of the results to the theory of atomic structure”. В: The London, Edinburgh, and Dublin Philosophical Magazine and Journal of Science 7.39 (1904), с. 237—265. doi: 10.1080/14786440409463107. url: https://doi.org/10.1080/14786440409463107.
[22] A. Krizhevsky and G. Hinton. “Learning multiple layers of features from tiny images”. В: Master’s thesis, Department of Computer Science, University of Toronto (2009). [23] Ya Le and Xuan S. Yang. “Tiny ImageNet Visual Recognition Challenge”. В: (2015). url: https://api.semanticscholar.org/CorpusID:16664790.
[24] Diederik P. Kingma and Jimmy Ba. “Adam: A Method for Stochastic Optimization”. В: CoRR abs/1412.6980 (2014). url: https://api.semanticscholar.org/CorpusID:6628106.
[25] Ekin Dogus Cubuk et al. “RandAugment: Practical Automated Data Augmentation with a Reduced Search Space”. В: Advances in Neural Information Processing Systems. Ed. by H. Larochelle et al. Vol. 33. Curran Associates, Inc., 2020, с. 18613—18624. url: https://proceedings.neurips.cc/paper_files/paper/2020/file/
d85b63ef0ccb114d0a3bb7b7d808028f-Paper.pdf.
[26] Kaiming He et al. “Deep Residual Learning for Image Recognition”. В: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Черв. 2016.

Paper Received 1/10/2026
Paper Accepted 2/26/2026
Published Online 2/26/2026
Cite
ACS Style
Lukashov , D. Hyperspherical Mixed Prototypes Networks. Bukovinian Mathematical Journal. 2026, 14 https://doi.org/https://doi.org/10.31861/bmj2026.01.06
AMA Style
Lukashov D. Hyperspherical Mixed Prototypes Networks. Bukovinian Mathematical Journal. 2026; 14(1). https://doi.org/https://doi.org/10.31861/bmj2026.01.06
Chicago/Turabian Style
Dmytro Lukashov . 2026. "Hyperspherical Mixed Prototypes Networks". Bukovinian Mathematical Journal. 14 no. 1. https://doi.org/https://doi.org/10.31861/bmj2026.01.06
Export

The journal is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International

We use own, third-party cookies, and localStorage files to analyze web traffic and page activities. Privacy Policy Settings