tech_surveillance2041 wordsRead on Arc Codex

Certainty-Aware Partition and Sufficient Utilization of Noisy Samples

Abstract Large-scale annotated image datasets commonly suffer from label noise, which can lead deep neural networks to overfit and degrade in performance. Learning with Noisy Labels (LNL) has gained prominence as a practical strategy. For most classical LNL methods, a core technical route primarily involves conducting a partition to separate samples into clean and noisy subsets. Then, Semi-Supervised Learning (SSL) is employed to fully exploit the information from both clean and noisy samples. However, these approaches still face three major challenges: (i) imperfect separation between clean and noisy samples, (ii) underutilization of uncertain unlabeled data, and (iii) distribution discrepancies between subsets caused by different class-wise annotation difficulties. To address these issues, we propose CAPSUN, a robust framework that improves the precision of clean sample selection and mitigates distribution bias through alignment among subsets. CAPSUN automatically selects the most discriminative perspective from multiple two-dimensional noise-associated feature spaces to achieve reliable clean-noisy separation. Also, it employs an adaptive weighting strategy to preserve the utility of uncertain unlabeled samples, ensuring that each clean sample benefits training even if it is misclassified into the noisy subset. Inspired by domain adaptation in transfer learning, we further design a distribution alignment module to adjust the class distribution contrast of labeled and unlabeled subsets to mitigate class distribution discrepancies. Experimental results across a wide range of synthetic and real-world noisy datasets verify the robustness of CAPSUN, showing consistent improvements over state-of-the-art baselines, particularly under severe noise conditions. Data Availability No datasets were generated or analysed during the current study. References Arazo, E., & Ortego, D (2019). Unsupervised label noise modeling and loss correction. International Conference on Machine Learning (ICML) (pp. 312–321) Bai, Y., Yang, E., Han, B., Yang, Y., Li, J., Mao, Y., & Liu, T. (2021). Understanding and improving early stopping for learning with noisy labels. Advances in Neural Information Processing Systems (NeurIPS) (Vol. 34, pp. 24392–24403). Cao, K., Wei, C., Gaidon, A., Arechiga, N., & Ma, T. (2019). Learning imbalanced datasets with label-distribution-aware margin loss. Advances in Neural Information Processing Systems (Vol. 32) Cheng, L., Tian, J., Zhao, Y., Chang, H., Du, Y., & L. (2025). Cgmatch: A different perspective of semi-supervised learning. Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR) (pp. 15381–15391) Cheng, H., Zhu, Z., Li, X., Gong, Y., Sun, X., Liu, Y. (2021). Learning with instance-dependent label noise: A sample sieve approach. International Conference on Learning Representations (ICLR). Cordeiro, F. R., Belagiannis, V., Reid, I., & Carneiro, G. (2021). Propmix: Hard sample filtering and proportional mixup for learning with noisy labels. arxiv preprint arxiv:2110.11809 Cordeiro, F. R., Sachdeva, R., Belagiannis, V., Reid, I., & Carneiro, G. (2023). Longremix: Robust learning with high confidence samples in a noisy label environment. Pattern Recognition, 133, Article 109013. Cubuk, E. D., Zoph, B., Shlens, J., & Le, Q. V. (2020). Randaugment: Practical automated data augmentation with a reduced search space. IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) (pp. 702–703) Deng, L., Yang, B., Kang, Z., & Xiang, Y. (2024). Invariant feature based label correction for dnn when learning with noisy labels. Neural Networks, 172, Article 106137. Feng, C., Ren, Y., & Xie, X. (2023). Ot-filter: An optimal transport filter for learning with noisy labels. Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR) (pp. 16164–16174) Fooladgar, F., To, M. N. N., Mousavi, P., & Abolmaesumi, P. (2024). Manifold dividemix: A semi-supervised contrastive learning framework for severe label noise. Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR) (pp. 4012–4021) Han, B., Yao, Q., Yu, X., Niu, G., Xu, M., Hu, W.. Sugiyama, M. (2018). Co-teaching: robust training of deep neural networks with extremely noisy labels. Proceedings of the 32nd International Conference on Neural Information Processing Systems (p.8536–8546). Red Hook, NY, USA: Curran Associates Inc. Hodson, T. O. (2022). Root-mean-square error (rmse) or mean absolute error (mae): when to use them or not. Geoscientific Model Development, 15(14), 5481–5487. Huang, Z., Zhang, J., & Shan, H. (2023). Twin contrastive learning with noisy labels. Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR) (pp. 11661–11670) Jiang, Z., Bai, J., Wang, X., Meng, W., & D. (2024). Which is more effective in label noise cleaning, correction or filtering? Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 38, pp. 12866–12873) Jiang, Z., Leung, Z., Li, T., Fei-Fei, L. J., & L. (2018). Mentornet: Learning data-driven curriculum for very deep neural networks on corrupted labels. Proceedings of the 35th International Conference on Machine Learning (Vol. 80, pp. 2304–2313) PMLR. Karim, N., Rizve, M.N., Rahnavard, N., Mian, A., Shah, M. (2022). Unicon: Combating label noise through uniform selection and contrastive learning. Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR) (p.9676-9686). Krizhevsky, A., & Hinton, G. (2009). Learning multiple layers of features from tiny images. Handbook of systemic autoimmune diseases (Vol. 1, Li, Chang, T W., Kuang, K., Li, X., Chen, L., Zhou, J. (2025). Learning causal transition matrix for instance-dependent label noise. Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 39, pp. 18305–18313). Li, Han, H., Shan, S., Chen, X. (2023). Disc: Learning from noisy labels via dynamic instance-specific selection and correction. Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR) (p.24070-24079). Li, Liu, T., Han, B., Niu, G., Sugiyama, M. (2021). Provably end-to-end label-noise learning without anchor points. International conference on machine learning (pp. 6403–6413). Li, S., Hoi, R., & S. C. (2020). DivideMix: Learning with Noisy Labels as Semi-supervised Learning. International Conference on Learning Representations. Li, Xia, X., Ge, S.,& Liu, T. (2022). Selective-supervised contrastive learning with noisy labels. Proceedings of the ieee/cvf conference on computer vision and pattern recognition (pp. 316–325). Xia, X., Zhu, F., Liu, T., Zhang, X.-Y., & Liu, C.-L. (2023). Dynamics-aware loss for learning with label noise. Pattern Recognition, 144, Article 109835. Xiong, C.,& Hoi, S.C (2021). Learning from noisy data with robust representation learning. Proceedings of the ieee/cvf international conference on computer vision (pp. 9485–9494) Lin, Y., Yao, Y., & Liu, T. (2024). Learning the latent causal structure for modeling label noise. Advances in Neural Information Processing Systems (Vol. 37, pp. 120549–120577) Liu, & Guo, H. (2020). Peer loss functions: Learning from noisy labels without knowing noise rates. International Conference on Machine Learning (pp. 6226–6236) Liu, S., Niles-Weed, J., Razavian, N., Fernandez-Granda, C. (2020). Early-learning regularization prevents memorization of noisy labels. Advances in Neural Information Processing Systems (Vol. 33, pp. 20331–20342). Lukasik, M., Bhojanapalli, S., Menon, A., & Kumar, S. (2020). Does label smoothing mitigate label noise? International Conference on Machine Learning (pp. 6448–6458) Northcutt, C., Jiang, L., & Chuang, I. (2021). Confident learning: Estimating uncertainty in dataset labels. Journal of Artificial Intelligence Research, 70, 1373–1411. Ortego, D., Arazo, E., Albert, P., O’Connor, N.E., McGuinness, K. (2021). Multi-objective interpolation training for robustness to label noise. Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR) (pp. 6606–6615). Patrini, G., Rozza, A., Krishna Menon, A., Nock, R., & Qu, L. (2017). Making deep neural networks robust to label noise: A loss correction approach. Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR) (pp. 1944–1952) Reed, S., Lee, H., Anguelov, D., Szegedy, C., Erhan, D., & Rabinovich, A. (2014). Training deep neural networks on noisy labels with bootstrapping. arxiv preprint arxiv:1412.6596. Simonyan, K., & Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. arxiv preprint arxiv:1409.1556. Smart, B., & Carneiro, G. (2023). Bootstrapping the relationship between images and their clean and noisy labels. Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR) (pp. 5344–5354) Sohn, K., Berthelot, D., Carlini, N., Zhang, Z., Zhang, H., Raffel, C. A., & Li, C.-L. (2020). Fixmatch: Simplifying semi-supervised learning with consistency and confidence. Advances in Neural Information Processing Systems, 33, 596–608. Song, H., Kim, M., Lee, J G. (2019). "SELFIE": Refurbishing Unclean Samples for Robust Deep Learning. International Conference on Machine Learning (ICML). Tanaka, D., Ikami, D., Yamasaki, T., & Aizawa, K. (2018). Joint optimization framework for learning with noisy labels. Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR) (pp. 5552–5560) Wang, X., Lan, L., Wu, X., Yu, J., Yang, W., & Liu, T. (2024). Tackling noisy labels with network parameter additive decomposition. IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(9), 6341–6354. Wang, Y., Ma, X., Chen, Z., Luo, Y., Yi, J., Bailey, J. (2019). Symmetric Cross Entropy for Robust Learning With Noisy Labels. 2019 IEEE/CVF International Conference on Computer Vision (ICCV) (p.322-330). Wei, Feng, L., Chen, X., An, B. (2020). Combating noisy labels by agreement: A joint training method with co-regularization. Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR) (p.13723-13732). Wei, Shi, J. X., Tu, W. W., & Li, Y. F. (2021). Robust long-tailed learning under label noise. arxiv preprint arxiv:2108.11569. Wei, & zhang, Y. (2022). Learning with noisy labels revisited: A study using real-world human annotations. International Conference on Learning Representations (ICLR). https://openreview.net/forum?id=TBWA6PLJZQm Wei, J., Liu, H., Liu, T., Niu, G., & Liu, Y. (2021). Understanding (generalized) label smoothing when learning with noisy labels. CoRR, arxiv:2106.04149 Wu, Y., Shu, J., Xie, Q., Zhao, Q., & Meng, D. (2021). Learning to purify noisy labels via meta soft label corrector. Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 35, pp. 10388–10396) Xia, X., Liu, T., Wang, N., Han, B., Gong, C., Niu, G., & Sugiyama, M. (2019). Are anchor points really indispensable in label-noise learning? Advances in Neural Information Processing Systems (Vol. 32) Xiao, T., Xia, T., Yang, Y., Huang, C., & Wang, X. (2015). Learning from massive noisy labeled data for image classification. Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR) (pp. 2691–2699) Yi, K., & Wu, J. (2019). Probabilistic end-to-end noise correction for learning with noisy labels. Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR) (pp. 7017–7025) Yu, X., Han, B., Yao, J., Niu, G., Tsang, I., & Sugiyama, M. (2019). How does disagreement help generalization against label corruption? Proceedings of the 36th International Conference on Machine Learning (Vol. 97, pp. 7164–7173) PMLR. Zhang, B. S., Hardt, M., Recht, B., & Vinyals, O. (2021). Understanding deep learning (still) requires rethinking generalization. Commun. ACM, 64(3), 107–115. Zhang, C., Dauphin, M., Lopez-Paz, Y. N., & D. (2017). mixup: Beyond empirical risk minimization. arxiv preprint arxiv:1710.09412. Zhang, & Sabuncu, M.R. (2018). Generalized cross entropy loss for training deep neural networks with noisy labels. In: Proceedings of the 32nd International Conference on Neural Information Processing Systems (p.8792–8802). Red Hook, NY, USA: Curran Associates Inc. Zhang, S. B., Wang, H., Han, B., Liu, T., Liu, L., & Sugiyama, M. (2024). Badlabel: A robust perspective on evaluating and enhancing label-noise learning. IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(6), 4398–4409. Zhang, Z., Wu, S., Goswami, P., & Chen, M. C. (2021). Learning with feature-dependent label noise: A progressive approach. arxiv preprint arxiv:2103.07756. Zhao, G., Li, G., Qin, Y., Liu, F., Yu, Y. (2022). Centrality and consistency: Two-stage clean samples identification for learning with instance-dependent noisy labels. Avidan, S., Brostow, G., CissΓ©, M., Farinella, G.M., & Hassner, T. (Eds.), Computer Vision – ECCV 2022 (pp. 21–37). Springer Nature Switzerland. Zhu, Z., Liu, T., Liu, Y. (2021). A second-order approach to learning with instance-dependent label noise. Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR) (pp. 10113–10123). Funding This work was supported in part by the National Natural Science Foundation of China under Grant U21A20513, Grant 62476157, Grant 62276161 and Grant 62576201. Author information Authors and Affiliations Contributions All authors contributed to the conception and design of the proposed method. Theoretical analysis were completed by Jia Zhang, Gaoxia Jiang. Experimental were completed by Jia Zhang. The manuscript was supervised and edited by Gaoxia Jiang, Wenjian Wang. Corresponding author Additional information Editor: Peng Zhao. Publisher's Note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations. Rights and permissions Springer Nature or its licensor (e.g. a society or other partner) holds exclusive rights to this article under a publishing agreement with the author(s) or other rightsholder(s); author self-archiving of the accepted manuscript version of this article is solely governed by the terms of such publishing agreement and applicable law. About this article Cite this article Zhang, J., Jiang, G., Hou, S. et al. Certainty-Aware Partition and Sufficient Utilization of Noisy Samples. Mach Learn 115, 204 (2026). https://doi.org/10.1007/s10994-026-07138-3 Received: Revised: Accepted: Published: Version of record: DOI: https://doi.org/10.1007/s10994-026-07138-3

How it works

Once you click Generate, Ollama reads this article and crafts 5 comprehension questions. Your answers are graded against the article content β€” general knowledge won't be enough. Score 70+ to count toward your certificate.

Questions are cached β€” you'll always get the same 5 for this article.