TopoMoE: a topology-fidelity discrete mixture-of-experts framework for multi
Abstract
Multi-modal knowledge graph completion (MMKGC) predicts missing facts by integrating structured triples with heterogeneous visual and textual information. Large-scale MMKGC is also computation-intensive because fine-grained token encoding, expert transformation, contrastive optimization, and candidate-entity scoring must be repeatedly performed over high-dimensional representations. However, conventional MoE routing and contrastive learning remain insufficient when used side by side. Smooth dense routing may mix noisy heterogeneous signals across multiple expert outputs and weaken specialization, leading to what we term “gradient blurring,” whereas overly sharp contrastive optimization may excessively separate structurally related entities and distort local neighborhoods, a potential effect termed “topology tearing.” Routing controls how features are transformed but not their inter-entity geometry, while contrastive learning shapes that geometry without controlling expert allocation. To bridge this gap, we propose TopoMoE, a topology-fidelity discrete mixture-of-experts framework that coordinates input purification, sparse routing, contrastive optimization, and algebraic scoring. Modality-aware gating and an asymmetric information bottleneck first suppress unreliable modality-specific signals. Gumbel-Softmax exploration followed by discrete Top-K selection then assigns each token to selected expert pathways, promoting differentiated feature processing. Lower-bounded adaptive temperatures and curriculum-driven hybrid hard-negative mining subsequently regulate the sharpness and sample composition of contrastive supervision. Finally, a zero-MLP TuckER scorer directly evaluates the optimized representations without additional MLP-based projection layers. This composition connects cleaner routing inputs and differentiated expert processing with bounded contrastive optimization, while retaining batched tensor operations suitable for accelerator execution. Across five independent runs, TopoMoE achieves an MRR of \(39.18 \pm 0.21\%\) and a Hit@1 of \(31.36 \pm 0.24\%\) on DB15K, and an MRR of \(37.78 \pm 0.11\%\) and a Hit@1 of \(31.80 \pm 0.19\%\) on MKG-W. Supplementary reference-value t-tests yield \(p<0.01\) across the evaluated metrics. Direct Topological Neighborhood Preservation analysis shows the clearest fusion-related improvement in the DB15K Degree \(>30\) group, where TNP increases from \(12.90\%\) to \(19.80\%\). In the controlled single-GPU efficiency comparison, TopoMoE records 155.01 G estimated FLOPs and 16.02 s per epoch, compared with 228.24 G and 17.01 s for the dense baseline. These results support a favorable accuracy–computation trade-off and topology-preservation behavior, particularly in dense structural regions, within the evaluated MMKGC settings.
Data availability
No datasets were generated or analyzed during the current study.
References
Liang W, Meo PD, Tang Y, Zhu J (2024) A survey of multi-modal knowledge graphs: technologies and trends. ACM Comput Surv 56(11):1. https://doi.org/10.1145/3656579
Liu X, Mao T, Shi Y, Ren Y (2024) Overview of knowledge reasoning for knowledge graph. Neurocomput 585(C):1. https://doi.org/10.1016/j.neucom.2024.127571
Cao J, Fang J, Meng Z, Liang S (2024) Knowledge graph embedding: a survey from the perspective of representation spaces 56(6). https://doi.org/10.1145/3643806
Ju W, Fang Z, Gu Y, Liu Z, Long Q, Qiao Z, Qin Y, Shen J, Sun F, Xiao Z et al (2024) A comprehensive survey on deep graph representation learning. Neural Netw 173:106207
Ge X, Wang Y-C, Wang B, Kuo C-CJ (2023) Knowledge graph embedding: an overview. arXiv preprint arXiv:2309.12501
Luvembe AM, Li W, Li S, Liu F, Wu X (2024) Caf-odnn: complementary attention fusion with optimized deep neural network for multimodal fake news detection. Inf Process Manag 61(3):103653
Li M, Zhuang X, Bai L, Ding W (2024) Multimodal graph learning based on 3d haar semi-tight framelet for student engagement prediction. Inf Fusion 105:102224
Park S, Lee D, Park H (2024) Enhancing knowledge tracing with concept map and response disentanglement. Knowl Based Syst 302:112346. https://doi.org/10.1016/j.knosys.2024.112346
Wang Y, Gou X, Xu X, Geng Y, Ke X, Wu T, Yu Z, Chen R, Wu X (2024) Scalable community search over large-scale graphs based on graph transformer. In: SIGIR ’24, Association for Computing Machinery, New York, NY, USA, pp 1680–1690. https://doi.org/10.1145/3626772.3657771
Li Y, Ji H, Yu F, Cheng L, Che N (2025) Temporal multi-modal knowledge graph generation for link prediction. Neural Netw 185:107108
Lu H-Y, Yu H-K, Fan C, Zhan Q, Fang W, Wu X-J (2025) Temme: Temporal knowledge graph completion using multi-grade multivector embeddings. In: Hadfi R, Anthony P, Sharma A, Ito T, Bai Q (eds), PRICAI 2024: Trends in Artificial Intelligence. Springer, Singapore
Balazevic I, Allen C, Hospedales T, Inui K, Jiang J, Ng V (2019) TuckER: tensor factorization for knowledge graph completion. In: Wan X (ed), Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). Association for Computational Linguistics, Hong Kong, China, pp 5185–5194. https://doi.org/10.18653/v1/D19-1522
Guo Y, Ma Q, Li H, Ning Q, Zhan F, Gu Y, Yu G, Guo S (2026) Lbmkgc: large model-driven balanced multimodal knowledge graph completion. Adv Neural Inf Process Syst 38:111874–111896
Geng Y, Chen J, Zeng Y, Chen Z, Zhang W, Pan JZ, Wang Y, Xu X (2025) Prompting disentangled embeddings for knowledge graph completion with pre-trained language model. Expert Syst Appl 268:126175
Lu H-Y, Li X-F (2025) A novel mllms-based two-stage model for zero-shot multimodal sentiment analysis. In: Hadfi R, Anthony P, Sharma A, Ito T, Bai Q (eds), PRICAI 2024: Trends in Artificial Intelligence, Springer, Singapore, pp 436–449
Cui Y, Sun Z, Hu W (2024) A prompt-based knowledge graph foundation model for universal in-context reasoning. Adv Neural Inf Process Syst 37:7095–7124. https://doi.org/10.52202/079017-0227
Zhang Y, Huang Q, Wang H (2026) M2f-net: multi-scale multi-frequency fusion network for image compressed sensing. Expert Syst Appl 319:132079. https://doi.org/10.1016/j.eswa.2026.132079
Shang B, Zhao Y, Liu J (2024) Learnable convolutional attention network for knowledge graph completion. Knowl Based Syst 285:111360
Wang J, Li W, Liu F, Wang Z, Luvembe AM, Jin Q, Pan Q, Liu F (2024) Conee: global and local context-enhanced embedding for inductive knowledge graph completion. Expert Syst Appl 246:123116
Wang H, Song D, Wu Z, Tian Y, Xu J (2025) Look one step ahead through first-order aggregation in reinforcement learning-based knowledge graph reasoning. Inf Sci 718:122373
Kim S, Yun S, Lee J, Chang G, Roh W, Sohn D-N, Lee J-T, Park H, Kim S (2024) Self-supervised multimodal graph convolutional network for collaborative filtering. Inf Sci 653:119760
Wang J, Xie H, Zhang S, Qin SJ, Tao X, Wang FL, Xu X (2025) Multimodal fusion framework based on knowledge graph for personalized recommendation. Expert Syst Appl 268:126308
Zhang Y, Chen Z, Guo L, Xu Y, Hu B, Liu Z, Zhang W, Chen H (2025) Tokenization, fusion, and augmentation: towards fine-grained multi-modal entity representation. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol 39, pp 13322–13330
Zhang Y, Chen Z, Guo L, Hu B, Liu Z, Zhang W, Chen H et al (2025) Multiple heads are better than one: mixture of modality knowledge experts for entity representation learning. In: International Conference on Learning Representations, vol 2025, pp 54811–54828
Sanga P, Singh J, Chakraborti T (2026) KG-MoE: multimodal Knowledge Graph Grounded Mixture of Experts for Fair Visual Question Answering. https://openreview.net/forum?id=muJ0EYIXFC
Du E, Liu S, Zhang Y (2025) Mixture of length and pruning experts for knowledge graphs reasoning. In: Christodoulopoulos C, Chakraborty T, Rose C, Peng V (eds), Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, Suzhou, China, pp 432–453. https://doi.org/10.18653/v1/2025.emnlp-main.23
Zhang Y, Wang H, Xia Y (2026) Mmc-cs: multi-branch multi-stage contrastive learning for self-supervised compressed sensing. Neural Netw 197:108475. https://doi.org/10.1016/j.neunet.2025.108475
Huang Z, Chen H, Wen Z, Zhang C, Li H, Wang B, Chen C (2023) Model-aware contrastive learning: towards escaping the dilemmas. In: International Conference on Machine Learning, PMLR, pp 13774–13790
Qian Y, Ma T, Zhang C, Ye Y (2024) Adaptive Temperature Enhanced Dual-level Hypergraph Contrastive Learning. https://openreview.net/forum?id=scxDIx6StY
Kim BJ, Kim SW (2026) Temperature-free loss function for contrastive learning. Neural Netw 109222:1
Wang Y, Huang C, Li M, Huang Q, Wu X, Wu J (2024) Ag-meta: adaptive graph meta-learning via representation consistency over local subgraphs. Pattern Recogn 151:110387. https://doi.org/10.1016/j.patcog.2024.110387
Li R, Zhong J, Hu W, Dai Q, Wang C, Wang W, Li X (2024) Adaptive class augmented prototype network for few-shot relation extraction. Neural Netw 169:134–142. https://doi.org/10.1016/j.neunet.2023.10.025
Zhang H, Zhang J, Molybog I (2024) Hasa: hardness and structure-aware contrastive knowledge graph embedding. In: Proceedings of the ACM Web Conference 2024. WWW ’24, Association for Computing Machinery, New York, NY, USA, pp 2116–2127. https://doi.org/10.1145/3589334.3645564
Qiao Z, Ye W, Yu D, Mo T, Li W, Zhang S (2023) Improving knowledge graph completion with generative hard negative mining. In: Rogers A, Boyd-Graber J, Okazaki N (eds), Findings of the Association for Computational Linguistics: ACL 2023. Association for Computational Linguistics, Toronto, Canada, pp 5866–5878. https://doi.org/10.18653/v1/2023.findings-acl.362
Takamoto M, Oñoro-Rubio D, Rim WB, Maruyama T, Kotnis B (2025) Optimal Embedding Guided Negative Sample Generation for Knowledge Graph Link Prediction. arXiv:org/abs/2504.03327
Zhang Y, Zhou Y, Jiang H, Wang H (2026) Cas-distillcs: confidence-aware self-distillation for self-supervised compressed sensing. Knowl Based Syst 349:116457. https://doi.org/10.1016/j.knosys.2026.116457
Bordes A, Usunier N, Garcia-Duran A, Weston J, Yakhnenko O (2013) Translating embeddings for modeling multi-relational data. Adv Neural Inf Process Syst 26:1
Yang B, Yih W-T, He X, Gao J, Deng L (2014) Embedding entities and relations for learning and inference in knowledge bases. arXiv preprint arXiv:1412.6575
Trouillon T, Welbl J, Riedel S, Gaussier E, Bouchard G (2016) Complex embeddings for simple link prediction. In: Balcan MF, Weinberger KQ (eds), Proceedings of The 33rd International Conference on Machine Learning. Proceedings of Machine Learning Research, vol 48, PMLR, New York, New York, USA, pp 2071–2080
Sun Z, Deng Z-H, Nie J-Y, Tang J (2019) Rotate: knowledge graph embedding by relational rotation in complex space. https://doi.org/10.48550/arXiv.1902.10197. arXiv preprint arXiv:1902.10197
Xie R, Liu Z, Luan H, Sun M (2016) Image-embodied knowledge representation learning. arXiv preprint arXiv:1609.07028
Mousselly-Sergieh H, Botschen T, Gurevych I, Roth S (2018) A multimodal translation-based approach for knowledge graph representation learning. In: Nissim M, Berant J, Lenci A (eds), Proceedings of the Seventh Joint Conference on Lexical and Computational Semantics, Association for Computational Linguistics, New Orleans, Louisiana, pp 225–234. https://doi.org/10.18653/v1/S18-2027
Wang Z, Li L, Li Q, Zeng D (2019) Multimodal data enhanced representation learning for knowledge graphs. In: 2019 International Joint Conference on Neural Networks (IJCNN), pp 1–8. https://doi.org/10.1109/IJCNN.2019.8852079
Lu X, Wang L, Jiang Z, He S, Liu S (2022) Mmkrl: a robust embedding approach for multi-modal knowledge graph representation learning. Appl Intell 52(7):7480–7497. https://doi.org/10.1007/s10489-021-02693-9
Wang M, Wang S, Yang H, Zhang Z, Chen X, Qi G (2021) Is visual context really helpful for knowledge graph? A representation learning perspective. In: Proceedings of the 29th ACM International Conference on Multimedia. MM ’21, Association for Computing Machinery, New York, NY, USA, pp 2735–2743. https://doi.org/10.1145/3474085.3475470
Zhang Y, Zhang W (2022) Knowledge graph completion with pre-trained multimodal transformer and twins negative sampling. arXiv preprint arXiv:2209.07084
Cao Z, Xu Q, Yang Z, He Y, Cao X, Huang Q (2022) Otkge: multi-modal knowledge graph embeddings via optimal transport. Adv Neural Inf Process Syst 35:39090–39102
Zhang Y, Chen Z, Zhang W (2023) Maco: a modality adversarial and contrastive framework for modality-missing multi-modal knowledge graph completion. In: Liu F, Duan N, Xu Q, Hong Y (eds), Natural Language Processing and Chinese Computing, Springer, Cham, pp 123–134
Li X, Zhao X, Xu J, Zhang Y, Xing C (2023) Imf: interactive multimodal fusion model for link prediction. In: Proceedings of the ACM Web Conference 2023. WWW ’23, Association for Computing Machinery, New York, NY, USA, pp 2572–2580. https://doi.org/10.1145/3543507.3583554
Wang X, Meng B, Chen H, Meng Y, Lv K, Zhu W (2023) Tiva-kg: a multimodal knowledge graph with text, image, video and audio. In: Proceedings of the 31st ACM International Conference on Multimedia. MM ’23, Association for Computing Machinery, New York, NY, USA, pp 2391–2399. https://doi.org/10.1145/3581783.3612266
Lee J, Chung C, Lee H, Jo S, Whang JJ (2023) VISTA: visual-textual knowledge graph representation learning. In: Bouamor H, Pino J, Bali K (eds), Findings of the Association for Computational Linguistics: EMNLP 2023, Association for Computational Linguistics, Singapore, pp 7314–7328. https://doi.org/10.18653/v1/2023.findings-emnlp.488
Zhang Y, Chen Z, Liang L, Chen H, Zhang W (2024) Unleashing the power of imbalanced modality information for multi-modal knowledge graph completion. In: Calzolari N, Kan M-Y, Hoste V, Lenci A, Sakti S, Xue N (eds), Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), ELRA and ICCL, Torino, Italia, pp 17120–17130
Zhang Y, Chen M, Zhang W (2023) Modality-aware negative sampling for multi-modal knowledge graph embedding. In: 2023 International Joint Conference on Neural Networks (IJCNN), pp 1–8. https://doi.org/10.1109/IJCNN54540.2023.10191314
Xu D, Xu T, Wu S, Zhou J, Chen E (2022) Relation-enhanced negative sampling for multimodal knowledge graph completion. In: Proceedings of the 30th ACM International Conference on Multimedia. MM ’22, Association for Computing Machinery, New York, NY, USA, pp 3857–3866. https://doi.org/10.1145/3503161.3548388
Funding
This research was supported by the National Natural Science Foundation of China (Grant Nos. 62376089, 62302153, 62302154), the key Research and Development Program of Hubei Province, China (Grant No. 2023BEB024) and the Young and Middle-aged Scientific and Technological Innovation Team Plan in Higher Education Institutions in Hubei Province, China (Grant No. T2023007).
Author information
Authors and Affiliations
Contributions
All authors contributed equally to the research and preparation of this manuscript.
Corresponding author
Ethics declarations
Conflict of interest
The authors declare no financial or personal conflict of interest.
Ethical approval
This study does not involve human or animal subjects. Ethical approval and informed consent are not applicable.
Additional information
Publisher's Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Rights and permissions
Springer Nature or its licensor (e.g. a society or other partner) holds exclusive rights to this article under a publishing agreement with the author(s) or other rightsholder(s); author self-archiving of the accepted manuscript version of this article is solely governed by the terms of such publishing agreement and applicable law.
About this article
Cite this article
Chen, H., Liu, C. & Shao, P. TopoMoE: a topology-fidelity discrete mixture-of-experts framework for multi-modal knowledge graph completion. J Supercomput 82, 697 (2026). https://doi.org/10.1007/s11227-026-08856-0
Received:
Accepted:
Published:
Version of record:
DOI: https://doi.org/10.1007/s11227-026-08856-0
How it works
Once you click Generate, Ollama reads this article and crafts 5 comprehension questions. Your answers are graded against the article content — general knowledge won't be enough. Score 70+ to count toward your certificate.
Questions are cached — you'll always get the same 5 for this article.