中文 | English
合肥工业大学校徽 合肥工业大学学报自科版

导航菜单

基于张量分解的多模态知识图谱嵌入

Multimodal knowledge graph embedding based on tensor decomposition

期刊信息

合肥工业大学(自然科学版),2026年5月,第49卷第5期:629-636

DOI: 10.3969/j.issn.1003-5060.2026.05.008

作者信息

胡雨琪 $ ^{1} $,杨依忠 $ ^{1} $,张晴宇 $ ^{2} $

(1. 合肥工业大学微电子学院, 安徽 合肥 230601; 2. 合肥工业大学计算机与信息学院, 安徽 合肥 230601)

摘要和关键词

摘要: 由于不同模态信息之间存在异构性,简单的多模态信息融合方式无法有效捕捉深层语义信息,难以辅助多模态知识在低维连续空间中的张量嵌入。为了探究视觉模态信息优化表示学习的有效方法,文章提出一种基于张量分解的多模态知识表示学习模型,旨在解决多源知识融合和多模态知识嵌入的问题。该模型首先利用多模态数据筛选模块对视觉信息进行基础编码,并计算图像的公共相似度过滤潜在的信息干扰;然后在多模态数据融合模块中设置权重将不同模态信息进行融合,从而得到初始化的多模态知识编码;最后在多模态知识嵌入模块中通过张量分解原理,将多模态知识编码嵌入到低维连续张量空间;此外还设计了基于多预测任务的联合损失函数,以提高模型训练效率和知识嵌入效果。链接预测实验结果表明,该模型在多个指标上均超越了经典基线模型,说明通过合理的数据过滤、编码与嵌入方法,多模态信息能够显著提高知识表示效果。

关键词: 多模态信息;知识图谱;知识表示;信息融合;知识挖掘

Authors

HU Yuqi $ ^{1} $, YANG Yizhong $ ^{1} $, ZHANG Qingyu $ ^{2} $

(1. School of Microelectronics, Hefei University of Technology, Hefei 230601, China; 2. School of Computer Science and Information Engineering, Hefei University of Technology, Hefei 230601, China)

Abstract and Keywords

Abstract: Due to the heterogeneity between different modal information, simple multimodal information fusion methods cannot effectively capture the deep semantic information, which makes it difficult to assist the tensor embedding of multimodal knowledge in low-dimensional continuous space. In order to explore an effective method for optimal representation learning of visual modal information, this paper proposes a multimodal knowledge representation learning model based on tensor decomposition, aiming to solve the problems of multi-source knowledge fusion and multimodal knowledge embedding. The model first uses a multimodal data filtering module to perform basic encoding on the visual information, and calculates the common similarity of images to filter potential information interference. Then, the weights are set in the multimodal data fusion module to fuse the information from different modalities to obtain the initialized multimodal knowledge encoding. Finally, the multimodal knowledge encoding is embedded into a low-dimensional continuous tensor space through the tensor decomposition principle in the multimodal knowledge embedding module. The paper also designs a joint loss function based on the multi-prediction task to improve the model training efficiency and knowledge embedding effect. The results of the link prediction experiments show that the model exceeds the classical baseline model in several indexes, indicating that the multimodal information can significantly improve the effect of knowledge representation through reasonable data filtering, encoding and embedding methods.

Keywords: multimodal information; knowledge graph; knowledge representation; information fusion; knowledge mining

基金信息

安徽省自然科学基金资助项目(2208085MF177)

个人中心