TY - GEN
T1 - Are a Thousand Words Better Than a Single Picture? Beyond Images - A Framework for Multi-modal Knowledge Graph Dataset Enrichment
AU - Zhang, Pengyu
AU - Zaporojets, Klim
AU - Liu, Jie
AU - Huang, Jia Hong
AU - Groth, Paul
N1 - Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Switzerland AG 2026.
PY - 2026
Y1 - 2026
N2 - Multi-Modal Knowledge Graphs (MMKGs) benefit from visual information, yet large-scale image collection is hard to curate and often excludes ambiguous but relevant visuals (e.g., logos, symbols, abstract scenes). We present Beyond Images, an automatic data-centric enrichment pipeline with optional human auditing. This pipeline operates in three stages: (1) large-scale retrieval of additional entity-related images, (2) conversion of all visual inputs into textual descriptions to ensure that ambiguous images contribute usable semantics rather than noise, and (3) fusion of multi-source descriptions using a large language model (LLM) to generate concise, entity-aligned summaries. These summaries replace or augment the text modality in standard MMKG models without changing their architectures or loss functions. Across three public MMKG datasets and multiple baseline models, we observe consistent gains (up to +7% Hits@1 overall). Furthermore, on a challenging subset of entities with visually ambiguous logos and symbols, converting images into text yields large improvements (+201.35% MRR and +333.33% Hits@1). Additionally, we release a lightweight Text–Image Consistency Check Interface for optional targeted audits, improving description quality and dataset reliability. Our results show that scaling image coverage and converting ambiguous visuals into text is a practical path to stronger MMKG completion. Code, datasets, and supplementary materials are available at https://github.com/pengyu-zhang/Beyond-Images.
AB - Multi-Modal Knowledge Graphs (MMKGs) benefit from visual information, yet large-scale image collection is hard to curate and often excludes ambiguous but relevant visuals (e.g., logos, symbols, abstract scenes). We present Beyond Images, an automatic data-centric enrichment pipeline with optional human auditing. This pipeline operates in three stages: (1) large-scale retrieval of additional entity-related images, (2) conversion of all visual inputs into textual descriptions to ensure that ambiguous images contribute usable semantics rather than noise, and (3) fusion of multi-source descriptions using a large language model (LLM) to generate concise, entity-aligned summaries. These summaries replace or augment the text modality in standard MMKG models without changing their architectures or loss functions. Across three public MMKG datasets and multiple baseline models, we observe consistent gains (up to +7% Hits@1 overall). Furthermore, on a challenging subset of entities with visually ambiguous logos and symbols, converting images into text yields large improvements (+201.35% MRR and +333.33% Hits@1). Additionally, we release a lightweight Text–Image Consistency Check Interface for optional targeted audits, improving description quality and dataset reliability. Our results show that scaling image coverage and converting ambiguous visuals into text is a practical path to stronger MMKG completion. Code, datasets, and supplementary materials are available at https://github.com/pengyu-zhang/Beyond-Images.
KW - Dataset Enrichment
KW - Entity Representation
KW - Image-to-Text
KW - Link Prediction
KW - Multi-modal Knowledge Graphs
UR - https://www.scopus.com/pages/publications/105039921966
U2 - 10.1007/978-3-032-25156-5_5
DO - 10.1007/978-3-032-25156-5_5
M3 - Contribución a la conferencia
AN - SCOPUS:105039921966
SN - 9783032251558
T3 - Lecture Notes in Computer Science
SP - 82
EP - 101
BT - The Semantic Web - 23rd European Semantic Web Conference, ESWC 2026, Proceedings
A2 - Acosta, Maribel
A2 - van Erp, Marieke
A2 - Rudolph, Sebastian
A2 - Hartig, Olaf
A2 - Spahiu, Blerina
A2 - Rula, Anisa
A2 - Garijo, Daniel
A2 - Osborne, Francesco
PB - Springer Science and Business Media Deutschland GmbH
T2 - 23rd European Semantic Web Conference, ESWC 2026
Y2 - 10 May 2026 through 14 May 2026
ER -