The ability to reason with and integrate different sensory inputs is the foundation underpinning human intelligence, and it is the reason for the growing interest in modeling multi-modal information within Knowledge Graphs.
Multi-Modal Knowledge Graphs (MMKGs) extend traditional Knowledge Graphs by linking entities to representations across multiple media (e.g., text, images, audio, video) so that these representations can holistically contribute to an entity’s semantics.
Despite the growing interest, the notion of modality is modeled heterogeneously across existing ontologies: some approaches conflate modality with content, while others separate modality types from their concrete artifacts and attach modality-specific metadata at different levels.
This heterogeneity complicates reuse and interoperability.
In this PhD position, we propose a deep investigation into the principles, modeling methodologies, and applications of the Multi-modal Ontology Design Pattern (ODP).
The main objectives are to provide a domain-agnostic structural backbone for multi-modal representation by separating (i) a multi-modal entity as an information object, (ii) its concrete digital realizations as modal descriptors, and (iii) reusable modality specifications, including support for modality composition.
The study path is driven by competency questions derived from a scoping review and is accompanied by query-shaped operationalizations and reuse guidelines to support consistent instantiation.
It is expected to validate the ODP by integrating it into a set of existing MMKGs and aligning them with representative multimodal ontologies from different domains, thereby demonstrating its coverage of recurring modeling structures and its extensibility in several industrial settings.