Knowledge Graph Completion Models
1. Definition and Components of Knowledge Graphs
Definition and Components of Knowledge Graphs
A knowledge graph (KG) is a structured representation of real-world entities and their relationships, typically modeled as a directed, labeled multigraph. Formally, a KG is defined as a tuple G = (E, R, T), where:
- E is a set of entities (nodes), representing objects, concepts, or instances.
- R is a set of relation types (edge labels), describing how entities are connected.
- T is a set of triples (edges) of the form (h, r, t), where h, t ∈ E and r ∈ R.
This triple structure encodes factual knowledge, with h (head) and t (tail) entities linked by relation r. For example, the triple (Albert_Einstein, won, Nobel_Prize) asserts a factual relationship between two entities.
Semantic Foundations
Knowledge graphs are grounded in formal semantics, often implementing the Resource Description Framework (RDF) data model. Each triple corresponds to an RDF statement, where:
This aligns with first-order logic, where relations are binary predicates. KGs extend simple graphs by allowing:
- Node and edge attributes (e.g., temporal validity, confidence scores)
- Type hierarchies through ontological constructs like rdf:type and rdfs:subClassOf
- Inference rules via formalisms like OWL axioms or rule-based systems
Core Components
1. Entity Representations
Entities are typically disambiguated through unique identifiers (URIs) and may include:
- Lexical information: Names, synonyms, descriptions in multiple languages
- Structured attributes: Numerical properties (e.g., population for cities), categorical types
- Embeddings: Vector representations learned via KG completion models
2. Relation Taxonomy
Relations exhibit hierarchical and logical structures:
- Domain/Range constraints: bornIn may link Person → Place
- Property chains: hasChild ◦ hasSpouse ⇒ hasSonInLaw
- Inverse relations: employedBy inverse of employs
3. Ontological Layer
Upper-level schemas define logical constraints:
Such axioms enable automated reasoning—for instance, inferring that all professors are persons without explicit assertions.
Practical Implementation
Modern KGs like Wikidata and Google's Knowledge Graph implement these components at scale:
- Wikidata: 100M+ entities with crowdsourced assertions and provenance tracking
- Enterprise KGs: Combine internal data (product catalogs) with external sources (DBpedia)
- Embedding-based: Translate components to vector spaces for machine learning compatibility
The graph structure enables efficient path queries (e.g., "List scientists educated in Germany who won Nobel Prizes") through SPARQL or graph traversal algorithms.

1.2 Common Knowledge Graph Datasets and Benchmarks
Standard Datasets for Knowledge Graph Completion
Knowledge graph completion (KGC) models are typically evaluated on well-established datasets that provide structured triples (head entity, relation, tail entity) alongside train/validation/test splits. The most widely used datasets include:
- FB15k & FB15k-237: Derived from Freebase, FB15k contains 15,000 entities with 1,345 relation types. FB15k-237 is a reduced subset eliminating inverse relations, making link prediction more challenging.
- WN18 & WN18RR: Based on WordNet, WN18 contains 18 relations and 40,943 entities. WN18RR removes reversible relations to prevent test leakage, serving as a more rigorous benchmark.
- YAGO3-10: A subset of YAGO with 123,182 entities and 37 relations, focusing on entities with at least 10 descriptive facts.
Specialized and Domain-Specific Benchmarks
Beyond general-purpose datasets, specialized benchmarks evaluate KGC models in constrained settings:
- CoDEx: A family of datasets (CoDEx-S, CoDEx-M, CoDEx-L) with textual descriptions and hard negative samples to test contextual understanding.
- OGBL-BioKG: A biomedical knowledge graph with 93,773 entities and 107 relations, used for drug discovery and biological pathway prediction.
- DBpedia50: A dense subset of DBpedia with fine-grained entity typing, useful for hierarchical relation learning.
Evaluation Metrics and Protocols
Standard evaluation employs ranking-based metrics:
where Q is the set of test queries and ranki is the position of the correct answer. The filtered setting removes all other valid triples during ranking to avoid artificial inflation of metrics.
Challenges and Dataset Biases
Common pitfalls in benchmark interpretation include:
- Inverse relation leakage: Models may exploit reversible patterns (e.g., bornIn vs. cityOfBirth) unless datasets are carefully pruned.
- Degree imbalance: Long-tail entity distributions can lead to overfitting on high-degree nodes.
- Temporal drift: Static snapshots (e.g., Freebase-based datasets) may not reflect evolving real-world knowledge.
Emerging Benchmarks
Recent datasets address these limitations:
- WikiKG90Mv2: A large-scale benchmark with 90M+ entities and 1,387 relations, featuring multilingual labels and hard negative mining.
- TimeKG: Incorporates temporal scopes for relations, enabling time-aware completion tasks.
- KG-BERT: Benchmarks combining structural and textual evidence using pre-trained language models.
1.3 Applications of Knowledge Graphs in AI
Semantic Search and Information Retrieval
Knowledge graphs enhance search engines by modeling relationships between entities, enabling semantic rather than keyword-based retrieval. Google's Knowledge Graph, for instance, resolves ambiguous queries by leveraging entity disambiguation and contextual relationships. The underlying mechanism involves graph traversal algorithms that compute relevance scores based on path length and node centrality. For a query q, the relevance score R(e|q) of an entity e is often computed using a modified PageRank algorithm:
where N(e) denotes neighboring nodes, α is a damping factor, and P(e|q) is a prior probability derived from query-entity co-occurrence statistics.
Question Answering Systems
Knowledge graphs power dynamic QA systems by mapping natural language queries to structured subgraphs. Systems like IBM Watson employ graph embeddings (e.g., TransE) to project entities and relations into a continuous vector space, enabling operations like:
where h, r, and t are head entity, relation, and tail entity embeddings, respectively. This allows answering complex queries (e.g., "Which scientists worked on quantum mechanics and were born in Germany?") through multi-hop reasoning.
Recommendation Systems
Graph-based collaborative filtering extends matrix factorization by incorporating side information (e.g., item attributes, user demographics) as additional nodes. The recommendation score for user u and item i is computed via graph convolutional networks (GCNs):
where z(l) denotes node embeddings at layer l, and N(u) represents neighboring items. This approach achieves 15–30% higher precision@k than traditional MF in benchmarks like MovieLens.
Drug Discovery and Biomedical Research
Knowledge graphs integrate heterogeneous biological data (protein-protein interactions, drug targets, pathways) to predict drug repurposing candidates. Models like KG-DDI use graph neural networks to infer unknown drug-drug interactions by learning from known triplets in biomedical KGs (e.g., DrugBank, Bio2RDF). The prediction head typically employs a DistMult scoring function:
yielding AUC-ROC scores >0.92 on benchmark datasets.
Enterprise Data Integration
Large corporations use knowledge graphs to unify siloed databases across departments. A financial institution might link customer profiles (CRM), transaction records (ERP), and market data (external APIs) into a unified graph, enabling fraud detection through temporal graph pattern mining. Dynamic graph embeddings (e.g., DySAT) capture evolving relationships:
where At is the adjacency matrix at time t, and Xt contains node features.
2. Link Prediction and Entity Resolution
Link Prediction and Entity Resolution
Link Prediction in Knowledge Graphs
Link prediction aims to infer missing edges between entities in a knowledge graph (KG). Given a KG G = (E, R, T), where E is the set of entities, R the set of relation types, and T the set of triples (h, r, t), the task is to predict whether a candidate triple (h', r', t') holds. This is typically framed as a scoring problem, where a model assigns a plausibility score f(h, r, t) to each triple.
Here, σ is the sigmoid function, h and t are entity embeddings, Mr is a relation-specific transformation matrix, and br is a relation-specific bias term. The model is trained to maximize scores for observed triples while minimizing scores for negative samples.
Entity Resolution and Disambiguation
Entity resolution (ER) identifies when two nodes in a KG refer to the same real-world entity. This is critical for merging KGs or cleaning noisy data. Given two entities e1 and e2, ER computes a similarity metric:
where α and β are weighting factors, simname compares entity names (e.g., using Levenshtein distance), and simneighbors compares their relational contexts (e.g., Jaccard similarity over adjacent nodes).
Joint Learning Frameworks
Recent approaches unify link prediction and entity resolution into a single optimization. For example, the AlignE model jointly learns entity embeddings while minimizing the distance between aligned entities across KGs:
where A is a set of pre-aligned entity pairs, γ is a margin hyperparameter, and λ controls the alignment loss weight. This enables mutual reinforcement between link prediction and entity resolution.
Practical Applications
- Cross-lingual KG alignment: Resolving equivalent entities in multilingual KGs like DBpedia and Wikidata.
- Drug discovery: Predicting protein-drug interactions by completing biomedical KGs.
- Recommendation systems: Linking user profiles across platforms to enrich collaborative filtering.
Evaluation Metrics
Standard benchmarks evaluate models using:
- Mean Reciprocal Rank (MRR): Average reciprocal rank of correct entities in ranked predictions.
- Hits@k: Percentage of correct entities ranked in the top k positions.
- Precision/Recall: For entity resolution, measured over predicted alignments.

Evaluation Metrics for Knowledge Graph Completion
Rank-Based Metrics
Rank-based metrics evaluate the performance of knowledge graph completion models by assessing the ranking of true triples against corrupted ones. Given a test triple (h, r, t), the model generates scores for the original triple and its corrupted variants (e.g., (h', r, t) or (h, r, t')). The rank of the true triple is then determined by sorting all scores in descending order.
Mean Rank (MR): The average rank of true triples across the test set. Lower values indicate better performance, but MR is sensitive to outliers.
Mean Reciprocal Rank (MRR): The average of the reciprocal ranks, emphasizing correct predictions in higher positions.
Hits@k: The fraction of true triples ranked in the top k positions. Common values for k are 1, 3, and 10.
Threshold-Based Metrics
Threshold-based metrics classify predictions as correct or incorrect based on a predefined score threshold.
Precision: The fraction of predicted triples that are correct.
Recall: The fraction of true triples that are correctly predicted.
F1 Score: The harmonic mean of precision and recall, balancing both metrics.
AUC-ROC and AUC-PR
Area Under the Receiver Operating Characteristic Curve (AUC-ROC) and Area Under the Precision-Recall Curve (AUC-PR) provide aggregate measures of model performance across all possible thresholds.
AUC-ROC: Measures the trade-off between true positive rate (TPR) and false positive rate (FPR). A value of 1 indicates perfect classification.
AUC-PR: Focuses on the precision-recall trade-off, particularly useful for imbalanced datasets.
Filtered vs. Raw Metrics
In knowledge graph completion, metrics can be computed in raw or filtered settings:
- Raw Metrics: Evaluate ranks without removing other valid triples from the test set. This can artificially inflate ranks due to the presence of other true triples.
- Filtered Metrics: Remove all other valid triples from the ranking process, providing a more realistic assessment of model performance.
Practical Considerations
When selecting evaluation metrics, consider the following:
- Dataset Characteristics: Imbalanced datasets may favor AUC-PR over AUC-ROC.
- Task Requirements: Hits@k is often preferred for applications where top-ranked predictions are critical.
- Computational Efficiency: Rank-based metrics can be computationally intensive for large knowledge graphs.
2.3 Challenges in Knowledge Graph Completion
Incomplete and Sparse Data
Knowledge graphs (KGs) are inherently incomplete due to the open-world assumption, where missing facts do not imply falsehood. This sparsity arises because real-world knowledge is vast and continuously evolving, making exhaustive data collection impractical. For example, Freebase contains only 22% of person-place-of-birth facts compared to ground truth. The incompleteness manifests as:
- Structural sparsity: Long-tail entities with few connections dominate KG statistics.
- Semantic gaps: Missing relationship types between otherwise well-connected entities.
Long-Tail Entity Distribution
Entity frequency in KGs follows a power-law distribution where a small fraction of entities (e.g., "Barack Obama") have disproportionately many connections, while most entities appear in very few triples. This creates learning bias where models achieve high accuracy on frequent entities but fail on tail entities. The performance drop can be quantified by:
where \(A_{head}\) and \(A_{tail}\) represent accuracy on head vs. tail entities respectively. State-of-the-art models show gaps exceeding 40% on benchmarks like FB15k-237.
Multi-Relational Complexity
Relations in KGs exhibit diverse properties that challenge modeling:
- Compositionality: Paths like \(e_1 \xrightarrow{r_1} e_2 \xrightarrow{r_2} e_3\) may imply \(e_1 \xrightarrow{r_3} e_3\)
- Symmetry/Antisymmetry: \(r(h,t) \Rightarrow r(t,h)\) vs. \(r(h,t) \Rightarrow \neg r(t,h)\)
- Hierarchy: \(r_1(h,t) \Rightarrow r_2(h,t)\) where \(r_2\) is a super-relation
No single embedding space can optimally capture all these patterns simultaneously. For instance, TransE handles compositionality well but fails on symmetric relations, while RotatE models symmetry but struggles with hierarchies.
Temporal Dynamics
Over 30% of facts in temporal KGs like Wikidata change over time. The quadruple \((h,r,t,\tau)\) requires modeling:
where entity/relation embeddings \(\mathbf{h},\mathbf{r},\mathbf{t}\) are time-dependent functions. Current approaches like HyTE and TA-DistMult discretize time into bins, losing continuous temporal granularity.
Scalability vs. Expressiveness Trade-off
Matrix factorization methods like RESCAL achieve high expressiveness with \(O(d^2)\) parameters per relation but become computationally intractable for large KGs. In contrast, translational models (TransE, RotatE) use \(O(d)\) parameters but cannot model complex relations. The trade-off is quantified by the parameter-efficiency ratio:
where MRR is the mean reciprocal rank. Current models cluster into high-\(\rho\) but low-MRR (translational) vs. low-\(\rho\) but high-MRR (tensor factorization) groups.

3. Translational Models: TransE, TransH, and TransR
Translational Models: TransE, TransH, and TransR
TransE: The Foundational Translational Model
TransE (Translational Embedding) is the simplest and most widely used knowledge graph completion model. It represents entities and relations as vectors in the same space, enforcing the translational principle: if a triple (h, r, t) holds, then the embedding of the tail entity t should be close to the embedding of the head entity h plus the relation vector r. Mathematically, this is expressed as:
The scoring function measures the plausibility of a triple using the L1 or L2 distance:
where p is typically 1 or 2. TransE performs well for one-to-one relations but struggles with one-to-many, many-to-one, and many-to-many relations due to its simplistic vector space assumption.
TransH: Handling Complex Relations
TransH (Translational Hyperplane) addresses TransE's limitations by projecting entities onto relation-specific hyperplanes. Each relation r has a hyperplane defined by its normal vector wr, and the entity embeddings are projected as follows:
The scoring function then becomes:
This allows entities to have different representations in different relation contexts, improving performance on complex relational patterns.
TransR: Entity-Relation Separation
TransR takes this further by modeling entities and relations in separate vector spaces. Entities are mapped from entity space to relation space via a projection matrix Mr:
The scoring function is then:
where λ is a regularization term. TransR achieves superior performance on complex relations but at higher computational cost due to the matrix-vector operations.
Comparative Analysis
Key differences between these models include:
- Expressiveness: TransE < TransH < TransR in modeling complex relational patterns.
- Computational Complexity: TransE is the most efficient, while TransR requires significant resources.
- Training Stability: TransE is easier to train, while TransR requires careful hyperparameter tuning.
Empirical studies show that TransH and TransR outperform TransE on benchmarks like FB15k and WN18, particularly for multi-mapping relations. However, recent work has shown that simpler models with careful training can sometimes match or exceed the performance of these more complex architectures.

Semantic Matching Models: RESCAL, DistMult, and ComplEx
Tensor Factorization and Semantic Matching
Semantic matching models operate by decomposing knowledge graphs into latent representations through tensor factorization. Given a knowledge graph represented as a third-order binary tensor X ∈ {0,1}Ne×Nr×Ne, where Ne is the number of entities and Nr is the number of relations, these models learn low-dimensional embeddings that capture the underlying semantics.
RESCAL: Bilinear Tensor Decomposition
The RESCAL model factorizes the knowledge graph tensor using a bilinear approach. Each relation rk is represented as a matrix Rk ∈ ℝd×d, while entities are embedded as vectors ei, ej ∈ ℝd. The scoring function for a triple (ei, rk, ej) is given by:
This formulation captures pairwise interactions between entities through the relation-specific matrix, enabling asymmetric relations but requiring O(d2) parameters per relation, which can be computationally expensive for large knowledge graphs.
DistMult: Diagonal Relation Matrices
DistMult simplifies RESCAL by constraining relation matrices to be diagonal, reducing the number of parameters to O(d) per relation. The scoring function becomes:
While computationally efficient, this symmetric formulation limits DistMult to modeling only symmetric relations, as fijk = fjik for all triples.
ComplEx: Complex-Valued Embeddings
ComplEx extends DistMult by introducing complex-valued embeddings, enabling asymmetric relations while maintaining linear time complexity. Entities and relations are represented as vectors in ℂd, and the scoring function uses the Hermitian dot product:
Here, Re(·) denotes the real part, and e̅j is the complex conjugate of ej. This formulation allows ComplEx to model both symmetric and antisymmetric relations efficiently.
Comparative Analysis
The trade-offs between these models are evident in their expressiveness and computational requirements:
- RESCAL captures complex relational patterns but scales quadratically with embedding dimension.
- DistMult is computationally efficient but limited to symmetric relations.
- ComplEx achieves a balance, modeling asymmetric relations with linear complexity through complex embeddings.
Empirical studies show ComplEx often outperforms both RESCAL and DistMult on standard benchmarks like FB15k and WN18, particularly for antisymmetric relations such as hypernymy and meronymy in WordNet.
Training and Optimization
These models are typically trained using negative sampling, where valid triples are contrasted with corrupted ones. The optimization objective minimizes a margin-based ranking loss:
where γ is a margin hyperparameter, 𝒯 is the set of valid triples, and 𝒯' contains negative samples generated by corrupting either the subject or object entity.

3.3 Path-Based and Rule-Based Approaches
Path-Based Reasoning
Path-based approaches leverage multi-hop relational paths between entities to infer missing links. Given a knowledge graph G = (E, R, T) where E represents entities, R relations, and T triples, these methods model the probability of a target triple (h, r, t) by aggregating information from all paths connecting h to t. The scoring function typically takes the form:
where P(h,t) denotes the set of paths between h and t, and σ is a sigmoid function. Path-ranking algorithms like PRA (Path Ranking Algorithm) learn weights for different path types through logistic regression, while neural variants such as DeepPath employ reinforcement learning to discover informative paths.
Rule-Based Inference
Rule-based methods utilize logical rules mined from the knowledge graph to perform completion. A Horn clause rule takes the form:
where the body consists of antecedent relations and the head is the consequent relation. Rule mining systems like AMIE+ employ confidence and support metrics to extract high-quality rules:
Neural theorem provers such as Neural-LP differentiable the rule application process, enabling gradient-based optimization of rule weights.
Hybrid Neuro-Symbolic Approaches
Recent work combines path-based and rule-based reasoning through neural-symbolic integration. Models like DRUM learn vector representations of rules while maintaining interpretability, with the scoring function:
where wR are learnable rule weights and countR(h,t) measures how many instantiations of rule R support the triple. Graph neural networks can simultaneously learn path representations and rule applications through message passing over the knowledge graph structure.
Practical Considerations
Path-based methods excel at capturing long-range dependencies but suffer from computational complexity in dense graphs. Rule-based approaches provide interpretability but require careful handling of noisy or incomplete data. Hybrid systems address these limitations by:
- Using neural networks to generalize beyond observed paths/rules
- Incorporating attention mechanisms to weight relevant evidence
- Jointly optimizing symbolic and subsymbolic components end-to-end
Applications include drug discovery (predicting protein interactions through biochemical pathways) and recommendation systems (inferring user preferences via behavior patterns). Current research focuses on scaling these approaches to billion-edge knowledge graphs while maintaining reasoning fidelity.

4. Graph Neural Networks for Knowledge Graph Completion
Graph Neural Networks for Knowledge Graph Completion
Graph Neural Networks (GNNs) have emerged as a powerful framework for knowledge graph completion by learning low-dimensional embeddings that capture both structural and semantic relationships. Unlike traditional embedding methods like TransE or ComplEx, GNNs leverage message-passing mechanisms to aggregate neighborhood information, enabling richer representations of entities and relations.
Message-Passing Framework
The core operation in GNNs is the message-passing step, where each node updates its representation by aggregating features from its neighbors. For a knowledge graph with entities E and relations R, the update rule at layer l can be expressed as:
where hv(l) is the embedding of entity v at layer l, W(l) is a learnable weight matrix, ruv is the relation-specific transformation, and σ is a non-linear activation function. The AGGREGATE function can be implemented as mean pooling, max pooling, or attention-weighted summation.
Relation-Aware Aggregation
Key variants like R-GCN (Relational Graph Convolutional Networks) extend this framework by introducing relation-specific transformations:
where cv,r is a normalization constant (e.g., |𝒩r(v)|), and Wr(l) are relation-specific weight matrices. This allows the model to distinguish between different relation types during aggregation.
Attention Mechanisms
Models like KB-GAT (Knowledge Graph Attention Networks) further refine this by computing attention weights αuv(r) for each neighbor:
where 𝐚 is a learnable attention vector and ∥ denotes concatenation. The attention mechanism dynamically prioritizes more relevant neighbors during aggregation.
Decoding and Scoring
After L GNN layers, the final embeddings are used to score triples (h, r, t) via a decoder such as DistMult or ConvKB. For example, the DistMult scoring function computes:
where d is the embedding dimension. The model is trained using margin-based or cross-entropy loss over positive and negative triples.
Practical Considerations
- Computational Efficiency: Sampling techniques like neighborhood sampling or graph partitioning are essential for scaling to large knowledge graphs.
- Relation Dependencies: Stacking multiple GNN layers captures higher-order relational dependencies but risks over-smoothing.
- Edge Features: Additional edge attributes (e.g., temporal information) can be incorporated via feature concatenation or dedicated edge-update mechanisms.

4.2 Convolutional and Attention-Based Models
Convolutional Models for Knowledge Graph Completion
Convolutional Neural Networks (CNNs) have been adapted for knowledge graph completion by treating entities and relations as embeddings and applying convolutional filters to capture local structural patterns. Given a triple (h, r, t), the embeddings h, r, and t are reshaped into 2D matrices and convolved with learnable filters. The score function for ConvE is defined as:
where σ is the sigmoid activation, * denotes convolution, ω represents the filters, and W is a linear transformation matrix. The overline notation indicates 2D reshaping of the embeddings. The model captures local interactions between entity and relation embeddings through the convolutional operation, enabling it to learn complex relational patterns.
Attention Mechanisms in Knowledge Graph Completion
Attention-based models, such as KBGAT, leverage graph attention networks to dynamically weigh the importance of neighboring nodes when aggregating information. For a given entity h, the attention coefficient αij between entities i and j is computed as:
where a is a learnable attention vector, W is a weight matrix, and ∥ denotes concatenation. The aggregated representation for entity i is then computed as a weighted sum of its neighbors' embeddings, enabling the model to focus on the most relevant connections in the knowledge graph.
Hybrid Convolutional-Attention Models
Recent advancements combine convolutional and attention mechanisms to leverage their complementary strengths. For instance, Conv-TransE integrates convolutional feature extraction with translational embedding learning. The score function is:
where CNN applies convolutional filters to the concatenated embeddings of h and r. The attention mechanism is then used to refine the convolutional features by focusing on the most informative dimensions. This hybrid approach achieves superior performance on benchmark datasets like FB15k-237 and WN18RR by capturing both local and global relational patterns.
Practical Applications
Convolutional and attention-based models excel in scenarios requiring fine-grained relational reasoning, such as biomedical knowledge graphs for drug discovery or recommendation systems in e-commerce. For example, in drug-drug interaction prediction, attention mechanisms can prioritize relevant biochemical pathways, while convolutional filters capture local structural similarities between molecular entities.

Transformer-Based Knowledge Graph Embeddings
Architectural Foundations
Transformer-based knowledge graph embeddings leverage the self-attention mechanism to model relationships between entities and relations in a knowledge graph. Unlike traditional embedding methods such as TransE or RotatE, which rely on fixed geometric transformations, transformers dynamically weigh the importance of different entities and relations through attention scores. The core architecture consists of:
- Multi-head attention layers that compute attention scores between entities and relations.
- Position-wise feed-forward networks that apply non-linear transformations to the embeddings.
- Layer normalization and residual connections to stabilize training.
The self-attention mechanism computes a weighted sum of all entities in the graph, where the weights are determined by the compatibility of queries and keys:
Here, Q, K, and V represent queries, keys, and values derived from entity and relation embeddings, and dk is the dimension of the key vectors.
Knowledge Graph-Specific Adaptations
Standard transformer architectures require modifications to handle knowledge graphs effectively. Key adaptations include:
- Relation-aware attention: Relations are incorporated into the attention mechanism by extending queries and keys to include relation embeddings.
- Structural embeddings: Positional encodings are replaced with structural embeddings that capture graph topology, such as random walk probabilities or graph neural network outputs.
- Edge-type masking: Attention between unrelated entities is masked to enforce sparsity and improve computational efficiency.
The relation-aware attention score between entity ei and ej with relation r is computed as:
where ⊕ denotes concatenation and Wq, Wk are learned projection matrices.
Training Objectives
Transformer-based knowledge graph models are typically trained using a combination of link prediction and contrastive learning objectives. The primary loss functions include:
- Negative sampling loss: Maximizes the score of true triples while minimizing the score of corrupted triples.
- Margin-based ranking loss: Enforces a margin between positive and negative triple scores.
- Reconstruction loss: Measures the accuracy of reconstructing masked entities or relations.
The margin-based ranking loss for a triple (h, r, t) is given by:
where γ is the margin, f is the scoring function, and (h', r, t') are negative samples.
Applications and Performance
Transformer-based models achieve state-of-the-art performance on knowledge graph completion benchmarks such as FB15k-237 and WN18RR. Key advantages include:
- Scalability: Efficient attention mechanisms enable processing of large-scale graphs.
- Expressiveness: Dynamic attention weights capture complex relational patterns.
- Transfer learning: Pre-trained transformers can be fine-tuned for downstream tasks.
Recent variants like KG-BERT and CoKE further enhance performance by incorporating pre-trained language models and contextualized embeddings.

5. Incorporating Temporal and Contextual Information
5.1 Incorporating Temporal and Contextual Information
Traditional knowledge graph completion models often treat relations as static, ignoring the dynamic nature of real-world knowledge. Temporal and contextual information significantly enhances the predictive power of these models by capturing how facts evolve over time or depend on specific conditions. This subsection explores advanced techniques for integrating such dynamic elements into knowledge graph embeddings.
Temporal Knowledge Graph Embeddings
Temporal knowledge graphs extend standard triples (h, r, t) to quadruples (h, r, t, τ), where τ represents a timestamp or time interval. The key challenge lies in modeling how relation semantics shift across time. Two dominant approaches are:
- Temporal Decomposition: Represents relations as time-dependent linear transformations. For a timestamp τ, the scoring function becomes:
where Wr,τ and rτ are learned temporal projections. TComplEx extends this by factorizing the temporal component in the complex space:
- Recurrent Event Modeling: Uses sequential architectures like Temporal Graph Networks (TGNs) to process knowledge graph snapshots. The entity representation h(k) at step k updates via:
Context-Aware Relation Learning
Contextual dependencies—such as spatial constraints or situational conditions—require modeling relation-specific contexts c. HypERContext employs hypernetworks to generate relation embeddings dynamically:
where Hr is a relation-specific hypernetwork. For multi-context scenarios, attention mechanisms weight relevant contexts:
Joint Temporal-Contextual Models
State-of-the-art approaches like TeLM (Temporal-Logical Memory) unify both dimensions through tensor factorization. The scoring function decomposes into temporal and contextual components:
HyTE projects entities and relations onto a time-specific hyperplane, while CENET uses contrastive learning to distinguish contextually valid triples from invalid ones. Evaluation on benchmarks like ICEWS18 shows a 12-15% MRR improvement over static baselines when incorporating both temporal and contextual signals.
Implementation Considerations
Efficient training requires:
- Time-window partitioning for temporal graphs to avoid memory overload
- Negative sampling strategies that account for temporal/contextual constraints
- Regularization terms to prevent overfitting on sparse temporal slices

5.2 Multi-Modal Knowledge Graph Completion
Traditional knowledge graph completion (KGC) models rely solely on structured triples (head entity, relation, tail entity), but multi-modal KGC integrates heterogeneous data sources such as text, images, and audio to enhance relational reasoning. This approach addresses the symbol grounding problem by anchoring abstract entities to real-world sensory data, improving both link prediction and entity alignment.
Architectural Components
Multi-modal KGC models typically consist of three core modules:
- Modality Encoders: Transform raw data (e.g., ResNet for images, BERT for text) into dense vector representations.
- Cross-Modal Alignment: A contrastive loss ensures consistency between structured knowledge and unstructured data embeddings.
- Joint Reasoning: Fuses multi-modal signals using attention mechanisms or tensor factorization.
where \(f_{kg}\) encodes KG entities, \(g_m}\) processes modality \(m\), and \(\mathcal{D}\) contains aligned entity-modality pairs.
Representative Models
MKGAT (Multi-modal Knowledge Graph Attention)
Uses graph attention networks to dynamically weight contributions from different modalities. The attention coefficient \(\alpha_{ij}\) between entity \(i\) and modality feature \(j\) is computed as:
TransMM
Extends TransE with modality-specific projections. The score function for a triple \((h,r,t)\) with associated image \(I_h\) becomes:
where \(\phi\) is a CNN encoder and \(\lambda\) controls modality alignment strength.
Training Paradigms
Two dominant strategies emerge:
- End-to-End Joint Training: All modalities are processed simultaneously with gradient updates propagating through the entire network.
- Pre-training + Fine-tuning: Modality encoders are pre-trained separately (e.g., vision models on ImageNet) before KGC-specific fine-tuning.
Evaluation Metrics
Beyond standard KGC metrics (MRR, Hits@K), multi-modal models require additional validation:
- Cross-modal Retrieval Accuracy: Measures whether image/text queries correctly retrieve corresponding KG entities.
- Modality Ablation Studies: Quantifies performance degradation when specific modalities are removed.
Applications
Multi-modal KGC enables novel applications such as:
- Visual Question Answering: Grounding questions about images (e.g., "What instrument is this?") to KG musical instrument hierarchies.
- Medical Diagnosis: Correlating radiology images with structured symptom-disease relations in biomedical KGs.
Current challenges include modality imbalance (some entities lack certain modalities) and computational complexity from processing high-dimensional sensory data.

5.3 Combining Symbolic and Neural Approaches
Recent advances in knowledge graph completion have demonstrated that hybrid architectures combining symbolic reasoning with neural networks outperform purely neural or purely symbolic approaches. These models leverage the complementary strengths of both paradigms: neural networks provide robust pattern recognition and generalization capabilities, while symbolic methods offer interpretability and precise logical constraints.
Architectural Paradigms
Three dominant architectures have emerged for combining symbolic and neural approaches:
- Neural-Symbolic Integration: Directly incorporates symbolic rules as differentiable constraints in neural architectures. For example, the loss function can be augmented with a term enforcing logical consistency:
where λ controls the trade-off between data fitting and rule satisfaction.
- Neural-Symbolic Cooperation: Uses neural networks to learn probabilistic representations that feed into symbolic reasoners. The neural component handles noisy data while the symbolic component performs exact inference.
- Neural-Symbolic Iteration: Alternates between neural prediction and symbolic refinement in a closed loop, with each component improving the other's output.
Differentiable Rule Injection
A key innovation is the development of differentiable implementations of first-order logic operators. For instance, the t-norm fuzzy logic provides a smooth approximation of logical conjunction:
These operators allow symbolic rules to be directly embedded in neural networks through fuzzy satisfiability. For a rule ∀x,y: r1(x,y) ⇒ r2(x,y), the satisfaction degree becomes:
where E is the set of entity pairs and r1(x,y), r2(x,y) are the neural network's prediction scores.
Case Study: Neural-LP
The Neural Logic Programming (Neural-LP) framework demonstrates this approach by implementing differentiable forward chaining. It represents Horn clauses as tensor operations:
where Pt is the predicate matrix at step t, Wt encodes the rule weights, and σ is a sigmoid activation. The model learns to compose rules through gradient descent while maintaining interpretable rule structures.
Performance Considerations
Hybrid models show particular advantages in:
- Data Efficiency: Symbolic constraints act as regularizers, reducing the need for large training sets
- Compositional Generalization: Ability to combine known rules in novel ways
- Few-shot Learning: Rapid adaptation to new relations with minimal examples
However, they introduce computational overhead from symbolic operations and require careful balancing between neural and symbolic components. The choice of architecture depends on the knowledge graph's characteristics - rule-based approaches excel in domains with clear ontological structures, while neural components handle noisy, incomplete data more effectively.

6. Popular Libraries and Frameworks for Knowledge Graph Completion
6.1 Popular Libraries and Frameworks for Knowledge Graph Completion
Library Selection Criteria
When evaluating libraries for knowledge graph completion (KGC), key considerations include scalability, support for heterogeneous graphs, ease of integration with deep learning frameworks, and availability of pre-trained models. Libraries optimized for sparse tensor operations and GPU acceleration are particularly valuable given the computational demands of KGC tasks.
PyKEEN (Python Knowledge Embeddings)
PyKEEN provides a unified interface for 50+ knowledge graph embedding models including TransE, RotatE, and ComplEx. Its modular design separates model implementation from training pipelines, enabling rapid experimentation. The library supports:
- Automatic hyperparameter optimization via Optuna integration
- Negative sampling strategies tailored for KGC
- Standardized evaluation protocols (MRR, Hits@k)
where h, r, t represent head entity, relation, and tail entity embeddings respectively in the TransE implementation.
AmpliGraph
Developed by Accenture Labs, AmpliGraph specializes in large-scale knowledge graph embeddings with TensorFlow backend. Notable features include:
- Distributed training on multi-GPU clusters
- Built-in support for relational graph convolutional networks (R-GCN)
- Active learning components for human-in-the-loop training
DGL-KE (Deep Graph Library Knowledge Embedding)
Built on the Deep Graph Library, DGL-KE optimizes for billion-scale graphs through:
- Partitioning algorithms for out-of-core training
- Mixed CPU-GPU execution pipelines
- Specialized kernels for sparse relation operations
GraphVite
This high-performance framework accelerates graph embedding tasks through:
- Parallel sampling with lock-free data structures
- SIMD-optimized scoring functions
- Approximate nearest neighbor search for efficient negative sampling
Comparative Performance
Benchmarks on FB15k-237 show varying throughput (triples processed/second):
- PyKEEN: 12k (single GPU)
- AmpliGraph: 18k (multi-GPU)
- DGL-KE: 42k (distributed CPU cluster)
Specialized Frameworks
KGNN (Knowledge Graph Neural Network)
Extends PyTorch Geometric for graph neural network-based KGC with:
- Differentiable rule injection
- Attention-based relation aggregation
- Temporal graph support
LibKGE
Research-focused library featuring:
- Memory-efficient batch processing
- Custom loss functions (e.g., self-adversarial negative sampling)
- Extensible sampling strategies
Integration with Deep Learning Ecosystems
Modern KGC libraries increasingly support interoperability with major ML frameworks:
- PyTorch Lightning integration in PyKEEN
- TensorFlow Serving deployment in AmpliGraph
- ONNX export capabilities in DGL-KE
6.2 Step-by-Step Implementation of a Basic Model
Model Architecture: Translational Embeddings (TransE)
TransE is a foundational knowledge graph completion model that represents entities and relations as vectors in a low-dimensional space. The core idea is that for a true triplet (h, r, t), the embedding of the tail entity t should be close to the embedding of the head entity h plus the relation vector r. The scoring function is:
where h, r, t are embeddings of the head, relation, and tail, respectively, and L2 norm enforces geometric proximity.
Step 1: Data Preparation
Load a standard dataset (e.g., FB15k-237 or WN18RR) with triplets (h, r, t). Preprocess the data by:
- Mapping entities and relations to unique integer IDs.
- Splitting into train/validation/test sets (e.g., 80%/10%/10%).
- Generating negative samples via corruption (replace h or t with random entities).
Step 2: Embedding Initialization
Initialize embeddings for entities and relations with dimensions d (typically 50–200). Use Xavier initialization:
where E represents entity or relation embeddings.
Step 3: Training Loop
Optimize the margin-based ranking loss:
where γ is the margin hyperparameter (e.g., 1.0), and T' contains negative samples. Use stochastic gradient descent (SGD) or Adam with learning rates between 0.001–0.01.
Key Hyperparameters
- Embedding dimension (d): Trade-off between expressiveness and computational cost.
- Margin (γ): Controls separation between positive and negative triplets.
- Batch size: Typically 128–1024 for stable gradients.
Step 4: Evaluation Metrics
Assess model performance using:
- Mean Rank (MR): Average rank of true entities in corrupted triplets.
- Hits@k: Percentage of true entities ranked in the top k (e.g., Hits@10).
Step 5: PyTorch Implementation
Below is a minimal TransE implementation in PyTorch:
import torch
import torch.nn as nn
import torch.optim as optim
class TransE(nn.Module):
def __init__(self, num_entities, num_relations, embed_dim, margin=1.0):
super(TransE, self).__init__()
self.embed_dim = embed_dim
self.margin = margin
self.entity_emb = nn.Embedding(num_entities, embed_dim)
self.relation_emb = nn.Embedding(num_relations, embed_dim)
nn.init.xavier_uniform_(self.entity_emb.weight)
nn.init.xavier_uniform_(self.relation_emb.weight)
def forward(self, h, r, t):
h_emb = self.entity_emb(h)
r_emb = self.relation_emb(r)
t_emb = self.entity_emb(t)
return -torch.norm(h_emb + r_emb - t_emb, p=2, dim=1)**2
def loss(self, pos_score, neg_score):
return torch.mean(torch.relu(self.margin - pos_score + neg_score))
Practical Considerations
- Normalization: Constrain entity embeddings to unit L2 norm to prevent scaling drift.
- Negative Sampling: Use Bernoulli sampling to balance head/tail corruption.
- Early Stopping: Monitor validation Hits@10 to avoid overfitting.

6.3 Optimizing and Scaling Knowledge Graph Models
Distributed Training for Large-Scale Knowledge Graphs
Training knowledge graph embedding models on large-scale graphs (e.g., Freebase, Wikidata) requires distributed optimization techniques to handle billions of triples efficiently. Two dominant approaches are:
- Parameter Server Architecture: Centralized servers maintain global embeddings while worker nodes compute gradients on graph partitions. The update rule for entity e at step t follows:
where K workers compute partial gradients gk with learning rate η. Synchronous updates ensure consistency but introduce communication overhead.
- Decentralized Training: Entities are partitioned across nodes with periodic synchronization. The consensus update for node i uses:
where wij are mixing weights and 𝒩(i) denotes neighboring partitions.
Negative Sampling Optimization
Traditional negative sampling uniformly corrupts triples, but adaptive strategies improve efficiency:
where α controls hardness weighting. Self-adversarial sampling (Sun et al., 2019) uses the current model's predictions to focus on challenging negatives.
Quantization and Pruning
Model compression techniques reduce memory footprint without significant accuracy loss:
- Quantization: Embeddings are stored as 8-bit integers with a shared 32-bit scaling factor per dimension:
- Structured Pruning: Removes entire embedding dimensions based on ℓ2-norm importance scores:
Hardware Acceleration
GPU/TPU optimizations exploit parallel computation patterns:
- Batched triple scoring with grouped matrix operations
- Kernel fusion for combining lookup, scoring, and loss computation
- Mixed-precision training (FP16/FP32) with loss scaling
Dynamic Graph Updates
For evolving knowledge graphs, incremental training strategies include:
- Warm-starting from previous embeddings
- Selective retraining of affected subgraphs
- Meta-learning approaches for fast adaptation
where β is the adaptation learning rate and 𝒟new contains new triples.

7. Key Research Papers and Surveys
7.1 Key Research Papers and Surveys
- Unifying Large Language Models and Knowledge Graphs: A Roadmap - arXiv.org — Although there are some surveys on knowledge-enhanced LLMs [47, 48, 49], ... Knowledge Graph Completion (KGC) refers to the task of inferring missing facts in a given knowledge graph. ... KGT5 introduces a novel KGC model that fulfils four key requirements of such models: scalability, quality, versatility, and simplicity. To address these ...
- PDF Knowledge Graph Refinement: A Survey of Approaches and Evaluation Methods — The Cyc knowledge graph is one of the oldest knowledge graphs, dating back to the 1980s [57]. Rooted in traditional artificial intelligence research, it is a curated knowledge graph, developed and main-tained by CyCorp Inc.7 OpenCyc is a reduced version of Cyc, which is publicly available. A Semantic Web
- (PDF) A Survey on Temporal Knowledge Graph Completion: Taxonomy ... — Index Terms —Knowledge Gr aphs, T emporal Knowledge Graphs, Knowledge Graph Completion, Interpolation, Extrapolation. 1 I NTRODUCTION K NO WLE DG E Graph (KGs) are structur ed multi-
- PDF LLM-based Multi-Level Knowledge Generation for Few-shot Knowledge Graph ... — 2.2 Multi-modal Knowledge Graph Completion Multi-modal Knowledge Graph Completion (MMKGC) mod-els primarily focus on incorporating visual information to aug-ment structural-only or text-only KGC tasks [Chen et al., 2024; Fang et al., 2022]. Recent MMKG-based models [Liang et al., 2022a; Zhang et al., 2023a] often process visual and structural
- Multi-perspective knowledge graph completion with global and ... — A knowledge graph (KG) is a knowledge base of a semantic network consisting of nodes, edges and attributes. Recently, technologies associated with the KGs domain have been rapidly developed and widely used for other downstream tasks, including question answering [1], [2], intelligent education [3], and recommendation systems [4], [5].According to the resource description framework (RDF ...
- Construction of Knowledge Graphs: Current State and Challenges - MDPI — With Knowledge Graphs (KGs) at the center of numerous applications such as recommender systems and question-answering, the need for generalized pipelines to construct and continuously update such KGs is increasing. While the individual steps that are necessary to create KGs from unstructured sources (e.g., text) and structured data sources (e.g., databases) are mostly well researched for their ...
- A survey on augmenting knowledge graphs (KGs) with large ... - Springer — Integrating Large Language Models (LLMs) with Knowledge Graphs (KGs) enhances the interpretability and performance of AI systems. This research comprehensively analyzes this integration, classifying approaches into three fundamental paradigms: KG-augmented LLMs, LLM-augmented KGs, and synergized frameworks. The evaluation examines each paradigm's methodology, strengths, drawbacks, and ...
- MHEC: One-shot relational learning of knowledge graphs completion based ... — The metric-based methods lack interpretability of results, while the algorithms based on path interaction are not suitable for few-shot scenarios and the availability of the model is limited in sparse knowledge graphs. In this paper, we propose a one-shot relational learning of knowledge graphs completion based on multi-hop information ...
- Towards Semantically Enriched Embeddings for Knowledge Graph Completion — If the work on knowledge graph completion is to go beyond simply data graph completion, the effort will need to focus on including the semantics of the KG. This can be as simply accounting for known relations between types such as the subsumption hierarchy or known disjointness relations between types to more sophisticated reasoning about the ...
- Towards electronic health record-based medical knowledge graph ... — Fig. 2 reveals that there has been a steady rise in the number of published articles in the knowledge graph domain that incorporated the use of EHR data. The number of articles shown in Fig. 2 are the articles selected for this literature survey using the search methodology mentioned in Section 3.The purpose of this paper is to review the literature on EHR-based medical knowledge graphs, which ...
7.2 Recommended Books and Online Courses
- CS224W: Machine Learning with Graphs - Stanford University — Towards Foundation Models for Knowledge Graph Reasoning ; InGram: Inductive Knowledge Graph Embedding via Relation Graphs ; Colab 5 out: Homework 3 due: Tue, 11/19: 16. Deep generative models for graphs : GraphRNN: Generating Realistic Graphs with Deep Auto-regressive Models ; Graph Convolutional Policy Network for Goal-Directed Molecular Graph ...
- PDF KnowledgeGraphs - content.e-bookshelf.de — area, defines the concept of a "knowledge graph", and provides a high-level overview of how knowledge graphs are currently being used. Chapter 2 presents and contrasts popular graph models that are commonly used to represent data as graphs, and the languages by which they
- PDF Knowledge Graph Refinement: A Survey of Approaches and Evaluation Methods — an existing knowledge graph and try to increase its coverage and/or correctness by various means. Since such works are reviewed in this survey, the focus of this survey is not knowledge graph construction, but knowledge graph refinement. For this survey, we view knowledge graph construc-tion as a construction from scratch, i.e., using a set of
- PDF Using Knowledge Graph for Explainable Recommendation of External ... — ected in the student model. 4 Building the Knowledge Graph We built a graph structure to represent the underlying knowledge layer of our system. The entities and relationships in this graph demonstrate the connection between the textbook content, Wikipedia and the student model. The knowledge graph is hosted on a native graph database (Neo4j ...
- A review: Knowledge reasoning over knowledge graph — Section 7 further explores the applications of such knowledge in downstream tasks, such as knowledge graph completion, question-answer systems, and recommendation systems. Section 8 discusses future research directions of knowledge graph reasoning. Finally, concluding remarks end the paper.
- Knowledge Graphs | The MIT Press - Ublish — The field of knowledge graphs, which allows us to model, process, and derive insights from complex real-world data, has emerged as an active and interdisciplinary area of artificial intelligence over the last decade, drawing on such fields as natural language processing, data mining, and the semantic web. ... III. KNOWLEDGE GRAPH COMPLETION (pg ...
- Knowledge graph construction for product designs from large CAD model ... — The use of domain specific text-driven semantic knowledge graphs has been explored for representing explicit and tacit design knowledge [46] and efficient process planning using a Process Knowledge Graph obtained from CAD/CAM systems [47]. Helping designers retrieve assemblies based on functional semantics was shown in the work by [48].
- Using Knowledge Graph for Explainable Recommendation of External ... — Rahdari et al. [3] used a graph database, Neo4j, to store the constructed knowledge graph that contains Wikipedia articles, textbook content, and the student model, and used Neo4j's internal full ...
- Meta concept recommendation based on knowledge graph — Massive Open Online Courses (MOOCs) are playing a key role in improving educational ways. Abundant learning resources make it difficult for online users to find suitable learning content. The current personalized service in the field of online education relies more on course recommendations. However, coarse-grained recommendations cannot help users discover the defects of their knowledge ...
- The construction of knowledge graphs based on associated ... - Springer — The full implementation of MOOCs in online education offers new opportunities for integrating multidisciplinary and comprehensive STEM education. It facilitates the alignment between online learning content and learning behaviors. However, it also presents new challenges, such as a high rate of STEM dropouts. Many learners struggle to establish effective learning behaviors and fail to complete ...
7.3 Open Datasets and Code Repositories
- PDF Graph-Based Entity Resolution and Completion for Academic Knowledge Graphs — 2.7 Graph-Based Techniques for Entity Resolution. . . . . . . . . . . . . .7 3 Data9 ... makes them particularly suitable for tasks such as graph completion, where the goal is to infer missing links or predict future ones based on observed data. In the con- ... institutional repositories, and open academic data sources such as OpenAlex.
- Path-based Explanation for Knowledge Graph Completion - arXiv.org — A graph distance metric based on the maximal common subgraph. Pattern recognition letters 19, 3-4 (1998), 255-259. Chang et al. (2023) Heng Chang, Jie Cai, and Jia Li. 2023. Knowledge Graph Completion with Counterfactual Augmentation. In Proceedings of the ACM Web Conference 2023. 2611-2620. Chang et al.
- PDF Open-World Taxonomy and Knowledge Graph Co-Learning - Emory University — Open-World Knowledge Bases. Existing KB completion models implicitly follow the closed-world assumption [Reiter,1981] in which all entities and relations have been observed and only missing links of known relations between existing entities can be discovered. Un-fortunately, closed-world KB completion models fail to adapt to new emerging ...
- PDF LLM-based Multi-Level Knowledge Generation for Few-shot Knowledge Graph ... — 2.2 Multi-modal Knowledge Graph Completion Multi-modal Knowledge Graph Completion (MMKGC) mod-els primarily focus on incorporating visual information to aug-ment structural-only or text-only KGC tasks [Chen et al., 2024; Fang et al., 2022]. Recent MMKG-based models [Liang et al., 2022a; Zhang et al., 2023a] often process visual and structural
- Fact Embedding through Difusion Model for Knowledge Graph Completion — ods in three benchmark datasets. Especially on FB15k-237, FDM achieves a 16.8% relative improvement in MRR scores compared to the state-of-the-art methods. ... explore the potential of diffusion models for knowledge graph completion tasks. •A learnable Fact Embedding Module is proposed to bridge the gap between continuous diffusion models and ...
- XLORE 3: A Large-Scale Multilingual Knowledge Graph from Heterogeneous ... — In knowledge representation and reasoning, a KG is a knowledge base that uses a graph-structured data model to capture, integrate, and manage data from diverse sources at large scale. 2 With the development of research, the resource description framework (RDF) and web ontology language were released and became important standards of KGs .
- (PDF) Revolutionizing Knowledge Graphs with Multi-Agent Systems AI ... — Linked Data Movement - Connecting open datasets to form a global knowledge . ... Reducing reliance on single data repositories, ... 3.4.1 AI-Powered Knowledge Graph Completion Models. ...
- Fine-Grained Evaluation of Rule- and Embedding-Based Systems for ... — Within our experiments we focussed mainly on the three datasets that have been extensively used to evaluate embedding-based models for knowledge graph completion: the WordNet dataset WN18 described in , the FB15k dataset, which is a subset of FreeBase, described in and FB15k-237, which has been designed in as a harder and more realistic variant ...
- Knowledge graph construction for product designs from large CAD model ... — The resulting Product Design Knowledge Graph and the data used for construction is released as a Linked Open Knowledge base for product design and manufacturing [13], [14]. Finally, use of the KG to provide efficient design recommendations from prior data is shown, along with its potential to use KGs for complex decision making.







