Skip to content
AI.info

Research

Elastic Weight Consolidation for Knowledge Graph Continual Learning: An Empirical Evaluation

Overview Research area: Continual learning for knowledge graph embeddings (link prediction), with a focus on regularization-based methods and evaluation protocol design. Technical level: Intermediate.

Elastic Weight Consolidation for Knowledge Graph Continual Learning: An Empirical Evaluation
arXiv
2512.01890
Published
2025-12-01
Authors
Gaganpreet Jhajj, Fuhua Lin

AI summary

Overview

Research area: Continual learning for knowledge graph embeddings (link prediction), with a focus on regularization-based methods and evaluation protocol design.

Technical level: Intermediate. The paper assumes familiarity with knowledge graph embeddings and continual learning terminology, but its methodology and conclusions are explained clearly enough for readers with a general machine learning background.

Scope: A focused empirical evaluation of Elastic Weight Consolidation applied to TransE embeddings on the FB15k-237 benchmark across four sequential tasks under two different task-partitioning strategies.

What This Paper Is About

Knowledge graphs are constantly updated with new information, but the neural embedding models used to reason over them tend to forget previously learned facts when they are trained on new data sequentially — a problem known as catastrophic forgetting. This paper tests whether Elastic Weight Consolidation (EWC), a technique that protects the model parameters most important for old tasks, can reduce that forgetting when the model is TransE trained for link prediction on FB15k-237. It also investigates a second question the authors consider underaddressed: how much the measured forgetting depends on the way the training data is split into sequential tasks.

Key Contributions

  1. A focused empirical evaluation of EWC on knowledge graph link prediction using TransE embeddings on FB15k-237, across multiple regularization strengths and five random seeds, compared against naive sequential training and replay-based baselines.
  2. A direct comparison of two task-partitioning strategies — relation-based partitioning (grouping triples by relation type) and random partitioning — showing that task construction substantially changes measured forgetting (a 9.8 percentage-point difference under naive training).
  3. Evidence that EWC outperforms replay-based methods in this setting, including the finding that random replay (13.78% forgetting) performed worse than naive sequential training (12.62%).
  4. A hyperparameter sensitivity analysis showing that the optimal EWC regularization strength depends on how tasks are constructed, with λ = 10 best for relation-based partitioning and λ = 0.1 best for random partitioning.

Main Findings

  • EWC reduces forgetting on relation-based tasks: Naive sequential training produced 12.62% forgetting (std 0.35), while EWC with λ = 10 reduced this to 6.85% (std 0.33), a 45.7% reduction. Final MRR improved from 0.206 (std 0.006) to 0.242 (std 0.004).

  • Intermediate regularization strengths show a monotonic trend: On relation-based partitioning, EWC at λ = 0.1 gave 10.44% forgetting and at λ = 1.0 gave 7.51% forgetting, both worse than λ = 10 but better than naive training.

  • Replay-based methods underperformed: Random replay reached 13.78% forgetting (final MRR 0.196), worse than naive training, and wave replay reached 12.54% (final MRR 0.216). EWC combined with wave replay reached 9.91% forgetting (final MRR 0.234) — better than replay alone but worse than pure EWC.

  • Task partitioning strongly affects measured forgetting: Under naive training, relation-based partitioning yielded 12.62% forgetting versus 2.81% for random partitioning, a 9.8 percentage-point difference (Table 2 reports 9.81 pp). The authors attribute this to relation-based tasks creating distinct distribution shifts, while random partitioning creates relation-level overlap across tasks.

  • EWC narrows the partitioning gap: Under EWC, the difference between relation-based (6.85%) and random (5.08%) partitioning was only 1.77 percentage points.

  • Stronger regularization hurts on random partitioning: On random partitioning, naive training gave 2.81% forgetting, while EWC at λ = 0.1 gave 2.88%, λ = 1.0 gave 3.88%, and λ = 10 gave 5.08%. Final MRR on random partitioning was 0.277 for naive, 0.270 for λ = 0.1, 0.244 for λ = 1.0, and 0.222 for λ = 10.

  • Best performance-forgetting trade-off: The paper reports that EWC with λ = 10 occupies the optimal region of the performance-forgetting trade-off, while replay methods cluster in the high-forgetting, low-performance region.

Methodology in Plain English

The researchers took a standard knowledge graph benchmark, FB15k-237, which contains 14,505 entities, 237 relations, and 272,115 triples, and trained TransE embeddings (50 dimensions, margin γ = 1.0, L2 distance) using the Adam optimizer with a learning rate of 0.001, batch size 256, and 20 epochs per task.

Rather than training on all the data at once, they split it into four sequential tasks so the model had to learn one after another. They used two different splitting strategies:

  • Relation-based partitioning: They sorted the 237 relations by frequency and assigned them in round-robin order to four tasks, so each task got approximately 59 relations and all triples sharing a relation stayed together. This means each task is about a different slice of relation types.
  • Random partitioning: They shuffled all 272,115 training triples and divided them into four chunks of roughly 68,000 triples each, so most relation types appear in every task.

After training on task i, they measured Mean Reciprocal Rank (MRR) using filtered rankings, and after each subsequent task they recorded how much performance on earlier tasks dropped. Forgetting is defined as the difference between a task's MRR right after it was learned and its MRR at the end of all training, averaged across the earlier tasks.

EWC works by computing the Fisher Information Matrix diagonal, which estimates how important each parameter is for the previous task, using all triples from that task processed in mini-batches of 256. When training on a new task, a penalty term pulls important parameters back toward their previous values, with λ controlling how strong the pull is. They compared naive sequential training against EWC at λ values of 0.1, 1.0, and 10.0, EWC combined with wave-based experience replay (500 examples per task), and replay-only baselines. All experiments were run with five random seeds (42, 123, 456, 789, 2024), and results are reported as means and standard deviations.

Hardware was an NVIDIA RTX 3070 Ti with 8GB, with approximately 20 hours of total computation for 80 experiments (8 methods × 5 seeds × 2 partitioning strategies). Code was implemented in PyTorch 1.13. Training used uniform random tail corruption with a 1:1 ratio of positive to negative samples, and evaluation was filtered.

Why This Matters

Impact on research. The paper argues that continual learning for knowledge graph link prediction is underexplored relative to image classification and NLP, and that a classic regularization method like EWC has not received a focused empirical analysis in this setting with explicit attention to task partitioning. Its most distinctive claim is methodological: reported forgetting rates depend heavily on how tasks are constructed, so evaluation protocols need to state and justify their partitioning strategy. The paper frames this as a caution for how continual learning benchmarks are designed.

Real-world applications:

  • Knowledge graph-based AI agents: The authors frame preserving KG knowledge representations as essential for agents that rely on KG-based memory and reasoning over long autonomous operation.
  • Question answering and recommendation systems: Both are cited as applications built on knowledge graphs that evolve continuously as information changes.
  • Educational knowledge graphs: The authors propose that continual learning matters where knowledge evolves as curriculum content updates and student learning data accumulates.
  • Temporal and evolving KGs: Any setting where new facts arrive and existing knowledge is refined, matching the paper's motivating premise.

Industry relevance. Systems that maintain and update knowledge graphs in production face exactly the trade-off studied here: incorporating new facts without degrading performance on existing ones. The finding that replay alone underperformed and even matched or exceeded naive forgetting in one case suggests that simply keeping a limited buffer of old examples is not a reliable solution when memory is constrained, whereas a parameter-protection approach was more effective in these experiments. The task-partitioning result also matters operationally, since how incoming data is batched and sequenced into training rounds is a practical design decision that this paper shows can change measured outcomes by roughly 9.8 percentage points.

Future Directions

  1. Generalize beyond one model and dataset. The authors explicitly call for evaluating EWC across other embedding methods (RotatE, ComplEx, TuckER) and other datasets (WN18RR, YAGO, Wikidata), since the current study uses only TransE on FB15k-237.

  2. Scale to longer task sequences. The paper notes that scaling studies with 10 or more tasks would reveal the long-term dynamics of continual learning, which the current four-task setup cannot address.

  3. Formalize the relationship between task construction and forgetting. The authors propose systematic investigation of partitioning strategies (including entity-based and domain-based alternatives to their round-robin frequency assignment) to formalize how task boundaries affect measured forgetting.

  4. Educational knowledge graphs and neuromorphic hardware. The authors plan to test EWC on educational KGs updated with new learning resources, pedagogical relations, and student interaction data, and to explore whether spike-timing-dependent plasticity in spiking neural networks could offer a natural mechanism for consolidation, noting their preliminary consumer-GPU experiments on this were inconclusive and that specialized neuromorphic hardware (Intel Loihi, IBM TrueNorth, SpiNNaker) would be needed.

Target Audience

Researchers and practitioners working on continual learning, knowledge graph embeddings, and link prediction will get the most from this paper, particularly those designing continual learning benchmarks or evaluation protocols. It is also relevant to engineers building systems that must update knowledge graph representations over time, and to readers interested in how regularization-based methods like EWC transfer from image classification to structured knowledge representations. The paper assumes at least passing familiarity with KG embedding models such as TransE and with the concept of catastrophic forgetting.

Authors’ abstract

Knowledge graphs (KGs) require continual updates as new information emerges, but neural embedding models suffer from catastrophic forgetting when learning new tasks sequentially. We evaluate Elastic Weight Consolidation (EWC), a regularization-based continual learning method, on KG link prediction using TransE embeddings on FB15k-237. Across multiple experiments with five random seeds, we find that EWC reduces catastrophic forgetting from 12.62% to 6.85%, a 45.7% reduction compared to naive sequential training. We observe that the task partitioning strategy affects the magnitude of forgetting: relation-based partitioning (grouping triples by relation type) exhibits 9.8 percentage points higher forgetting than randomly partitioned tasks (12.62% vs 2.81%), suggesting that task construction influences evaluation outcomes. While focused on a single embedding model and dataset, our results demonstrate that EWC effectively mitigates catastrophic forgetting in KG continual learning and highlight the importance of evaluation protocol design.

Read the original paper