Skip to content
AI.info

Research

MemEIC: A Step Toward Continual and Compositional Knowledge Editing

MemEIC: A Step Toward Continual and Compositional Knowledge Editing Overview Research area: Multimodal knowledge editing for large vision-language models (LVLMs), sitting at the intersection of knowle

MemEIC: A Step Toward Continual and Compositional Knowledge Editing
arXiv
2510.25798
Published
2025-10-29
Authors
Jin Seong, Jiyun Park, Wencke Liermann, Hongseok Choi, Yoonji Nam, Hyun Kim, Soojong Lim, Namhoon Lee

AI summary

MemEIC: A Step Toward Continual and Compositional Knowledge Editing

Overview

Research area: Multimodal knowledge editing for large vision-language models (LVLMs), sitting at the intersection of knowledge editing, retrieval-augmented generation, and parameter-efficient fine-tuning.

Technical level: Advanced. The paper assumes familiarity with LoRA adapters, feed-forward-network-as-memory views of transformers, retrieval-based editing (SERAC/IKE style), catastrophic forgetting, and multimodal benchmarks.

Scope: The paper introduces a benchmark (CCKEB) and a method (MemEIC) for continually editing both visual and textual knowledge in LVLMs and for answering queries that require combining those edits.

Publication details: arXiv:2510.25798v1 [cs.LG], 29 October 2025, licensed CC BY-NC-ND 4.0. Authors are affiliated with the Electronics and Telecommunications Research Institute, POSTECH, and Sungkyunkwan University, Republic of Korea. A project page is listed at https://github.com/MemEIC/MemEIC.

What This Paper Is About

Large vision-language models store factual knowledge in both a textual and a visual pathway, but existing knowledge-editing methods typically modify only one modality at a time and are evaluated on a single edit rather than a long stream of edits. The authors argue this is unrealistic: real updates arrive incrementally, alternate between images and facts, and sometimes require a model to combine a visual edit with a textual edit to answer one question. The paper's goal is to define that harder problem, build a benchmark for it, and propose an editor that survives it.

Key Contributions

  1. CCKEB benchmark: The authors introduce what they describe as the first benchmark for continual and compositional multimodal knowledge editing, combining sequential visual edits with associated textual edits, along with a new metric, Compositional Reliability (CompRel), for evaluation.
  2. MemEIC framework: A multimodal editing framework that combines external memory retrieval with internal model editing using modality-specific adapters inspired by brain lateralization, joined by a brain-inspired Knowledge Connector.
  3. Modality-aware external memory (Mem-E): A retrieval memory split into two storage units, one for textual QA edits and one holding visual edits as (image, question, answer) triples, so that visual and textual cues are both used at retrieval time.
  4. Empirical results: Experiments on LLaVA-1.5 and MiniGPT-4 showing interference-free knowledge updates and improvements over prior methods on edit success and compositional reasoning.

Main Findings

  • Text-only retrieval is insufficient for visual edits: In the external-memory ablation on VLKEB under sequential editing with LLaVA-1.5, Mem-E with textual cues only reached 48.02 reliability and 4.02 image locality, while adding visual cues raised these to 96.51 reliability and 57.10 image locality. The SERAC baseline scored 75.00 reliability and 1.91 image locality.
  • Separating visual and textual memories helps, especially for locality: On the CCKEB validation set, Mem-I (Dual-LoRA, r:8 x 2) reached 80.93 on visual-edit T-Loc and 12.45 on I-Loc, versus 63.16 and 9.59 for Mem-I (Single-LoRA, r:16) — differences of -17.77% and -2.86% for the single-adapter version. For textual edits, the dual version led on every metric (Rel +2.70%, Gen +3.24%, Loc +5.06%). Experiments were repeated five times, and a paired t-test confirmed significance at p < 0.05.
  • Internal-memory baselines forget under long edit streams: Fine-tuning and LoRA achieve perfect visual and textual reliability at the 0-gap setting, but over long sequences of edits they overwrite prior knowledge, causing a drop of nearly 30 points in visual and textual reliability.
  • Retrieval alone caps compositional reasoning: Even with perfect oracle retrieval, compositional reliability leveled off at 64.93%, because retrieved facts sit alongside conflicting internal facts with no mechanism to reconcile them.
  • Separate adapters improve but drift over gaps: Dual-LoRA+RAG reached 78.16% compositional reliability at gap=0 but declined to 63.39% at gap=100. Without retrieval, Dual-LoRA scored 70.05% at gap=0, below Single-LoRA's 77.47%.
  • The Knowledge Connector stabilizes composition: Adding the attention-based Knowledge Connector to Dual-LoRA+RAG raised compositional reliability to 99.21% at gap=0 and maintained 97.01% at gap=100.
  • Headline comparison against baselines: Averaged across all edit–test gaps on LLaVA-1.5, MemEIC outperformed WISE by +16.94 in visual reliability and +32.35 in compositional reliability. Gap-averaged visual and textual reliability reached 98.93 and 92.48, and average compositional reliability reached 80.56 — a +18.51 improvement over the best baseline, LoRA (62.05).
  • External methods stay stable but weak at fusion: SERAC showed little difference between the 0 gap and 100 gap settings, but its text-only retrieval failed to locate visual edits reliably and its retrieved facts remained isolated, limiting compositional reliability.

Methodology in Plain English

The authors split the editing problem into four coordinated pieces.

Query decomposition. Every incoming question is automatically separated into its visual part and its textual part, which simplifies the query and lets each modality go down its own processing pipeline. The authors used GPT-4o for this decomposition.

External memory (Mem-E). Edits are stored outside the model in two separate stores: one for textual question–answer edits and one for visual edits kept as image, question, and answer triples. Textual retrieval uses DistilBERT [CLS] representations; visual retrieval combines a CLIP image encoder's [CLS] representation with a textual match using a weight of α = 0.5. For a compositional query, the model first retrieves the target entity using the visual sub-query, then substitutes that entity into the textual sub-query to retrieve more accurately from textual memory. Retrieved context is prepended to the query prompt.

Internal memory (Mem-I). Rather than embedding all edits into one shared parameter space, the authors attach two parallel LoRA adapters to the frozen feed-forward network — one for visual knowledge updates and one for textual knowledge updates — drawing an analogy to the left hemisphere's specialization for language and the right hemisphere's for visual and spatial information. Only the adapter matching the edit's modality is updated, so a visual edit cannot distort the textual representations and vice versa.

Knowledge Connector. Because separately maintained adapters can hold modality-biased token representations that interact poorly, the authors add LoRA adjustments to the self-attention query and key projections, gated by an indicator that turns on only when both adapters are active. For unimodal queries the connector defaults to an identity operation, preserving independent representations; for compositional queries it lets the two streams exchange information, in analogy to the corpus callosum.

Training and testing. Training has two stages. In stage one the base LVLM is frozen and only the external memory module is trained on the CCKEB training set. In stage two both adapters are active and the Knowledge Connector is trained using an adversarial retriever that supplies a mix of factual and counterfactual evidence, so the connector learns selective integration rather than over-reliance on external memory. At deployment the connector is frozen; new edits update the relevant adapter, and new evidence is simply appended to external memory. Testing performs 500 editing steps sequentially, with edit retention evaluated at gaps of 0, 10, 20, 50, and 100.

Why This Matters

Real deployments of vision-language models face information that keeps changing, and correcting one pathway while leaving the other stale produces inconsistent behavior. This paper matters because it reframes multimodal knowledge editing as a continual, compositional problem rather than a one-shot single-modality update, and because it shows that neither pure retrieval nor pure parameter editing is sufficient on its own.

Real-world applications named or implied by the paper:

  • Correcting critical visual recognition errors, such as a model misidentifying Donald Trump as Boris Johnson.
  • Updating outdated factual statements about an entity (for example, a person becoming a new president) at the same time as fixing a changed or mislabeled identity in a photo.
  • Keeping models current as real-world facts such as job titles, affiliations, and roles change over time.
  • Handling identity changes over time and correction of misidentified individuals in deployed image understanding systems.

Industry relevance: Any product that serves model answers about people, places, or events has to ingest updates without retraining and without corrupting earlier knowledge. The paper's finding that an internal-only editor loses roughly 30 points of reliability over a long edit stream, and that a connector-gated dual-adapter design keeps compositional reliability at 97.01% at gap=100, is directly relevant to teams that must schedule frequent incremental updates while preserving prior behavior.

Future Directions

  • Extending beyond paired editing. The authors deliberately limited experiments to a paired setting — a visual edit followed directly by its associated textual edit, then pairs of unrelated edits. The paper does not report results for other interleavings, leaving the ordering question open.
  • Scaling the evaluation. The provided content reports two backbone models (LLaVA-1.5 and MiniGPT-4) and 500 sequential editing steps. Whether the approach holds for other backbones, longer edit sequences, or larger knowledge bases is not reported.
  • Reducing dependence on GPT-4o and external knowledge graphs. Query decomposition and dataset construction rely on GPT-4o, with triples drawn from DB15K and FB15K and the MMKG knowledge graph; whether decomposition can be done internally is not reported.
  • Handling conflicting retrieval more broadly. The paper discusses counterfactual retrieval scenarios and states that its model avoids over-relying on incorrect external knowledge, but the paper's listed Limitations and Broader Impacts appendix is not included in the provided content, so its own stated limitations cannot be summarized here.

Target Audience

Researchers and engineers working on knowledge editing, retrieval-augmented generation, or multimodal model maintenance will benefit most. It is also relevant to practitioners who need to update deployed vision-language systems without retraining, and to benchmark designers interested in sequential evaluation and metrics such as CompRel. Readers should be comfortable with LoRA, transformer attention and feed-forward internals, and multimodal retrieval before approaching the method sections.

Authors’ abstract

The dynamic nature of information necessitates continuously updating large vision-language models (LVLMs). While recent knowledge editing techniques hint at promising directions, they often focus on editing a single modality (vision or language) in isolation. This prevalent practice neglects the inherent multimodality of LVLMs and the continuous nature of knowledge updates, potentially leading to suboptimal editing outcomes when considering the interplay between modalities and the need for ongoing knowledge refinement. To address these limitations, we propose MemEIC, a novel method for Continual and Compositional Knowledge Editing (CCKE) in LVLMs. MemEIC enables compositional editing of both visual and textual knowledge sequentially. Our approach employs a hybrid external-internal editor featuring a dual external memory for cross-modal evidence retrieval and dual LoRA adapters that facilitate disentangled parameter updates for each modality. A key component is a brain-inspired knowledge connector, activated selectively for compositional reasoning, that integrates information across different modalities. Experiments demonstrate that MemEIC significantly improves performance on complex multimodal questions and effectively preserves prior edits, setting a new benchmark for CCKE in LVLMs.

Read the original paper