Research
Memory-Integrated Reconfigurable Adapters: A Unified Framework for Settings with Multiple Tasks
Memory-Integrated Reconfigurable Adapters (MIRA): A Unified Framework for Settings with Multiple Tasks Overview Research area: Machine learning — parameter-efficient adaptation, biologically inspired

- arXiv
- 2512.00940
- Published
- 2025-11-30
- Authors
- Susmit Agrawal, Krishn Vishwas Kher, Saksham Mittal, Swarnim Maheshwari, Vineeth N. Balasubramanian
AI summary
Memory-Integrated Reconfigurable Adapters (MIRA): A Unified Framework for Settings with Multiple TasksOverview
- Research area: Machine learning — parameter-efficient adaptation, biologically inspired associative memory, domain generalization (DG), and continual learning (CIL/DIL), with a Vision Transformer (ViT) backbone.
- Technical level: Advanced. The paper assumes familiarity with ViT architectures, LoRA-style adapters, Hopfield/associative memory formulations, and the CIL/DIL/DG evaluation protocols.
- Scope (one sentence): The paper proposes MIRA, a framework that stores per-task LoRA adapters inside Hopfield-style associative memories attached to every ViT layer and retrieves per-sample combinations of them at inference, handling domain generalization and continual learning under a single architecture with only changes to task-specific objectives.
Note: the provided paper text is truncated in Section 4 (it ends partway through "Impact of Adapter Count"), so the numbers in Table 7 and the ablation/conclusion material are not available in the content supplied.
What This Paper Is About
Deep learning has separate research communities for domain generalization (robustness to distribution shift), domain-incremental learning, and class-incremental learning, even though all three ask a model to adapt quickly to new tasks or environments without destroying earlier knowledge. Biological brains handle these challenges conjointly through a single neural circuitry modulated by signals such as dopamine and acetylcholine, and associative memory is thought to be central to that ability. The paper's goal is to build one architecture, MIRA, that uses associative memory to store and recall per-task adapter weights, so that a single frozen backbone can switch between tasks rapidly while retaining what it learned before.
Key Contributions
-
A unified framework. MIRA is presented as one architecture that serves domain generalization, class-incremental learning, and domain-incremental learning, differing only in the task-specific objective functions.
-
A refinement of Hopfield-based retrieval. The technical core is embedding associative memory (Universal Hopfield Network / Modern Hopfield Network style) units inside every ViT layer, and—rather than using static indexing keys—learning the retrieval keys and per-layer query modules post hoc so that keys align with preceding layer activations.
-
Storing weights, not data. Unlike prior associative-memory work that stores raw samples or feature representations, MIRA stores adapter-weight updates as the values in memory and retrieves affine combinations of them per sample at inference.
-
Broad empirical evaluation. MIRA is tested on standard CIL, DIL, and DG benchmarks, with the authors reporting state-of-the-art results in multiple settings and up to roughly 10% improvement over task-specialized architectures in some cases.
Main Findings
-
Continual learning results (Table 2). On iDigits, the paper's text highlights an average accuracy of 83% with forgetting of just 10.62% for MIRA; the table's CIL column lists 83.00 ± 1.29 average accuracy and 10.62 ± 2.80 forgetting, and the DIL column lists 82.46 ± 0.12 average accuracy and 8.49 ± 0.43 forgetting, against an unlabeled final column value of 82.73. The strongest comparison given is ICON at 71.53 ± 0.68 (CIL column), 19.36 ± 1.17 forgetting, 84.83 ± 0.51 (DIL column), 12.67 ± 0.61 forgetting, and 78.18 overall.
-
CORe50 and DomainNet. MIRA reports 83.39 ± 0.24 average accuracy with 7.99 ± 1.43 forgetting and 93.89 ± 0.33 average accuracy with 0.00 ± 0.00 forgetting on CORe50 (final column 88.64), and 67.29 ± 0.19 with 7.60 ± 1.06 forgetting and 69.18 ± 0.10 with 4.07 ± 0.15 forgetting on DomainNet (final column 68.24).
-
Domain generalization (Table 3). MIRA reaches 97.01 ± 0.0 on PACS, 82.10 ± 0.5 on VLCS, 87.36 ± 0.3 on OfficeHome, and 61.19 ± 0.1 on DomainNet, for an average of 81.92. It is best on three of the four datasets; PEGO holds the best VLCS score at 83.20 ± 0.3. Other baselines include CoOp (79.88 average), GESTUR (80.48), PEGO (80.30), CLIP (79.85), and SWAD (74.33).
-
Harder and longer benchmarks. On ImageNet-R in class-incremental settings, MIRA reports 78.06 ± 0.76 for the 5-task split and 73.08 ± 0.46 for the 10-task split, versus 81.14 ± 0.34 for the Joint upper bound, 58.74 ± 1.28 / 46.07 ± 1.15 for Sequential training, and 75.85 ± 0.31 / 71.89 ± 0.45 for C-LoRA (Table 5). On the CDDB DIL dataset, MIRA reports 77.37 ± 0.21 average accuracy against 74.51 for S-iPrompts, 61.28 for L2P, 60.94 for LwF, 51.27 for DyTox, and 50.59 for EWC, with a Joint upper bound of 85.50 (Table 6).
-
Surpassing replay-based methods without replay. On the DN4IL dataset (DIL with a 200-exemplar buffer), MIRA reports 78.40 ± 0.29 last accuracy, compared with 44.45 ± 0.18 for DUCA, 44.11 ± 0.98 for DARE++, 41.70 ± 1.41 for CLS-ER, 40.59 ± 0.73 for DARE, 35.74 ± 0.67 for DER++, and 27.45 ± 0.94 for ER. The paper states MIRA beats methods that use a 200-exemplar replay buffer without using exemplar replay itself.
-
Separation functions matter, and the affine choice wins overall. The paper compares affine, Softmax, ReLU, and Tanh separation functions (and mentions Identity and Polynomial used by other memory models). It reports that the ability to assign negative weights—and thereby actively remove interfering information—is key in CIL, where the affine and Tanh variants perform best, while negative coefficients may hurt in DIL; for DG, non-uniform selection of adapters plus removal of interfering information is what drives out-of-distribution generalization. The affine function achieves the best performance overall. The specific numbers in Table 7 are not included in the supplied text.
Methodology in Plain English
MIRA starts from a frozen Vision Transformer (ViT-B/16 initialized with CLIP weights, following PEGO), and attaches small rank-4 LoRA adapters to the Query and Value matrices of every layer. The trick is what happens to those adapters afterward.
Training has two stages. In the Adaptation stage, the system trains adapters for one task at a time using cross-entropy loss, then writes each layer's
Authors’ abstract
Organisms constantly pivot between tasks such as evading predators, foraging, traversing rugged terrain, and socializing, often within milliseconds. Remarkably, they preserve knowledge of once-learned environments sans catastrophic forgetting, a phenomenon neuroscientists hypothesize, is due to a singular neural circuitry dynamically overlayed by neuromodulatory agents such as dopamine and acetylcholine. In parallel, deep learning research addresses analogous challenges via domain generalization (DG) and continual learning (CL), yet these methods remain siloed, despite the brains ability to perform them seamlessly. In particular, prior work has not explored architectures involving associative memories (AMs), which are an integral part of biological systems, to jointly address these tasks. We propose Memory-Integrated Reconfigurable Adapters (MIRA), a unified framework that integrates Hopfield-style associative memory modules atop a shared backbone. Associative memory keys are learned post-hoc to index and retrieve an affine combination of stored adapter updates for any given task or domain on a per-sample basis. By varying only the task-specific objectives, we demonstrate that MIRA seamlessly accommodates domain shifts and sequential task exposures under one roof. Empirical evaluations on standard benchmarks confirm that our AM-augmented architecture significantly enhances adaptability and retention: in DG, MIRA achieves SoTA out-of-distribution accuracy, and in incremental learning settings, it outperforms architectures explicitly designed to handle catastrophic forgetting using generic CL algorithms. By unifying adapter-based modulation with biologically inspired associative memory, MIRA delivers rapid task switching and enduring knowledge retention in a single extensible architecture, charting a path toward more versatile and memory-augmented AI systems.