Research
Forget Less by Learning from Parents Through Hierarchical Relationships
Overview Research area: Computer vision and generative modeling — specifically continual learning for personalized (custom) text-to-image diffusion models. Technical level: Advanced. The paper builds

- arXiv
- 2601.01892
- Published
- 2026-01-05
- Authors
- Arjun Ramesh Kaushik, Naresh Kumar Devulapally, Vishnu Suresh Lokhande, Nalini K. Ratha, Venu Govindaraju
AI summary
Overview
Research area: Computer vision and generative modeling — specifically continual learning for personalized (custom) text-to-image diffusion models.
Technical level: Advanced. The paper builds on diffusion model fine-tuning (LoRA, Custom Diffusion), continual learning, and hyperbolic geometry (Lorentz model, entailment cones), and its core equations assume familiarity with Riemannian manifolds and denoising objectives.
Scope: The paper proposes FLLP, a framework that models parent-child relationships between sequentially learned concepts in hyperbolic space in order to reduce catastrophic forgetting in custom diffusion models, and evaluates it on three public datasets plus one synthetic benchmark.
What This Paper Is About
Custom diffusion models can be personalized to generate new concepts, but when concepts arrive one after another, learning a new one tends to erase earlier ones — catastrophic forgetting. Existing methods try to minimize interference between concepts; this paper instead tries to exploit relationships between concepts, treating previously learned concepts as "parents" that guide new ones. The goal is a continual personalization framework that retains prior concepts while still absorbing new ones.
Key Contributions
- FLLP framework: The authors present Forget Less by Learning from Parents (FLLP), which they describe as the first framework to leverage inter-concept interactions positively — rather than only mitigating interference — by projecting concept embeddings into hyperbolic space (a Lorentzian manifold) where tree-like hierarchies arise naturally.
- Hierarchical Parent Entailment Loss: A loss that uses a union-find algorithm to form a parent-chain and enforces that a new (child) concept lies within the entailment cone of its predecessor (parent) concepts.
- Continual training of CDMs with this hierarchy: The framework is used to train custom diffusion models continually, with the paper reporting state-of-the-art results across three public datasets and one synthetic benchmark.
- A controlled synthetic diagnostic: A one-dimensional Gaussian diffusion benchmark with five sequential tasks (each a Gaussian with a distinct mean and unit variance), used to isolate and quantify forgetting mechanisms and to propose a "forgetting rate" metric.
Main Findings
- Synthetic 1D benchmark: The vanilla continual learning baseline reaches a forgetting rate of 18.6 and converges near the mean-of-means (9.6); CIDM (Dong et al. 2024) lowers this to 13.2, and FLLP lowers it further to 11.4 (lower is better).
- Two mechanisms named for forgetting: The paper attributes the 1D Gaussian forgetting to parameter interference (gradient updates for a new task displacing solutions for old tasks) and task-specific representation collapse (token embeddings for distinct means converging toward a shared subspace).
- CIFC dataset: FLLP's reported table averages are 80.0 IA / 76.1 TA versus CIDM's 78.0 IA / 74.8 TA (difference row: +2.0 IA, +1.3 TA). The prose states that FLLP improves on CIDM for 9 out of 10 concepts and reports an overall gain of +1.2 IA and 1.1 TA, with +9.2 IA on C7 (Dog2) and +3.2 TA on C9 (Cat2).
- CelebA dataset: FLLP surpasses CIDM by +4.4 in Image Alignment and +2.0 in Text Alignment (table averages 77.7 IA / 60.8 TA versus 73.3 IA / 58.8 TA). Person 5 (C5) shows the largest IA gain at +8.9; Person 8 (C8) shows the largest TA gain at +4.2.
- ImageNet dataset: FLLP outperforms CIDM by +1.1 IA and +0.5 TA (table averages 82.3 IA / 79.0 TA versus 81.2 IA / 78.5 TA), with C9 (Wood Rabbit) showing the largest IA gain of +4.4.
- Image attention maps beat LoRA weights for the entailment loss: Applying the hyperbolic parent entailment loss directly to LoRA weights helps IA but hurts TA (a "zero-sum game"); image attention maps (IAM) are therefore used instead.
- Scalability to 35 concepts: On CustomConcept101, FLLP outperforms CIDM by +2.1 on average IA and +1.0 on average TA. The paper notes 35 concepts is the ceiling imposed by CLIP's tokenizer (77 tokens, with 2–3 tokens per concept).
- Lower parameter drift: Measured by Frobenius norm averaged across all layers over 35 concepts on CustomConcept101, FLLP shows 22% lower parameter drift than CIDM.
- Threshold sensitivity: Across the three datasets, each concept has a distinct optimal value of the threshold β; lower values enforce stronger parental guidance, higher values relax the constraint. Each concept is trained with a different β.
- Parent chains are interpretable: In the qualitative example, FLLP learns the CIFC concept C7 (dog) using a parent chain of Cat1 (C3), Rubber Duck (C2), and Painting (C6) — i.e., C2, C3, C6, and C7 form a group.
- Non-replay compliance: Storing timestep-weighted mean image attention maps for prior concepts is argued not to violate the no-replay constraint, since that constraint restricts storage of prior reference images.
Methodology in Plain English
The setting is taken from CIFC: concepts arrive one at a time, each with a few reference images and a text prompt, and the model may only see the current task plus everything before it. There are three constraints — all concepts are distinct, concepts must be learned in arrival order with no access to future tasks, and no reference images from past tasks may be stored.
The model is a pretrained denoising UNet adapted with LoRA for each new customization task. The distinctive step is what happens in the latent space. Instead of keeping concept embeddings in ordinary Euclidean space, FLLP lifts them onto a hyperbolic hyperboloid (the Lorentz model), where distances grow exponentially and hierarchies fit naturally. Image features and previously learned token embeddings are averaged and projected onto this manifold.
To find a parent, the method performs a recursive search: starting from the new concept's embedding, it computes Lorentzian geodesic distances to all previous concepts and picks the second-closest one (skipping the trivial self-match), adds it to a parent chain, and repeats until it hits an already-visited node or the new concept itself (loop detection). This produces a chain of semantically relevant "parent" concepts.
Training then adds a penalty term: if a child concept's image attention map falls outside the entailment cone of its parent, the violation is penalized, scaled by a threshold β. This is combined with the CIDM-style objective, which shares a layer-wise common subspace across tasks via a learnable projection matrix, plus the standard diffusion denoising loss — giving a weighted total loss. At inference, the learned LoRA weights are combined using Elastic Weight Aggregation, weighting each stored concept by its semantic similarity to the current prompt embedding, so all personalized concepts are consolidated into one updated UNet.
Why This Matters
Impact on research. The paper reframes continual learning for generative models: instead of viewing concepts as independent units that must not interfere, it treats them as nodes in a hierarchy that can transfer knowledge. It also connects hyperbolic representation learning — largely used statically in vision-language models — to dynamic continual adaptation, which the authors note existing hyperbolic generative work (e.g., HypDiff) does not address.
Potential real-world applications (inferred from the personalization setting):
- Personalizing image generators for individual users, faces, or identities over time without re-learning from scratch.
- Creative and design tools that accumulate a user's styles and objects across sessions.
- Product and brand imagery pipelines that must keep earlier catalog concepts intact while adding new ones.
- Continual adaptation of generative models in settings where past training data cannot be retained (the no-replay constraint).
Industry relevance. Any deployment of custom diffusion models that ships updates or accepts user-specific concepts faces the forgetting problem this paper targets. The 22% lower parameter drift and the reported gains across CIFC, CelebA, and ImageNet matter to practitioners because they suggest more stable, less costly continual updates — though the paper's scope is bounded by the CLIP tokenizer limit of roughly 35 concepts.
Future Directions
- Breaking the tokenizer ceiling: The authors state that scalability is capped at roughly 35 concepts because of CLIP's 77-token limit (2–3 tokens per concept); extending beyond that is an open question.
- Automatic threshold selection: Each concept has a distinct optimal β across the three datasets, so deriving β per concept without per-dataset tuning is a natural next step.
- Richer and dynamic hierarchies: The parent chains are built from pairwise geodesic distances over averaged embeddings; whether hierarchies should adapt as more concepts arrive, or capture multi-parent structure beyond a single chain, remains unexplored.
- Generalizing beyond CDMs: The paper positions this as the first contribution to leverage inter-concept interactions in continual learning for CDMs, leaving open whether the same entailment-based mechanism transfers to other generative or discriminative continual settings.
Target Audience
Researchers and graduate students working on diffusion models, personalization, and continual learning will get the most from this paper. It is also relevant to practitioners building customizable generative systems who need to understand the trade-offs of hyperbolic embeddings and entailment-cone constraints, and to anyone interested in the intersection of hyperbolic representation learning and generative modeling. Readers without a background in Riemannian geometry will find the mathematical sections demanding, while the empirical tables and ablation studies remain accessible. Some appendix material (per-dataset β figures, architectural details, extended tables) is referenced but truncated in the version summarized here.
Authors’ abstract
Custom Diffusion Models (CDMs) offer impressive capabilities for personalization in generative modeling, yet they remain vulnerable to catastrophic forgetting when learning new concepts sequentially. Existing approaches primarily focus on minimizing interference between concepts, often neglecting the potential for positive inter-concept interactions. In this work, we present Forget Less by Learning from Parents (FLLP), a novel framework that introduces a parent-child inter-concept learning mechanism in hyperbolic space to mitigate forgetting. By embedding concept representations within a Lorentzian manifold, naturally suited to modeling tree-like hierarchies, we define parent-child relationships in which previously learned concepts serve as guidance for adapting to new ones. Our method not only preserves prior knowledge but also supports continual integration of new concepts. We validate FLLP on three public datasets and one synthetic benchmark, showing consistent improvements in both robustness and generalization.