Research
VisionHOPE: Visual Backbones as Self-Modifying Learning Systems
VisionHOPE: Visual Backbones as Self-Modifying Learning Systems Overview Research area: Computer vision backbone architecture design, specifically recurrent visual operators that adapt during inferenc

- arXiv
- 2609.33325
- Published
- 2026-09-27
- Authors
- Siran Peng, Tianshuo Zhang, Tianyu Fu, Weisong Zhao, Haoyuan Zhang, Jiankuo Zhao, Minghui Wu, Ping Jiang, Xiangyu Zhu, Chenxu Zhao, Zhen Lei
AI summary
VisionHOPE: Visual Backbones as Self-Modifying Learning SystemsOverview
Research area: Computer vision backbone architecture design, specifically recurrent visual operators that adapt during inference (CNNs, Vision Transformers, State-Space Models, Test-Time Training layers) and Nested Learning theory.
Technical level: Advanced. The paper derives a proximal-problem formulation of gradient descent as associative memory, proves spectral-norm bounds on operator transitions, and relies on matrix-valued memory states.
Scope: The paper introduces a visual backbone whose content memory, key/value generators, learning rate, and retention factor all evolve within a single image, stabilized by a proven non-expansive step-size control, and validates it on classification, detection, instance segmentation, and semantic segmentation.
What This Paper Is About
Visual backbones have become progressively more input-adaptive, from CNNs with local kernels to SSMs with input-dependent state transitions and TTT layers that adapt an inner learner. In all of these, however, the rule that governs adaptation is fixed after training and never changes while the model processes an image.
VisionHOPE asks whether a backbone can change not just what it remembers but how it learns, by letting five coupled memories generate each other's update quantities as visual context accumulates along a scan.
Key Contributions
-
VisionHOPE as a self-modifying learning system. The first generic visual backbone formulated so that content, key, value, learning-rate, and retention memories co-evolve along visual scans. The backbone modifies both its stored content and its own learning rule.
-
Identification and removal of a stability barrier. The unconstrained self-referential update is shown to admit expansive transitions (the paper shows $\eta_t \lVert k_t \rVert_2 \lVert k_t + \delta_t \rVert_2 > 1 + \alpha_t$ is sufficient for expansion). The authors derive a stability-matched step-size control combining a soft cap on self-referential injection with a spectral clamp on the retained transition, and prove the resulting memory dynamics are non-expansive.
-
Adaptation of Nested Learning's chunk formulation to two-dimensional images. Chunks are aligned with image rows and columns across four directional scans (forward and reverse row-major, forward and reverse column-major), with independent states per direction and channel-wise fusion of directional outputs.
-
Instantiation in hierarchical and plain backbones with competitive results across model scales on ImageNet-1K, COCO, and ADE20K, plus an analysis of efficiency and design choices.
Main Findings
-
Image classification (hierarchical). At 224×224 on ImageNet-1K, VisionHOPE-T reaches 84.1% Top-1 with 27M parameters and 4.9G FLOPs; VisionHOPE-S reaches 85.2% (53M, 9.8G); VisionHOPE-B reaches 85.6% (91M, 17.3G). The authors state that across both backbone layouts VisionHOPE achieves the best or tied-best Top-1 accuracy in every scale panel. At the T scale, VisionHOPE-T (84.1) is level with RMT-S (84.1) and above H-ViT₃-T (84.0).
-
Image classification (plain/isotropic). P-VisionHOPE-T reaches 78.4% (6M, 1.2G), P-VisionHOPE-S reaches 82.3% (22M, 4.7G), and P-VisionHOPE-B reaches 83.4% (88M, 19.0G). These exceed the strongest listed TTT baselines at each scale (Vision-TTT-T 77.7, Vision-TTT-S 81.8, Vision-TTT-B 82.7).
-
Object detection and instance segmentation. With Mask R-CNN on COCO val2017 under the 1× schedule, VisionHOPE-T attains 47.9 box AP and 43.1 mask AP at 266G FLOPs; VisionHOPE-S attains 49.5 and 44.2 at 365G; VisionHOPE-B attains 50.5 and 45.0 at 516G. The paper reports best or tied-best box and mask AP across all three evaluated scales (VisionHOPE-B ties MILA-B at 50.5/45.0).
-
Semantic segmentation. With UPerNet on ADE20K at 512×2048, VisionHOPE-T reaches 49.4 mIoU (55M complete-model parameters, 942G FLOPs), VisionHOPE-S 50.3 mIoU (82M, 1043G), and VisionHOPE-B 51.8 mIoU (121M, 1200G). These are the best mIoU in each panel, ahead of H-ViT₃-T (48.0), H-ViT₃-S (50.2), and H-ViT₃-B (51.7).
-
Proven non-expansion. Proposition 1 establishes complementary operator bounds $\lVert A_t \rVert_2 \le \alpha_t$ and $\lVert B_t \rVert_2 < 1 - \alpha_t$ for the retained transition and the injection operator respectively. Corollary 1 extends this to $\lVert T_t \rVert_2 < 1$ and $\lVert M_t^{\square} \rVert_F \le \lVert M_{t-1}^{\square} \rVert_F$ for every memory and token. A separate argument in Appendix B.4 covers the chunk-wise case relative to each chunk's boundary states.
-
Parameter/FLOP trade-off versus TTT baselines. In the hierarchical panels, VisionHOPE uses fewer parameters than H-ViT₃ at all three scales (27 vs 29, 53 vs 54, 91 vs 94) but slightly more FLOPs at S and B (9.8 vs 8.8, 17.3 vs 16.7), with T matched at 4.9G. In the COCO and ADE20K hierarchical comparisons the same pattern appears (e.g., COCO B: 516G vs 510G; ADE20K S: 1043G vs 1026G). In the plain panels, P-VisionHOPE uses both fewer parameters and fewer FLOPs than the TTT baselines at every scale.
-
Efficiency analysis. The section compares per-layer interaction costs of DeiT, Vim, and VisionHOPE for a square feature map with $N$ tokens of width $D$. The detail of the comparison is not included in the available paper content.
Methodology in Plain English
The starting point is Nested Learning's description of a model as a stack of interconnected learning processes, each compressing its context into an internal state. Nested Learning casts plain gradient descent as writing into an associative memory, and introduces Delta Gradient Descent (DGD), which corrects each write using what the memory already returns for the current key. Adding a retention factor gives the retained DGD rule used here.
VisionHOPE keeps five memories per module: one matrix that stores content, two matrices that produce the key and value vectors, and two row vectors that produce a learning rate and a retention factor. The self-referential twist is that each memory builds its own regression target by transforming the value with its own current state, and then uses the resulting error as its learning signal. Because each memory generates quantities used in its own later updates, the stored content and the learning rule change together as the scan proceeds. A separate query projection, learned during normal training, reads the output from the content memory and stays fixed within the image.
Straight application of this update is unstable: the feedback loop can amplify memory norms. The authors first show analytically that an expansive transition is possible, then impose two controls. A soft cap limits how large a step the self-referential injection term may take, tapering toward a limit that shrinks as retention approaches one; a spectral clamp caps the step so the retained transition matrix has spectral norm at most the retention factor. Together these yield the two complementary bounds, hence non-expansive memory dynamics.
To handle 2D feature maps, the module is run four times per block over forward and reverse row-major and column-major scans. Chunks are aligned with rows for row-major scans (chunk length equals the width) and with columns for column-major scans (chunk length equals the height). Within a chunk, token-dependent quantities are computed in parallel from fixed boundary states, while state updates accumulate in token order; boundary states are refreshed between chunks. The four directional outputs are restored to the feature grid and merged with learned per-channel weights. The whole operator sits inside a pre-normalized residual block alongside input/output projections, a depthwise convolution, and a feed-forward network.
Why This Matters
For research, the paper makes a conceptual argument that has been building across CNNs, ViTs, SSMs, and TTT: adaptation granularity has been increasing, but the adaptation rule itself has stayed frozen. VisionHOPE pushes that argument one step further and, more importantly, shows that the resulting feedback loop is not merely unstable in principle but fixable with provable bounds. That combination — a self-referential formulation plus a proven non-expansion guarantee — is what distinguishes it from TTT-based backbones, which adapt an inner learner while leaving its representation maps and update rule largely outer-parameterized. The results also show the approach scales across hierarchical and plain layouts rather than fitting a single architecture niche.
Real-world applications, inferred from the tasks the paper evaluates (the paper itself does not enumerate deployments):
- Photo and video organization or content moderation, which depend on image classification accuracy at scale; VisionHOPE-T delivers 84.1% Top-1 at 27M parameters and 4.9G FLOPs, a size suitable for server-side pipelines.
- Autonomous driving and robotics perception, which rely on object detection and instance segmentation; VisionHOPE-B reaches 50.5 box AP and 45.0 mask AP on COCO.
- Scene parsing for AR, mapping, and satellite or aerial imagery, which depends on semantic segmentation; VisionHOPE-B reaches 51.8 mIoU on ADE20K.
- Medical and industrial inspection, where dense per-pixel labeling with a compact backbone matters and where the plain layout's low parameter and FLOP budgets (P-VisionHOPE-T at 6M parameters and 1.2G FLOPs) are attractive.
For industry, the practical profile is mixed but favorable. In hierarchical settings VisionHOPE edges out the TTT baselines on accuracy while using fewer parameters, at a small FLOPs premium. In plain isotropic settings it beats them on all three axes. The recurrent computation scales linearly with token count, which matters for high-resolution inputs where self-attention's quadratic cost becomes prohibitive.
Future Directions
-
Resolving the parameter/FLOP trade-off. In the hierarchical ImageNet, COCO, and ADE20K comparisons VisionHOPE consistently uses fewer parameters but slightly more FLOPs than H-ViT₃ at the S and B scales. Whether the operator can be made strictly dominant on both axes is open.
-
Replacing the linear memory recurrence. The paper notes that Nested Learning's HOPE architecture instantiates memories as MLPs with an $L_2$ regression objective, but that the corresponding recurrence is derived explicitly only for matrix-valued linear memories. VisionHOPE therefore adopts the linear recurrence. Extending the stability analysis to nonlinear memories is a natural next step.
-
Shrinking the performance gap in the plain layout. The plain panels are evaluated at smaller scales (6M, 22M, 88M parameters). Whether plain VisionHOPE retains its advantage at larger scales is not established by the reported results.
-
A fuller efficiency characterization. The efficiency analysis section compares per-layer interaction costs across DeiT, Vim, and VisionHOPE, but the outcomes of that comparison are not contained in the available paper content. The overhead of four directional scans, five memory states, and the step-size control relative to attention and standard SSM layers deserves a complete accounting.
Target Audience
Researchers and engineers working on visual backbone architecture, particularly those already familiar with Mamba-style selective state-space models, Vision Transformers, or Test-Time Training layers. Readers interested in Nested Learning or in the theory of associative memory and Delta Gradient Descent will find the stability analysis most relevant. Practitioners who need a backbone for detection, instance segmentation, or semantic segmentation will care about the transfer results, but should be prepared for the mathematical density of the methodology section. Beginners will need background in linear algebra, spectral norms, and proximal optimization to follow the derivations.
Authors’ abstract
Visual backbones have evolved from Convolutional Neural Networks (CNNs) with local aggregation to Vision Transformers (ViTs) with global interactions, State-Space Models (SSMs) with input-dependent state transitions, and Test-Time Training (TTT) layers that adapt an inner learner while processing an image. Across this progression, visual computation has become increasingly adaptive to each input, yet the rules governing that adaptation remain largely prescribed by the trained backbone. We introduce VisionHOPE, the first generic visual backbone formulated as a self-modifying learning system, in which what the model remembers and how it learns co-evolve within an image. Building on the self-referential construction of Nested Learning (NL), VisionHOPE realizes this co-evolution through five coupled memories that store content, generate key and value representations, and govern learning rate and retention. These memories evolve jointly as visual context accumulates along each scan. However, directly applying the unconstrained self-referential update to a visual backbone leads to instability. We therefore derive a stability-matched step-size control scheme that combines a soft cap on self-referential injection with a spectral clamp on the retained memory transition, and prove that the resulting memory dynamics are non-expansive along each scan. For two-dimensional feature maps, we adapt NL's chunk formulation by aligning chunks with image rows and columns across four directional scans. The proposed VisionHOPE achieves competitive results on ImageNet-1K, COCO, and ADE20K, establishing self-modifying learning systems as a practical foundation for general-purpose visual backbones. The code is available at https://github.com/PSRben/VisionHOPE.