Skip to content
AI.info

Research

FOVI: A biologically-inspired foveated interface for deep vision models

Overview Research area: Computer vision, with biological inspiration from primate retina and primary visual cortex (V1); covers foveated sensing, convolutional network design, and adaptation of large

FOVI: A biologically-inspired foveated interface for deep vision models
arXiv
2602.03766
Published
2026-02-03
Authors
Nicholas M. Blauch, George A. Alvarez, Talia Konkle

AI summary

Overview

Research area: Computer vision, with biological inspiration from primate retina and primary visual cortex (V1); covers foveated sensing, convolutional network design, and adaptation of large pre-trained vision transformers.

Technical level: Advanced. The paper assumes familiarity with convolutional networks, vision transformers, self-attention cost scaling, and low-rank adaptation (LoRA), and it uses cortical magnification functions from visual neuroscience.

Scope: The paper introduces FOVI (a foveated vision interface) that reformats a variable-resolution retina-like sensor array into a uniformly dense, V1-like manifold, shows how to perform convolution on that manifold, and evaluates it in both a custom CNN and an adapted DINOv3 ViT.

What This Paper Is About

Most computer vision systems encode images at uniform resolution, which makes full-field high-resolution processing expensive — especially for transformers, where cost grows doubly quadratically with image side length. Human vision instead samples with variable resolution that peaks at the center of gaze while covering a large (~180°) field of view. The paper builds a general-purpose foveated sensing interface that mimics this retino-cortical organization and shows it can be plugged into deep vision models to achieve competitive accuracy with far fewer sampled pixels and less computation.

Key Contributions

  1. A foveated vision interface (FOVI) modeled on the retino-cortical mapping. Visual space is sampled using the cortical magnification function M(r) = 1/(r + a) from Rovamo and Virsu (1984), producing sampling density that depends only on eccentricity, not polar angle. Cutting this manifold along the vertical meridian and flattening it yields what the authors describe as a strong first-order match to the retinotopic organization of human V1.

  2. A kNN-convolution and kernel mapping method for processing on the sensor manifold. Spatial receptive fields are defined as k-nearest-neighborhoods on the sensor manifold, and a novel kernel mapping technique transforms a reference kernel learned in a standard Cartesian grid into each neighborhood so that convolutional weights can be shared while preserving orientational alignment in visual space.

  3. A novel foveated convolutional neural network (FOVI-CNN) trained end-to-end. The authors vary the degree of foveation and show biologically plausible receptive field properties plus a performance advantage for intermediate foveation in image classification versus non-foveated controls.

  4. A foveated adaptation of a state-of-the-art pre-trained vision transformer. kNN-convolution implements foveated patch embedding inside an otherwise standard ViT, and LoRA-based fine-tuning functionally integrates the new sensor into two sizes of the DINOv3 ViT (ViT-S+ and ViT-H+).

Main Findings

  • Intermediate foveation performed best in ImageNet classification under a pixel budget. FOVI-CNNs were constrained to a 64x64 pixel budget sampled from a 256x256 ambient resolution, a 16-fold reduction in samples. At the maximum of 20 fixations, performance followed an inverted U-shaped function over the foveation parameter a, with peak performance at a = 0.5. Models with intermediate foveation outperformed the pixel-matched uniform model (a = 500).

  • FOVI-CNN receptive fields matched known primate properties. Receptive field diameter increased approximately linearly with eccentricity and then plateaued as padding units entered the field, and both the slope and intercept of that function increased with hierarchical layer. These dependencies were abolished in the non-foveated variant (a = 50), and the paper reports a qualitative match to human population receptive field sizes measured with fMRI (Dumoulin and Wandell, 2008).

  • FOVI-ViT-H+ approached full-resolution accuracy at roughly one-third the GFLOPs. Using 3976 pixels and 64 patches at 58.43 GFLOPs, FOVI-ViT-H+ @64 (a = 2.79) reached 0.844 accuracy with a single fixation, compared to 0.871 for the ViT-H+ uniform @ 224 baseline using 50176 pixels, 196 patches, and 172.39 GFLOPs. The paper states this is >96% of the baseline accuracy and describes it as state-of-the-art ImageNet-1K performance at a resolution of ≤4096 pixels (84%).

  • The smaller FOVI-ViT-S+ also beat matched low-resolution baselines. With a single fixation at 2.04 GFLOPs, FOVI-ViT-S+ @64 reached 0.700 accuracy, versus 0.693 for uniform @ 64, 0.693 for weak FOVI (a = 60.94), and 0.643 for log-polar @ 64. The paper reports it achieves 88.2% of the accuracy of the ViT-S+ uniform @ 224 baseline (0.794) in a single fixation. With 3 fixations, FOVI reached 0.735 versus 0.726 (uniform @ 64), 0.725 (weak FOVI), and 0.694 (log-polar).

  • LoRA over the patch embedding and first half of the network was the best adaptation strategy. On IN-100, this strategy outperformed frozen adaptation by approximately 30%, full fine-tuning by approximately 10%, and late-only adaptation by approximately 15%.

  • Higher-resolution reference kernels improved accuracy. Using a reference kernel of side length s = 2√k rather than s = √k improved ImageNet-1K accuracy by approximately 3%.

  • Attention costs dominate at high resolution. In ViT-S+, non-attention operations scaled approximately as O(m^1.76) empirically, while attention-based operations scaled as O(m^4), where m = √n is the pixels per side of a square image. For m < 400, non-attention costs outweighed attention costs; beyond that, attention costs became enormous.

  • Foveation can save orders of magnitude of compute at high resolution. Even with 20 fixations, processing at √n = 64 required two orders of magnitude fewer GFLOPs than a single pass at √n = 1024.

  • Efficiency gains were not uniform across metrics. The 1-fixation FOVI models showed significantly reduced GFLOPs, latency in both training and validation modes, and memory during training mode. In validation mode, forward-pass memory costs were largely outweighed by model size, and reductions in latency and memory were generally less pronounced than the reduction in FLOPs.

Methodology in Plain English

The authors start from a mathematical description of how densely human retina and cortex sample the visual field. A magnification function M(r) = 1/(r + a) controls how much resolution is devoted to the center versus the periphery: small values of a mean strong foveation, and as a grows the sampling becomes uniform. They integrate this function to get a "cortical distance," sample evenly along that dimension to obtain a set of eccentricities, and then choose the number of angular samples so that spacing between neighboring angles matches spacing between neighboring radii (local isotropy). The result is a set of sampling points that are evenly distributed on a curved manifold even though they are unevenly distributed in the original image.

Because that manifold is not a rectangular grid, standard convolution does not apply directly. Their solution is to define each output unit's receptive field as its k nearest neighbors on the manifold, then map a learned Cartesian kernel into each neighborhood using polar coordinates measured relative to that unit. This keeps filters orientationally aligned across the visual field — a vertical edge filter still detects vertical edges everywhere — while the same kernel is reused at every location on the manifold.

For the CNN experiments, they built an AlexNet-like model with five convolutional layers, three pooling layers, and two fully connected layers, with 96, 256, 384, 384, and 256 channels, global average pooling before the first fully connected layer, and two 1024-unit fully connected layers with ReLU and batch normalization. Training used ImageNet and a custom ImageNet-100 subset (100 random ImageNet categories, 500 training and 100 validation images per category), 100 epochs, a cosine-decay learning rate schedule, and no early stopping. Each model sampled 4 random fixations from a central region (radius 0.25 of image size in the main experiments, with 0.45 also tested) during training and up to 20 during validation, averaging logits across fixations.

For the transformer experiments, they replaced the standard patch embedding with a kNN-convolution-based foveated patch embedding in pre-trained DINOv3, targeting 64 patches with patch size 8, and adapted via LoRA over the patch embedding and first half of the network. They compared against four baselines: the original uniform model at 224 resolution, a weak-FOVI variant (a = 60.94), a uniform downsampled 64x64 variant, and a log-polar variant.

Why This Matters

Foveation offers a principled way to trade peripheral resolution for central detail, and this paper argues it is the first general-purpose foveated interface that is locally isotropic, biologically plausible, and compatible with diverse architectures — from CNNs to large pre-trained transformers. The paper positions foveation as a complement to patch-reduction techniques, noting that foveated perception also limits the number of pixels sampled rather than only the number of tokens processed.

Real-world applications the paper motivates or mentions:

  • Robotics and simulation rendering. Constraining ray tracing in environments such as IsaacLab to a foveated sensor array can reduce the compute needed to render each frame.
  • Self-driving cars and humanoid robotics. The introduction cites these as settings where processing at high resolution over a large field of view is critical.
  • Large field-of-view naturalistic scene processing. The future-directions section identifies high-resolution naturalistic scenes and high-resolution artificial or natural worlds as settings where foveation should help most.
  • Camera hardware. The paper suggests future sensors could directly sample in a foveated array, achieving high peak resolution with few sensor elements while reducing energy and bandwidth costs.
  • Computational modeling of human vision. FOVI is proposed as a tool for modeling peripheral phenomena such as crowding and how peripheral vision guides saccades.

Industry relevance is direct: the first author is employed by NVIDIA, which the paper states is developing extensions of FOVI, and LoRA adaptation of a pre-trained foundation model (DINOv3) is a practical route for vendors to make existing large models more efficient. Code and pre-trained models are released at github.com/nblauch/fovi and huggingface.co/fovi-pytorch.

Future Directions

  • Optimizing kNN convolution. The authors note that kNN convolution carries computational overhead relative to standard 2D convolution because each neighborhood has a different scattered layout, which is more problematic for CNNs (which pay the cost at every stage) than for ViTs (which pay it only at patch embedding). They plan more backend optimizations.
  • Extending beyond image recognition. Evaluations are limited to image recognition in a low-resolution setting with modest results; object detection and segmentation are not demonstrated, and the paper does not rigorously compare FOVI against non-foveated generic patch-reduction techniques.
  • Active vision and saccadic integration. Making foveation practical for real-world tasks requires advances in mechanisms for active vision and saccadic integration, motivated by the temporal, interactive, and embodied nature of robotic vision.
  • Sensing efficiency and hardware. Beyond perceptual efficiency, restricting sample counts could reduce rendering compute in simulation and enable camera hardware that samples natively in a foveated array.

Target Audience

Researchers and engineers working on efficient vision architectures, high-resolution or large-field-of-view perception, and vision foundation models, as well as computational neuroscientists interested in biologically grounded models of foveation and retinotopic organization. Readers need working knowledge of convolutional networks, vision transformers, and fine-tuning methods such as LoRA; the neuroscience content is presented in enough mathematical detail to be followed by those without a biology background, but the paper is not an introductory-level read.

Authors’ abstract

Human vision is foveated, with variable resolution peaking at the center of a large field of view; this reflects an efficient trade-off for active sensing, allowing eye-movements to bring different parts of the world into focus with other parts of the world in context. In contrast, most computer vision systems encode the visual world at a uniform resolution, raising challenges for processing full-field high-resolution images efficiently. We propose a foveated vision interface (FOVI) based on the human retina and primary visual cortex (V1), that reformats a variable-resolution retina-like sensor array into a uniformly dense, V1-like sensor manifold. Receptive fields are defined as k-nearest-neighborhoods (kNNs) on the sensor manifold, enabling kNN-convolution via a novel kernel mapping technique. We demonstrate two use cases: (1) an end-to-end kNN-convolutional architecture, and (2) a foveated adaptation of the DINOv3 ViT foundation model, leveraging low-rank adaptation (LoRA). These models provide competitive performance with a fraction of the pixels and computational cost of full resolution non-foveated baselines, opening pathways for efficient and scalable active sensing for high-resolution egocentric vision. Code (https://github.com/nblauch/fovi) and pre-trained models (https://huggingface.co/fovi-pytorch) are available.

Read the original paper