Skip to content
AI.info

Research

GazeFormer-MoE: Context-Aware Gaze Estimation via CLIP and MoE Transformer

Overview Research area: Computer vision, specifically 3D appearance-based gaze estimation using multimodal Transformers. Technical level: Advanced. The paper combines a pretrained CLIP vision encoder

GazeFormer-MoE: Context-Aware Gaze Estimation via CLIP and MoE Transformer
arXiv
2601.12316
Published
2026-01-18
Authors
Xinyuan Zhao, Xianrui Chen, Ahmad Chaddad

AI summary

Overview

  • Research area: Computer vision, specifically 3D appearance-based gaze estimation using multimodal Transformers.
  • Technical level: Advanced. The paper combines a pretrained CLIP vision encoder (ViT-B/32), a ResNet-50 CNN backbone, learnable semantic prototype banks, and a Mixture-of-Experts Transformer in a single architecture.
  • Scope (one sentence): The paper proposes GazeFormer-MoE, a single-frame model that conditions CLIP global features with learnable context prototypes and fuses them with CLIP patch tokens and high-resolution CNN tokens inside a routed/shared MoE Transformer, evaluated on four gaze benchmarks (MPIIFaceGaze, EYEDIAP, Gaze360, ETH-XGaze).

What This Paper Is About

Gaze estimation means recovering where a person is looking — a 3D line-of-sight vector — from a face image, and it is hard to do reliably when lighting, head pose, background, or a person's facial appearance change. Existing deep models either lack a way to modulate their features per sample using semantic context, or they fail to combine coarse global semantics with mid-level and fine-grained visual detail. GazeFormer-MoE addresses both gaps by injecting CLIP-aligned, learnable prototypical contexts into the global feature and by jointly attending prototype-enriched global vectors, CLIP patch tokens, and CNN tokens inside one Transformer encoder whose feed-forward layers are partly replaced by sparse and shared experts.

Key Contributions

  1. Method — semantics-modulated multi-scale pipeline. CLIP-aligned, learnable prototypes are injected into global features, and prototype-enriched global vectors are attended jointly with CLIP patch tokens and high-resolution CNN tokens inside a single Transformer encoder.
  2. Architecture — routed and shared MoE Transformer. The design combines specialized experts (route-dependent) for rare appearance sub-distributions with shared experts for base stability, increasing modeling capacity without uniformly increasing dense parameters.
  3. Evaluation — four benchmarks. Following the benchmark protocol of a recent gaze estimation review, the method reports new state-of-the-art angular errors on MPIIFaceGaze, EYEDIAP, Gaze360, and ETH-XGaze.
  4. Ablation evidence. Separate ablations isolate the contributions of feature combinations, the presence or absence of MoE, and learning-rate choice.

Main Findings

  • State-of-the-art angular errors on four benchmarks: 2.49° on MPIIFaceGaze (M), 3.22° on EYEDIAP (E), 10.16° on Gaze360 (G), and 1.44° on ETH-XGaze (Et), reported as up to a 64% relative improvement over previously reported results.
  • Comparison against prior best reported numbers: The paper states that on M and E it reduces the best previously reported errors (3.5° and 4.50°) to 2.49° and 3.22°, and that on G it improves the strongest prior result (10.34°) to 10.16° while preserving robustness at large yaw. Table 1 lists those prior values for GazeCLIP (3.50° on M), PCNet (4.50° on E), and MCA-PGI (10.34° on G).
  • Largest-gain claim: The paper states the largest absolute and relative gain appears on E (4.00° → 1.44°). Table 1 lists 4.00° for PCNet and 1.44° for this method on Et, not E.
  • Prototype conditioning alone is weak: In Table 2, using only the enriched vector f1 gives 7.66° (M), 10.25° (E), 28.43° (G), and 10.75° (Et); adding f2 gives 7.72°, 10.15°, 27.70°, and 10.40°, indicating coarse semantic priors without fine spatial structure under-express periocular micro-texture and shading.
  • High-resolution CNN tokens cause the largest error collapse: Adding T_cnn to f1 + f2 drops errors to 3.20° (M), 4.39° (E), 10.92° (G), and 1.66° (Et).
  • CLIP patch tokens give the final consistent refinement: Adding T_patch yields the full model's 2.49°, 3.22°, 10.16°, and 1.44°, improving M from 3.20° to 2.49° and E from 4.39° to 3.22°.
  • MoE routing matters: With MoE the errors are 2.49°, 3.22°, 10.16°, and 1.44°; without MoE they degrade to 4.20°, 5.78°, 10.72°, and 4.38°, with the largest degradations on E and M.
  • Learning rate is critical: In Table 4, 10⁻⁴ gives the best result on all four datasets (2.49°, 3.22°, 10.16°, 1.44°); 10⁻³ gives 8.46°, 11.30°, 10.90°, 4.66°; 10⁻⁵ gives 3.74°, 7.70°, 15.61°, 1.84°; 10⁻⁶ gives 5.39°, 9.81°, 15.89°, 3.70°; and 10⁻⁷ gives 5.79°, 10.37°, 17.38°, 4.23°.
  • Interpretation of gains: The discussion argues gains are not simple depth or parameter scaling, and that low-level texture, mid-level semantic structure, and prototype-guided context are complementary when co-attended in one sequence space rather than fused late.
  • Known limitations: Discrete argmax prototype selection with a learnable temperature sharpens routing but imposes finite granularity for continuous illumination gradients; a static vocabulary risks amplifying CLIP domain biases under sensor spectral shift (low-light color cast, infrared leakage); MoE brings routing variance, possible tail latency under unbalanced expert loads, and sensitivity to the load-balancing coefficient; and the formulation is strictly single-frame, ignoring temporal cues such as micro-saccades, blink dynamics, and head micro-motion.

Methodology in Plain English

The model takes a face image and extracts three kinds of features at once. A frozen CLIP vision encoder supplies a global image embedding plus a set of patch tokens; a trainable ResNet-50 supplies a high-resolution CNN feature map that is flattened into tokens.

Separately, the model keeps four learnable "prototype banks" — one each for illumination, head pose, background, and description/label. For a given image, the normalized global feature is compared against every prototype in a bank, a softmax with a learnable temperature produces similarity scores, and the single best-matching prototype per bank is selected (argmax). Two context-enriched vectors are then formed: one that adds the selected illumination, head-pose, and background prototypes to the global feature, and one that adds the selected label prototype. This gives each sample its own semantic prior without needing extra annotations.

Those two enriched vectors, the CLIP patch tokens, and the CNN tokens are concatenated into one sequence and fed to a Transformer encoder. In several feed-forward blocks, a gated Mixture-of-Experts layer replaces the dense feed-forward path: a router picks the top-K experts per token and their outputs are combined as a weighted sum, alongside a set of shared experts that always contribute base capacity. A load-balancing regularizer and weight decay are added.

The pooled Transformer output is mapped to a 3D gaze vector, and training uses an angular loss written as one minus the cosine similarity between prediction and ground truth — a form that is monotone with angular error and avoids computing arccos explicitly. Optimization uses AdamW with a learning rate annealed from 10⁻⁴ to 10⁻⁶ by cosine annealing, 100 epochs, and batch size 128, on an NVIDIA RTX 4090 with PyTorch 2.4.1+cu124.

Architecture specifics: CLIP ViT-B/32, ResNet-50, a 12-layer MoE Transformer with 8 heads, 512 dimension, and 2048 feed-forward dimension, Top 4 experts, 8 MoE experts (4 active, 4 shared, 1024 feed-forward), and a unified feature dimension of 768. The four benchmarks are split into training, validation, and test sets at an 8:1:1 ratio following the protocol of the cited gaze estimation review.

Why This Matters

Impact on research. The paper argues that semantic conditioning plus conditional (expert-routed) capacity is a concise recipe for robust fine-grained geometric estimation. Its ablations separate the effect of prototype conditioning, cross-scale token fusion, and MoE routing, which gives follow-up work a clearer picture of which ingredient produces which gain. It also extends a line of recent CLIP-based gaze estimation work (GazeCLIP, CLIP-DFENet) by adding per-sample prototype selection and sparse experts.

Real-world applications named in the paper:

  • Human-computer interaction, where gaze provides a non-intrusive input or attention signal.
  • Virtual and augmented reality, where headset-mounted cameras estimate gaze under changing illumination and pose.
  • Learning analytics, where gaze is used to infer attention.
  • General non-intrusive tracking of attention from images, the paper's framing of gaze estimation as recovering a 3D line of sight or 2D point of regard.

Industry relevance. Any product that needs gaze tracking without dedicated infrared hardware — AR/VR headsets, interaction research pipelines, and educational technology platforms — depends on robustness to lighting, head pose, and background shift, which is exactly the failure mode this paper targets. The paper reports gains specifically attributed to handling illumination and background diversity, and it releases code at https://github.com/AIPMLab/Gazeformer under a CC BY-NC-ND 4.0 license. The paper does not report latency, throughput, or memory measurements, so deployment cost is not quantified.

Future Directions

  1. Dynamic prototype evolution. The current prototype vocabulary is static and discrete; the authors propose making it evolve rather than leaving it fixed, which could better represent continuous illumination gradients.
  2. Temporally aware sparse experts. Because the current formulation is strictly single-frame, the authors suggest exploiting temporal coherence (micro-saccades, blink dynamics, head micro-motion) to regularize transient noise.
  3. Latency-oriented expert distillation. The authors name routing overhead — routing variance, potential tail latency under unbalanced expert loads, and sensitivity to the load-balancing coefficient — as a practical trade-off, and propose distilling experts to reduce latency.
  4. Mitigating CLIP domain bias. The paper raises the risk that a static prototype vocabulary amplifies CLIP domain biases under sensor spectral shift, such as low-light color cast or infrared leakage, which is left as an open question.

Target Audience

Researchers and graduate students working on gaze estimation, eye tracking, or multimodal vision-language models will get the most from this paper, particularly those interested in how CLIP features and Mixture-of-Experts routing can be adapted to a fine-grained geometric regression task. It is also relevant to engineers building gaze-aware interfaces for HCI, VR/AR, or learning analytics who need robustness across lighting, pose, and background conditions. It assumes familiarity with Transformer architectures, CLIP, and sparse MoE routing, so the methodology section is not aimed at beginners.

Authors’ abstract

We present a semantics modulated, multi scale Transformer for 3D gaze estimation. Our model conditions CLIP global features with learnable prototype banks (illumination, head pose, background, direction), fuses these prototype-enriched global vectors with CLIP patch tokens and high-resolution CNN tokens in a unified attention space, and replaces several FFN blocks with routed/shared Mixture of Experts to increase conditional capacity. Evaluated on MPIIFaceGaze, EYEDIAP, Gaze360 and ETH-XGaze, our model achieves new state of the art angular errors of 2.49°, 3.22°, 10.16°, and 1.44°, demonstrating up to a 64% relative improvement over previously reported results. ablations attribute gains to prototype conditioning, cross scale fusion, MoE and hyperparameter. Our code is publicly available at https://github. com/AIPMLab/Gazeformer.

Read the original paper