Skip to content
AI.info

Research

ViTaS: Visual Tactile Soft Fusion Contrastive Learning for Visuomotor Learning

Overview Research area: Robotics — visuomotor learning, multimodal (visual-tactile) representation learning for manipulation, spanning reinforcement learning (RL) and imitation learning (IL). Technica

ViTaS: Visual Tactile Soft Fusion Contrastive Learning for Visuomotor Learning
arXiv
2602.11643
Published
2026-02-12
Authors
Yufeng Tian, Shuiqi Cheng, Tianming Wei, Tianxing Zhou, Yuanhang Zhang, Zixian Liu, Qianwei Han, Zhecheng Yuan, Huazhe Xu

AI summary

Overview

Research area: Robotics — visuomotor learning, multimodal (visual-tactile) representation learning for manipulation, spanning reinforcement learning (RL) and imitation learning (IL).

Technical level: Advanced. The paper assumes familiarity with contrastive learning, conditional variational autoencoders, reinforcement learning objectives (PPO), and diffusion policy for imitation learning.

Scope: The paper proposes ViTaS, a visual-tactile fusion framework that replaces direct concatenation with a modified contrastive objective plus a CVAE reconstruction module, and evaluates it across 12 simulated tasks, auxiliary generalization variants, and 3 real-world robot tasks.

What This Paper Is About

Robots that rely only on cameras struggle when their view of an object is blocked by their own arm or when objects are transparent. Prior work that adds touch typically fuses vision and touch by simply concatenating their features, which ignores how the two senses relate and complement one another. ViTaS aims to learn a shared visual-tactile representation that exploits both the correspondence between the two modalities (via a "soft fusion contrastive" objective) and their complementarity (via a CVAE that reconstructs images from fused features), then uses that representation to drive manipulation policies.

Key Contributions

  1. A modified contrastive learning objective for cross-modal fusion. The authors extend the RGB single-modality framework of CoCLR to visual-tactile fusion, retrieving the top-K most similar visual samples to define positives for each tactile sample, then periodically swapping modality roles.
  2. The ViTaS framework. A representation learning paradigm combining soft fusion contrastive learning with a CVAE module that reconstructs visual observations from visuo-tactile embeddings, applied to guide visuomotor policy training in both RL and IL.
  3. A broad empirical evaluation. 12 simulated tasks across 5 distinct environments (plus 3 auxiliary generalization variants), 3 simulated imitation learning tasks, and 3 real-world tasks on a Galaxea-R1 humanoid robot, compared against 6 visuo-tactile representation baselines in simulation.
  4. Ablations and qualitative reconstruction analysis. Component-wise ablation (tactile input, unified encoder, soft fusion contrastive, CVAE, time contrastive, K values, loss weights) plus image reconstruction visualizations under noise and masking.

Main Findings

  • ViTaS leads on 11 of 12 reported simulation tasks, and dominates on average. In Table I, ViTaS achieves an average success rate of 91.4, ahead of CVT (71.5), MVT (70.3), VTT (47.5), M3L (46.6), Concat (40.6), and PoE (35.7). The single task where it is not best is Lift, where MVT scores 97.9 against ViTaS's 97.5.
  • The gap is largest on contact-rich rotation tasks. On Egg Rotate, ViTaS scores 85.7 versus 71.6 for CVT, 58.1 for M3L, and under 1.0 for PoE, VTT, and Concat. On Block Rotate, ViTaS scores 93.3 versus 70.4 for CVT and 69.1 for MVT. On Dual Arm Lift, ViTaS scores 100.0.
  • Stronger than Diffusion Policy in simulation, with less visual input. In Table II, DP with ViTaS averages 60.4 across Box Stack (53.3), Wiping (71.7), and Assembly (56.3), versus 38.2 for DP with a CNN encoder and 33.4 for DP with a Transformer encoder. Notably, ViTaS uses a single head camera plus tactile sensors, while the original DP receives inputs from 3 cameras.
  • Robustness to object geometry change, noise, and randomized targets. Under a Gaussian noise level of 0.3 in Insertion, ViTaS scores 89.2 versus 70.5 for CVT. On Pen Rotate with a randomized target, ViTaS drops to 78.4 but still leads MVT (47.5), CVT (51.3), and M3L (42.7); on the fixed-target version it scores 99.2.
  • Real-world advantage despite fewer cameras. In Table IV, ViTaS averages 46.0 across DAC (30.0), TPP-1 (42.0), TPP-3 (36.0), and FPP (76.0), against 30.0 for DP (20.0, 36.0, 24.0, 40.0) — an average improvement of 16%. ViTaS used only the head camera, whereas DP used all three cameras.
  • Tactile input is the single most important component. Ablating tactile information drops the average success rate by 32% (to 60.9, from 92.5). Ablating the soft fusion contrastive module yields 54.7, ablating CVAE yields 63.2, and replacing separate encoders with a unified encoder collapses the average to 27.1.
  • K = 10 beats both conventional and temporal contrastive alternatives. With K = 1 (standard cross-modal contrastive, same-timestep pairs only) the average is 78.8; with K = 20 it is 68.1; with K = 50 it is 67.3. Time contrastive, which uses temporally neighboring frames as positives, averages 70.6.
  • The chosen loss weights are validated. With λ = 1 and μ = 0.1 (the default), Egg Rotate reaches 85.7 and Block Rotate 93.3; every other tested setting of μ (0.01, 1, 10, 0.5, 0.05) and λ (0.1, 10, 100, 5, 0.5) scores lower on both tasks.
  • Reconstruction evidence for complementarity. In Figure 6, ViTaS reconstructs task-critical details such as the egg's location better than the token-based MAE in M3L, maintains reconstruction under higher noise, and performs better when heavy noise (noise level 0.5) is applied to one modality than to both. With core image regions masked, ViTaS still reconstructs observations using tactile and masked visual features.

Methodology in Plain English

ViTaS treats a trajectory as a sequence of paired visual and tactile observations. Two separate CNN encoders map each modality into its own latent space; the paper argues separate encoders matter because visual and tactile data differ fundamentally.

The first component is soft fusion contrastive learning. For a given tactile reading, the method looks up which images in the dataset are most visually similar to the image at that timestep, takes the top K of them (K = 10), and treats the tactile readings paired with those images as positives. Everything else is a negative. In this first stage only the tactile encoder is trained while the visual encoder is frozen; then the roles are swapped on a periodic schedule defined by a switching coefficient, so the visual encoder is updated too. This differs from standard cross-modal contrastive learning, which treats only the same-timestep pair as positive, and from temporal contrastive learning, which uses neighboring frames.

The second component is a conditional variational autoencoder. The concatenated visual-tactile feature is used as a condition to reconstruct the current camera image. The loss combines a reconstruction error term with a KL divergence term that pushes the latent distribution toward a standard normal. The idea is that forcing the model to rebuild vision from touch and vision together encourages the fused embedding to carry complementary information useful when vision is occluded. During training, the CVAE encoder, decoder, projector, and both modality encoders are optimized jointly; at inference only the two modality encoders are kept.

The final training objective adds the policy loss to these two representation losses, weighted by λ and μ. In RL the policy loss is PPO; in IL the policy loss is from Diffusion Policy.

For hardware, the authors use a 32×32×3 tactile map on a parallel gripper for several tasks, with channels 1 and 2 for normal force and channel 3 for shear force, and they augment dexterous hands with four 3×3×3 sensors per finger (proximal, middle, distal, tip) plus one on the palm, zero-padded to 32×32×3. In the real world they use 3D-ViTaC sensors producing 16×16×1 discrete haptic maps and adjust the tactile encoder accordingly. Real-world data consists of 100 expert trajectories per task collected via Meta Quest 3 teleoperation. Evaluation in RL runs 5 repetitions per environment with different random seeds and 3×10^6 timesteps; imitation learning uses 10^3 epochs with 50 expert demonstrations per task.

Why This Matters

Impact on research. The paper argues that alignment alone is insufficient and that complementarity must be modeled explicitly, and it supports that claim with ablations showing both the soft fusion contrastive module and the CVAE are necessary. It also provides a recipe for making a single head camera plus touch outperform a three-camera setup, which is a meaningful shift for embodied AI research that assumes rich vision is required.

Real-world applications:

  • Warehouse and logistics picking, where grippers occlude their own view of the object being grasped.
  • Household robots handling transparent or reflective items such as bottles and glassware, which are noted as a target scenario in the project framing.
  • Food handling and deformable or fragile object manipulation, where contact information matters more than appearance.
  • Refrigerator or shelf restocking, matching the Fridge Pick Place task, where the robot's own arm blocks the camera.

Industry relevance. The result that a cheaper sensing configuration (one camera plus tactile) can beat an expensive one (three cameras) has direct cost implications for robot fleet design. The reported real-world average success rate improvement of 16% over Diffusion Policy is measured on a humanoid platform (Galaxea-R1), so it speaks to deployability rather than simulation-only gains.

Future Directions

  • High-dynamic precise manipulation. The authors list real-world pen spinning as still challenging, citing physical capability limits and the bottleneck of RL in high-dimensional cases.
  • Deformable object manipulation. The authors state that the potential of visuo-tactile sensing for deformable objects warrants extra exploration.
  • Extending to more complex scenarios with advanced platforms. The conclusion commits to exploring fused visuo-tactile features in harder settings on both simulation and real-world platforms.
  • Open question on fusion scheduling. The switching schedule for alternating the two contrastive stages is controlled by a period T_switch, but the paper does not report the specific period value used, leaving how sensitive results are to that schedule an open detail.

Target Audience

Robotics and embodied AI researchers working on multimodal perception and manipulation policies, particularly those already familiar with contrastive representation learning, VAEs, PPO, and diffusion policy. It is also relevant to engineers building real-world robot systems who need robust performance under occlusion or with transparent objects, and to readers interested in how tactile sensing can substitute for additional cameras. Readers without a background in self-supervised representation learning or robot learning will find the method sections demanding.

Authors’ abstract

Tactile information plays a crucial role in human manipulation tasks and has recently garnered increasing attention in robotic manipulation. However, existing approaches mostly focus on the alignment of visual and tactile features and the integration mechanism tends to be direct concatenation. Consequently, they struggle to effectively cope with occluded scenarios due to neglecting the inherent complementary nature of both modalities and the alignment may not be exploited enough, limiting the potential of their real-world deployment. In this paper, we present ViTaS, a simple yet effective framework that incorporates both visual and tactile information to guide the behavior of an agent. We introduce Soft Fusion Contrastive Learning, an advanced version of conventional contrastive learning method and a CVAE module to utilize the alignment and complementarity within visuo-tactile representations. We demonstrate the effectiveness of our method in 12 simulated and 3 real-world environments, and our experiments show that ViTaS significantly outperforms existing baselines. Project page: https://skyrainwind.github.io/ViTaS/index.html.

Read the original paper