Research
AnyTouch 2: General Optical Tactile Representation Learning For Dynamic Tactile Perception
Overview Research area: Robotics — optical tactile sensing, tactile representation learning, and contact-rich manipulation. Technical level: Advanced. The paper assumes familiarity with masked autoenc
- arXiv
- 2602.09617
- Published
- 2026-02-10
- Authors
- Ruoxuan Feng, Yuxuan Zhou, Siyu Mei, Dongzhan Zhou, Pengwei Wang, Shaowei Cui, Bin Fang, Guocai Yao, Di Hu
AI summary
Overview
Research area: Robotics — optical tactile sensing, tactile representation learning, and contact-rich manipulation.
Technical level: Advanced. The paper assumes familiarity with masked autoencoders, contrastive multi-modal alignment (CLIP-style), video self-supervised learning, and robot policy learning (diffusion policy).
Scope: The paper introduces a five-tier "tactile dynamic pyramid" for organizing tactile data, releases a 2,426,174-sample hierarchical tactile dataset called ToucHD, and presents AnyTouch 2, a general optical tactile representation learning framework whose architecture mirrors the pyramid's tiers.
What This Paper Is About
Existing optical tactile datasets and pre-trained models focus on static, object-level properties such as material and texture, mostly collected through press-only or random-sliding actions, and therefore miss the fine-grained temporal deformations and contact forces that matter during real manipulation. The authors argue that advancing dynamic tactile perception needs a systematic hierarchy of perception capabilities that guides both data collection and model design. They build that hierarchy into a dataset (ToucHD) and a model (AnyTouch 2) that unifies object-level understanding with action-aware and force-aware dynamic perception across multiple sensors.
Key Contributions
-
A tactile dynamic pyramid. A five-tier framework ranking tactile data by the complexity of the dynamic perception it supports: (T5) Press Only, (T4) Random Action, (T3) Specific Action, (T2) Manipulation, and (T1) Force. The authors note most existing datasets sit in Tiers 4 and 5, while higher tiers remain scarce.
-
ToucHD, a large-scale hierarchical dynamic tactile dataset. 2,426,174 contact samples across three subsets: Sim (1,118,896 multi-sensor contact frames), Mani (584,842 contact frames from 46 manipulation tasks), and Force (722,436 touch–force pairs collected with 71 indenters). Sim and Mani are collected with five and two sets of optical tactile sensors respectively; the Force subset uses five tactile sensors.
-
AnyTouch 2, a general tactile representation learning framework. It combines masked video reconstruction, a new frame-difference reconstruction objective, multi-modal (tactile–visual and tactile–language) alignment, cross-sensor and action matching, and force plus delta-force prediction, trained with a curriculum task schedule where higher-level tasks start later with gradually increasing weights.
-
Evaluation across offline benchmarks and real-world manipulation. Benchmarks cover static object properties (TAG, Cloth), dynamic physical attributes (Sparsh, ToucHD Bench), and four real-world tasks spanning the pyramid tiers, across GelSight, DIGIT, and GelSight Mini sensors.
Main Findings
-
Object Bench results. AnyTouch 2 reaches 76.97% on TAG material classification and 42.31% on Cloth textile classification with GelSight. The paper describes this as comparable to AnyTouch 1 on Object Bench (80.82% on TAG; Cloth marked "Seen"), a benchmark emphasizing static semantic features. Removing multi-modal alignment drops TAG to 63.84% and Cloth to 37.61%, the largest object-level degradation in the ablation.
-
Dynamic and force benchmarks. AnyTouch 2 reports 86.66 F1 / 87.80 RMSE on Slip (DIGIT) and 97.96 F1 / 80.83 RMSE (GelSight Mini); 624.26 (DIGIT) and 202.14 (Mini) RMSE on Sparsh Force Prediction; and 894.32 (DIGIT) and 1051.03 (Mini) RMSE on ToucHD Bench Force Prediction. All RMSE values are in mN, and the paper states AnyTouch 2 consistently outperforms prior approaches on tasks requiring fine-grained dynamics and force-sensitive reasoning.
-
Multi-frame models beat single-frame models on dynamic tasks. Single-frame baselines (CLIP, UniTouch, T3) sometimes perform worse than CLIP on Force Prediction and Slip Detection, which the authors attribute to the absence of temporal position embeddings and thus an inability to capture tactile input ordering.
-
Data scale alone helps. Training MAE (Sparsh)† on the same data as AnyTouch 2, including ToucHD, improves results across all tasks even without additional objectives (for example, TAG rises from 59.47% to 63.32%, Cloth from 19.40% to 36.84%, and ToucHD Bench Force RMSE falls from 1953.82 to 1714.84 on DIGIT and from 3655.39 to 2467.42 on Mini). The authors present this as evidence of ToucHD's value as a high-tier dynamic dataset.
-
Ablation: each module matters for the capability it targets. Removing frame-difference reconstruction degrades all dynamic tasks; removing action matching hurts slip detection; removing force prediction hurts force and delta-force prediction. Removing the whole ToucHD dataset causes broad declines, as do the individual subsets for their corresponding tasks: dropping ToucHD (Sim) hurts the two slip tasks, and dropping ToucHD (Force) raises force RMSE substantially (for example, to 1792.49 on ToucHD Bench with DIGIT and 2424.68 with Mini).
-
A static-versus-dynamic trade-off. Removing multi-modal alignment improves most dynamic perception results (for example, Sparsh Force RMSE falls to 589.13 on DIGIT and 193.73 on Mini, and ToucHD Bench Force RMSE to 976.73 on Mini), while object-level performance drops sharply. The authors describe this as an inherent trade-off between static object properties and dynamic tactile features.
-
Real-world manipulation. Four tasks span the pyramid: Tactile Grasping (Tier 5), Whiteboard Wiping (Tiers 4 and 3), USB Insertion (Tier 2), and Chip Moving (Tier 1). With Diffusion Policy as the policy head and frozen tactile encoders, each task is tested 20 times and the average success rate is reported. The paper states that static single-frame models perform significantly worse than dynamic models on higher-tier tasks; models trained only on lower-tier data perform poorly on Tier 1 and Tier 2 tasks; MAE (S)† gains dynamic capability across Tier 2 and lower-tier tasks but not accurate Tier 1 force perception; and AnyTouch 2 outperforms all baselines across all four tasks, including Tier 1 Chip Moving. Numeric per-task success rates appear only in the figure, not as values in the text.
-
Sensor differences. GelSight Mini, with a cleaner background and sharper deformation imaging, outperforms DIGIT on the Tier-5 task with AnyTouch 2. DIGIT's higher acquisition frequency (30 Hz versus GelSight Mini's 18 Hz) provides more training samples and denser dynamic information, leading to better performance on higher-tier manipulation tasks.
-
Pre-training data composition. Contact samples were filtered from nine tactile datasets: Touch and Go (TAG), VisGel, ObjectFolder Real, TVL, YCB-Slide, SSVTP, Octopi, TacQuad, and ToucHD. Downstream evaluation used TAG and Cloth for object property understanding and Sparsh plus ToucHD Bench (10 unseen indenters, of which 3 are selected as testing indenters) for dynamic physical understanding.
Methodology in Plain English
The authors first define a ladder of tactile skills and use it to decide what data to gather. Press-only data is the easy bottom rung; force-paired data is the hard top rung. They then collect data at the top three rungs: simulated contacts of four atomic actions (sliding left, sliding right, rotating clockwise, rotating counterclockwise) on 1,043 objects from ObjectFolder and OmniObject3D using an IMPM-based simulator; real manipulation recordings made by modifying FastUMI so its two grippers carry different tactile sensors, across 46 tasks; and touch–force pairs generated by pressing 71 different indenters against five sensors while a wrist-mounted force sensor records 3D contact forces.
For the model, they start from a video masked autoencoder: tactile frames have the background frame subtracted, are split into 3D spatio-temporal tokens, partially masked, and reconstructed. A second decoder reconstructs frame differences (each frame minus the first frame) so the model becomes sensitive to subtle localized deformation rather than static appearance. On top of that, they add semantic objectives — contrastive alignment of tactile features with paired visual and textual features following the CLIP paradigm, plus cross-sensor matching that pulls together tactile videos of the same object captured by different sensors, and action matching that clusters the eight atomic actions (pressing, leaving, four sliding directions, two rotation directions) while separating different ones. Finally, force and delta-force decoders predict 3D contact force and frame-to-frame force change from the tactile video with an L1 loss, grounding the representation in physics. Training uses curriculum task scheduling: pixel-level reconstruction runs from the start with the highest weight, while alignment, matching, and force tasks switch on after task-specific start iterations with weights that ramp up linearly.
Why This Matters
For research, the paper reframes tactile pre-training around a hierarchy of dynamic capabilities rather than a flat collection of press-based datasets, and shows that the tier of training data a model sees determines which tier of task it can handle. The ablation also documents an explicit tension between static object semantics and dynamic force perception, which is a useful design constraint for future tactile foundation models.
Real-world applications:
- Dexterous manipulation and assembly — tasks such as USB insertion and precise chip moving require sensing slip and contact force, exactly the Tier 1 and Tier 2 capabilities the framework targets.
- Robotic grasping of unfamiliar objects — object-level identification and material classification support grasp and force planning.
- Multi-sensor robot hands — the cross-sensor matching objective and the inclusion of GelSight, DIGIT, and GelSight Mini aim at a representation that transfers across sensor hardware.
- Surface interaction tasks — wiping or cleaning a whiteboard-style surface depends on detecting transient friction and deformation changes as the tool moves.
Industry relevance: robot hardware vendors ship different tactile sensors with different imaging quality and frame rates, and the paper's cross-sensor design and its observation about DIGIT's 30 Hz versus GelSight Mini's 18 Hz suggest that a shared representation could reduce per-sensor re-training. Any large-scale dataset with paired touch and force signals is also directly useful for sim-to-real transfer and for learning force-aware control policies in factories and warehouses.
Future Directions
- Closing the static-versus-dynamic trade-off. The ablation shows that multi-modal alignment helps object properties but can hurt fine-grained dynamics; a principled way to get both without a trade-off is unresolved.
- Stronger Tier 1 force perception. Even with ToucHD, the authors note that models without the full framework still fail at accurate force perception for the Tier 1 Chip Moving task, and force RMSE on unseen indenters remains in the hundreds to thousands of mN.
- Broader sensor coverage. The evaluation spans three optical sensors (GelSight, DIGIT, GelSight Mini) with different imaging quality and frame rates; extending the hierarchy and representation to other tactile modalities is an open question.
- More diverse real-world manipulation data. The Mani subset covers 46 designed tasks; scaling to messier, longer-horizon, or multi-fingered manipulation would test whether the pyramid generalizes.
Target Audience
Robotics researchers working on tactile sensing and manipulation, self-supervised and multi-modal representation learning practitioners interested in non-visual modalities, and engineers building force-aware robot control or tactile sensor integration. Readers should be comfortable with masked autoencoders, contrastive learning, and imitation-learning policies such as diffusion policy to get the most from the method section.
Authors’ abstract
Real-world contact-rich manipulation demands robots to perceive temporal tactile feedback, capture subtle surface deformations, and reason about object properties as well as force dynamics. Although optical tactile sensors are uniquely capable of providing such rich information, existing tactile datasets and models remain limited. These resources primarily focus on object-level attributes (e.g., material) while largely overlooking fine-grained tactile temporal dynamics during physical interactions. We consider that advancing dynamic tactile perception requires a systematic hierarchy of dynamic perception capabilities to guide both data collection and model design. To address the lack of tactile data with rich dynamic information, we present ToucHD, a large-scale hierarchical tactile dataset spanning tactile atomic actions, real-world manipulations, and touch-force paired data. Beyond scale, ToucHD establishes a comprehensive tactile dynamic data ecosystem that explicitly supports hierarchical perception capabilities from the data perspective. Building on it, we propose AnyTouch 2, a general tactile representation learning framework for diverse optical tactile sensors that unifies object-level understanding with fine-grained, force-aware dynamic perception. The framework captures both pixel-level and action-specific deformations across frames, while explicitly modeling physical force dynamics, thereby learning multi-level dynamic perception capabilities from the model perspective. We evaluate our model on benchmarks that covers static object properties and dynamic physical attributes, as well as real-world manipulation tasks spanning multiple tiers of dynamic perception capabilities-from basic object-level understanding to force-aware dexterous manipulation. Experimental results demonstrate consistent and strong performance across sensors and tasks.