Skip to content
AI.info

Research

VisuoTactile 6D Pose Estimation of an In-Hand Object using Vision and Tactile Sensor Data

Overview Research area: Robotics perception — 6D object pose estimation using fused vision (RGB-D) and tactile sensor data, with synthetic dataset generation and sim-to-real transfer. Indexed by the a

VisuoTactile 6D Pose Estimation of an In-Hand Object using Vision and Tactile Sensor Data
arXiv
2601.01675
Published
2026-01-04
Authors
Snehal s. Dikhale, Karankumar Patel, Daksh Dhingra, Itoshi Naramura, Akinobu Hayashi, Soshi Iba, Nawid Jamali

AI summary

Overview

Research area: Robotics perception — 6D object pose estimation using fused vision (RGB-D) and tactile sensor data, with synthetic dataset generation and sim-to-real transfer. Indexed by the authors under Perception for Grasping and Manipulation, Deep Learning in Grasping and Manipulation, and Force and Tactile Sensing.

Technical level: Intermediate to Advanced. The paper assumes familiarity with convolutional networks, point-cloud representations, pixel-wise dense fusion, Siamese architectures, and pose/quaternion error metrics.

Scope: The paper proposes a tactile-sensor-agnostic representation and a two-channel visuotactile fusion network that estimates the 6D pose of an object grasped in a robot's hand, trained on a synthetically generated dataset of 11 YCB objects and deployed qualitatively on two real robot setups. (arXiv:2601.01675v1 [cs.RO], 04 Jan 2026.)

What This Paper Is About

When a robot holds an object, its own gripper and fingers block much of the camera's view, making it hard to tell exactly where the object is and how it is oriented. Most 6D pose estimation methods rely on vision (color and depth) alone and degrade under this self-occlusion. This paper adds tactile sensing at the fingertips — represented as a point cloud of the object surface touched by the fingers — and fuses it with vision to estimate the pose of the in-hand object, on the reasoning that the fingers that cause the occlusion are also the sensors that can see behind it in a different modality.

Key Contributions

  1. A tactile sensor invariant representation. Rather than using raw taxel readings (which vary wildly between sensor types, e.g., optical sensors that output RGB images versus randomly distributed pressure sensors), the authors represent tactile data as an object-surface point cloud at the finger-object contact location. Forward kinematics gives the 3D position of each contacting taxel. This makes the method independent of the tactile sensor type and the gripper type.
  2. A visuotactile network architecture. A two-channel network in which color and depth features are fused at the pixel level in a visual channel, while depth point-cloud features and tactile point-cloud features are fused at the point level in a tactile channel. A global feature fuses color, depth, and tactile embeddings to give the network global context. The pose estimator outputs a translation vector, a rotation vector, and a confidence level per feature, and selects the pose with maximum confidence.
  3. A synthetic visuotactile dataset generation method. The authors extended NVIDIA's Deep Learning Dataset Synthesizer (NDDS) in Unreal Engine 4 to produce photo-realistic color and depth images plus corresponding tactile object-surface point clouds and ground-truth 6D poses for 11 objects from the YCB Object and Model Set, with 20000 examples per object.
  4. Evidence that tactile data helps. Simulated experiments show the visuotactile network outperforms a vision-only baseline with statistical significance for all objects except the cracker box, and qualitative real-robot experiments show successful transfer of a synthetically trained network to two different physical setups.

Main Findings

  • Overall accuracy versus the vision-only baseline: For most objects the proposed network beats the baseline with statistical significance. On the adjustable wrench, position error is 0.36 cm versus 1.14 cm for the baseline. Other objects with large improvements include the scissors, potted meat can, and gelatin box. The cracker box is the one exception where the proposed method does not outperform the baseline.
  • Angular gains are smaller than positional gains: The spread between baseline and proposed method is less pronounced for angular error than for position error, which the authors attribute to the network relying on color features from the camera image to infer object orientation.
  • Heavy occlusion (above 85%): Position error for the proposed network is 0.4 ± 0.003 cm, roughly half the baseline's 0.78 ± 0.008 cm. Angular error is 11.5 ± 0.22 degrees versus 13.8 ± 0.24 degrees — a reduction of 2.3 degrees. Occlusion was binned into three levels: below 80%, between 80% and 85%, and above 85%.
  • Angular error still degrades under heavy occlusion: Although the proposed network outperforms the baseline, its angular error degrades substantially under heavy occlusion relative to its position error, suggesting texture is important for orientation inference (e.g., distinguishing the front from the back of a mustard bottle). Objects with large angular gains from tactile data, such as the adjustable wrench with uniform texture, depend heavily on shape captured by point-cloud data.
  • Fewer tactile contact points: The network was tested with 4, 12, 36, 60, and 240 tactile contact points. Reducing the number of contact points increases both position and angular error, but the tactile-augmented network still outperforms the vision-only baseline even with as few as 4 tactile points.
  • Ablation — Siamese network: Using the Siamese architecture gains 0.006 cm in position and 0.2 degrees in angular accuracy (Tactile-Channel Siamese: 0.299 ± 0.002 cm, 8.074 ± 0.105 degrees; Non-Siamese: 0.305 ± 0.002 cm, 8.269 ± 0.102 degrees).
  • Ablation — global feature: The global feature contributes a 0.57 cm gain in position and a 3.31 degree gain in angular accuracy (No Global Features: 0.876 ± 0.004 cm, 11.385 ± 0.123 degrees).
  • Ablation — visual data: Removing the visual channel (Tactile-Only) costs 1.55 cm in position accuracy and 51.19 degrees in angular accuracy (Tactile-Only: 1.849 ± 0.009 cm, 59.268 ± 0.282 degrees), consistent with the observation that orientation inference depends on color features.
  • Comparison with state of the art: PoseCNN trained on this dataset achieves 6.146 ± 0.023 cm position error and 10.896 ± 0.082 degrees angular error; PoseCNN+ICP achieves 6.158 ± 0.023 cm and 10.897 ± 0.083 degrees. The proposed Tactile-Channel network achieves 0.299 ± 0.002 cm and 8.074 ± 0.105 degrees. The authors note a 4.78 degree increase in angular error for the comparison method relative to theirs, and interpret the results as evidence that the dataset presents a challenging pose estimation problem.
  • Real-robot deployment (Setup-I): For the data presented in the accompanying video, the paper reports 36.88 ± 1.27 degrees and 2.69 ± 0.13 cm for the proposed network, compared to 64.57 ± 1.40 degrees and 4.76 ± 0.32 cm for the baseline. The proposed network produced a more stable frame-by-frame output, and the authors suggest it uses tactile information from fingers at the back of the object to improve the estimate. Under gradually increased occlusion produced by blocking the main camera with a piece of paper, the proposed network handled higher occlusion levels than the baseline.
  • Inference speed: Network inference alone takes 10.1 ms; the entire pipeline including the semantic segmentation network and ROS overhead takes 109.5 ms in deployment on a workstation with an NVIDIA Quadro RTX 6000.

Methodology in Plain English

The authors take three inputs: a color image and a depth image from a camera, plus a point cloud representing the object surface that the tactile sensors are touching. The tactile point cloud is built by using forward kinematics to place each contacting sensor element at its 3D location, which works whether the underlying sensor is optical or pressure based and regardless of how many fingers the gripper has.

The network splits the work into two channels. In the visual channel, the color image is semantically segmented to isolate the object, then cropped and passed through a CNN to produce per-pixel color embeddings; the depth image is masked by the same segmentation and turned into a 3D point cloud in the camera frame, then fed through its own CNN. The two are combined with pixel-wise dense fusion, a technique borrowed from prior work. In the tactile channel, the depth-derived point cloud and the tactile point cloud are fused at the point level. A third global feature combines color, depth, and tactile embeddings for scene-wide context.

A key detail is that only 1000 randomly selected embeddings and their corresponding 3D points are kept per channel, so the network focuses only on points that lie on the object. The pose estimator then emits up to 1000 pose estimates per channel, each with a confidence level, and the highest-confidence estimate wins. Training uses a loss that combines a point-wise pose error with a confidence loss whose expected confidence is 0.5 raised to the power of the point-wise loss.

Training data comes from simulation. In Unreal Engine 4, the authors place a 6-DoF robotic arm with a 4-fingered Allegro Hand equipped with tactile sensors, and attach a virtual camera to each phalanx with its origin on the phalanx surface and its z-axis pointing outward. Those virtual cameras capture the object's surface point cloud at contact. A separate main camera in front of the robot generates the color and depth images. Data collection cycles through 200 random trajectories, placing the object randomly, setting random orientations and a random gripper preshape, grasping, then recording 100 instances per grasp while moving the arm. For variation, 20% of the dataset uses no domain randomization and 80% is domain randomized using NDDS's Unreal Engine 4 plugin, randomizing background, floor surface, and a spotlight's position and color.

The dataset uses 11 YCB objects chosen for graspability and 3D model quality, with 20000 samples per object, 12 virtual tactile cameras per hand, 32x32 resolution for the finger cameras (proportional to the tactile sensor shape), and 640x480 for the main camera. Training used a 4:1 split (16000 training, 4000 test per object), 1500 epochs, batch size 16, learning rate 0.0002, hyper-parameter w of 0.015, and the Adam optimizer in PyTorch. Deployment used two real setups: Setup-I with a Weiss gripper and two GelSight tactile sensors on a Sawyer robot, a Kinect2 RGB-D camera mounted 4 m from the robot base at 45 degrees facing the workspace, and an OptiTrack motion capture system for ground truth; and Setup-II with an Intel RealSense camera and a 4-fingered Honda R&D hand with 16 joints (11 actuated) and a tactile sensor suite of 224 taxels.

Why This Matters

This paper is distinctive in extending tactile-vision fusion from shape estimation to full 6D pose estimation for in-hand manipulation, and in pairing that with a scalable way to generate synthetic tactile ground truth — a step that sidesteps the cost and wear-and-tear of collecting large real manipulation datasets.

Impact on research: The work offers a sensor-agnostic tactile representation (contact-surface point clouds) that other groups can reuse regardless of their tactile hardware, and treats synthetic visuotactile generation as a first-class research problem rather than an afterthought. It also identifies a concrete limitation, the dependence of orientation inference on color features, that frames follow-up work.

Real-world applications:

  • Robotic grasping and manipulation, where knowing an object's pose in the gripper is a prerequisite for placing, reorienting, or assembling it.
  • Assembly and industrial handling tasks where parts are gripped and must be positioned precisely under partial camera occlusion.
  • Prosthetic and multi-fingered robot hands with tactile arrays, which could use the same fused pipeline to become aware of object pose in hand.
  • Virtual and augmented reality, where the paper notes accurate 6D pose is useful, and planning systems that need a reliable object state estimate before acting.

Industry relevance: The work comes from Honda Research Institute USA with co-authors from Honda R&D Co., Ltd., and is validated on two different industrial-grade physical setups (a Sawyer with a Weiss gripper and GelSight sensors, and a proprietary Honda multi-fingered hand with 224 taxels), which indicates the approach is being evaluated for real hardware rather than only in simulation.

Future Directions

  • Temporal filtering: Use long short-term memory networks to exploit temporal coherence and filter out erroneous pose estimates.
  • Gripper-aware reasoning: Use the geometry of the gripper and the positions of the fingers to help the network reject physically implausible pose estimates.
  • Asymmetric objects: Extend the dataset with 3D printed asymmetrical objects, since most objects in the current dataset are symmetrical on at least one axis, and evaluate the method on them.
  • Environmental occlusion study: The authors observed their method is robust to environmental occlusions and propose a dedicated study testing it against different levels of environmental occlusion.

Target Audience

Robotics researchers working on pose estimation, grasping, and manipulation; engineers building tactile-sensing robot hands and grippers; and practitioners interested in synthetic data generation and sim-to-real transfer for manipulation. Readers with a background in deep learning for perception will get the most from the network architecture and ablation sections, while those focused on hardware will find the tactile representation and the two deployment setups most relevant.

Authors’ abstract

Knowledge of the 6D pose of an object can benefit in-hand object manipulation. In-hand 6D object pose estimation is challenging because of heavy occlusion produced by the robot's grippers, which can have an adverse effect on methods that rely on vision data only. Many robots are equipped with tactile sensors at their fingertips that could be used to complement vision data. In this paper, we present a method that uses both tactile and vision data to estimate the pose of an object grasped in a robot's hand. To address challenges like lack of standard representation for tactile data and sensor fusion, we propose the use of point clouds to represent object surfaces in contact with the tactile sensor and present a network architecture based on pixel-wise dense fusion. We also extend NVIDIA's Deep Learning Dataset Synthesizer to produce synthetic photo-realistic vision data and corresponding tactile point clouds. Results suggest that using tactile data in addition to vision data improves the 6D pose estimate, and our network generalizes successfully from synthetic training to real physical robots.

Read the original paper