Skip to content
AI.info

Research

Seq-DeepIPC: Sequential Sensing for End-to-End Control in Legged Robot Navigation

Overview Research area: Robotics — end-to-end (perception-to-control) navigation for legged robots, combining multi-modal sensing (RGB-D + GNSS), multi-task learning, temporal fusion, and imitation le

Seq-DeepIPC: Sequential Sensing for End-to-End Control in Legged Robot Navigation
arXiv
2510.23057
Published
2025-10-27
Authors
Oskar Natan, Jun Miura

AI summary

Overview

  • Research area: Robotics — end-to-end (perception-to-control) navigation for legged robots, combining multi-modal sensing (RGB-D + GNSS), multi-task learning, temporal fusion, and imitation learning.
  • Technical level: Advanced. The paper assumes familiarity with BEV representations, GRU recurrent models, multi-task loss weighting, imitation learning, and embedded deployment on edge hardware.
  • Scope in one sentence: Seq-DeepIPC is a lightweight, sequential end-to-end model that maps short sequences of RGB-D frames and GNSS/route-point inputs to semantic segmentation, depth, future waypoints, and position–orientation control commands for a quadruped robot navigating mixed road and grass terrain.

What This Paper Is About

Most end-to-end navigation models look at a single frame at a time, rely on an IMU compass for heading, and are demonstrated on wheeled robots on structured roads. That combination breaks down on a legged robot, whose gait causes constant high-frequency camera pitch oscillation, and in urban settings where magnetometer-based heading drifts. This paper builds a model that consumes a short sequence of RGB-D frames, fuses them into a temporally smoothed bird's-eye-view (BEV) representation, derives heading from consecutive GNSS fixes instead of an IMU, and validates the whole pipeline online on a Unitree Go2 robot dog over an extended campus route with stairs, grass, and asphalt.

Key Contributions

  1. Locomotion-aware temporal perception. Sequentially integrated RGB-D inputs (K = 3) are processed via a GRU, with the temporal window empirically tuned to damp the high-frequency camera pitch oscillations caused by legged gaits, stabilizing the BEV projection without mechanical stabilization.
  2. Magnetometer-free global heading. Bearing is derived solely from differential GNSS fixes using a geodesic formula, removing the drift that the paper attributes to hard- and soft-iron magnetic interference on IMU-based compasses in urban environments, and eliminating the noisy IMU from the sensor suite.
  3. A larger, more diverse dataset. Data was collected with a Unitree Go2, Stereolabs Zed 2i RGBD camera, and U-blox Zed-F9P RTK-GNSS receiver over an extended campus loop covering road and grass, partitioned into 26 trajectories (16 train, 5 validation, 5 test) with 5164 / 1799 / 1681 observation-set samples respectively, recorded under sunny and cloudy conditions.
  4. A compact multi-task architecture with a comprehensive legged-robot evaluation. Seq-DeepIPC jointly predicts 19-class Cityscapes-style semantic segmentation, depth, BEV semantics, five future waypoints, and (x, y, θ) controls using 12,291,227 parameters (47.5 MB), and is compared against Huang et al., AIM-MT, and the prior DeepIPC baseline with sequence-length ablations (K ∈ {1, 2, 3}).

Main Findings

  • Sequential input benefits the DeepIPC-derived models, not the baselines. DeepIPC improves from K = 1 to K = 3 (segmentation IoU 0.839 ± 0.002 to 0.847 ± 0.001, waypoint MAE 0.770 ± 0.017 to 0.729 ± 0.019, control MAE 0.055 ± 0.003 to 0.047 ± 0.001), and Seq-DeepIPC improves on depth MAE (0.088 ± 0.005 to 0.084 ± 0.001) and control MAE (0.063 ± 0.009 to 0.046 ± 0.001). By contrast, Huang et al. degrades (IoU 0.814 ± 0.003 to 0.809 ± 0.003; control MAE 0.061 ± 0.001 to 0.062 ± 0.004) and AIM-MT degrades slightly (IoU 0.847 ± 0.004 to 0.842 ± 0.003; waypoint MAE 0.710 ± 0.025 to 0.714 ± 0.026).
  • K = 3 was selected empirically as the optimal trade-off. The paper reports that shorter sequences (K < 3) failed to adequately smooth the camera shake caused by the robot's stepping gait.
  • Seq-DeepIPC is the smallest model of the four compared. 12,291,227 parameters / 47.5 MB, versus DeepIPC at 20,953,266 / 85.0 MB, AIM-MT at 27,967,063 / 112.1 MB, and Huang et al. at 74,953,360 / 300.2 MB. All models were deployed on a Jetson AGX Orin.
  • Best reported depth and control accuracy among the compared models. At K = 3, Seq-DeepIPC reports depth MAE 0.084 ± 0.001 (AIM-MT reports 0.090 ± 0.001 and 0.091 ± 0.003) and control MAE 0.046 ± 0.001 (Huang et al. reports 0.060–0.062, DeepIPC 0.047 ± 0.001).
  • Segmentation accuracy is competitive but not the top figure. Seq-DeepIPC reaches up to 0.844 ± 0.002 IoU at K = 3, compared with 0.847 reported for both AIM-MT at K = 1 and DeepIPC at K = 3.
  • Waypoint accuracy is comparable, with tighter variance. At K = 3, Seq-DeepIPC reports waypoint MAE 0.725 ± 0.010, against AIM-MT's 0.714 ± 0.026 and DeepIPC's 0.729 ± 0.019.
  • GNSS-only heading is drift-free but degrades near tall buildings. The comparison in Fig. 2 shows the magnetometer-based heading (purple) drifting over time while the GNSS-derived bearing (orange) stays globally consistent; the raw GNSS bearing is noisier at low speeds but is smoothed by the GRU policy. Performance degraded near tall buildings due to partial satellite blockage.
  • The model is "end-to-end" only at the navigation-policy level. Commands are executed by the robot's built-in low-level controller via inverse kinematics; leg locomotion dynamics are not learned by Seq-DeepIPC.

Methodology in Plain English

The robot receives a short window of K

Authors’ abstract

We present Seq-DeepIPC, a sequential end-to-end perception-to-control model for legged robot navigation in real-world environments. Seq-DeepIPC advances intelligent sensing for autonomous legged navigation by tightly integrating multi-modal perception (RGB-D + GNSS) with temporal fusion and control. The model jointly predicts semantic segmentation and depth estimation, giving richer spatial features for planning and control. For efficient deployment on edge devices, we use a lightweight model as the encoder, reducing computation while maintaining accuracy. Heading estimation is simplified by removing the noisy IMU and instead deriving global heading via differential analysis of sequential GNSS coordinates. We collected a larger and more diverse dataset that includes both road and grass terrains, and validated Seq-DeepIPC on a robot dog. Comparative and ablation studies show that sequential inputs improve perception and control in our models, while other baselines do not benefit. Seq-DeepIPC achieves competitive or better results with reasonable model size; although GNSS-only heading is less reliable near tall buildings, it is robust in open areas. Overall, Seq-DeepIPC extends end-to-end navigation beyond wheeled robots to more versatile and temporally-aware systems. To support future research, we will release the codes to our GitHub repo at https://github.com/oskarnatan/Seq-DeepIPC.

Read the original paper