Research
KineST: A Kinematics-guided Spatiotemporal State Space Model for Human Motion Tracking from Sparse Signals
KineST: A Kinematics-guided Spatiotemporal State Space Model for Human Motion Tracking from Sparse Signals Overview Research area: Computer vision and graphics — full-body human motion tracking/pose r

- arXiv
- 2512.16791
- Published
- 2025-12-18
- Authors
- Shuting Zhao, Zeyu Xiao, Xinrong Chen
AI summary
KineST: A Kinematics-guided Spatiotemporal State Space Model for Human Motion Tracking from Sparse SignalsOverview
Research area: Computer vision and graphics — full-body human motion tracking/pose reconstruction from sparse head-mounted display (HMD) signals for AR/VR.
Technical level: Intermediate. The paper builds on State Space Duality (SSD), the Mamba-2 framework, and Lie group SO(3) geometry, but its core design ideas can be described without advanced mathematics.
Scope (one sentence): The paper proposes KineST, a lightweight state space model that combines a kinematics-guided bidirectional scanning strategy, a mixed spatiotemporal representation learning mechanism, and a geometric angular velocity loss to reconstruct accurate and smooth full-body motion from the three sparse tracking signals (head and two hands) provided by AR/VR headsets.
What This Paper Is About
Head-mounted displays and controllers in AR/VR typically report only three tracking signals, from the head and the two hands, so inferring a full 22-joint body pose from them is severely under-constrained. Existing solutions either achieve good poses at high computational cost (large Transformer stacks, VQ-VAEs, diffusion models) or trade pose accuracy against motion smoothness by modeling spatial and temporal dependencies separately. KineST aims to be lightweight while delivering both high pose accuracy and temporally coherent, smooth motion.
Key Contributions
-
A kinematics-guided state space model (KineST) that extracts spatiotemporal information while integrating local and global pose perception within a lightweight framework (11M parameters in the reported comparison).
-
A Spatiotemporal Kinematic Flow Module (SKFM) that applies a Spatiotemporal Mixing Mechanism (STMM) to tightly couple spatial and temporal contexts, and a novel Kinematic Tree Scanning Strategy (KTSS) that injects kinematic priors into spatial feature capture. The scanning strategy is reformulated from the unidirectional scan of SSD into a kinematics-guided bidirectional scan, with two variants: Five-branch Kinematic Scan (FKS) and Unified Kinematic Scan (UKS).
-
A geometric angular velocity loss that constrains both the magnitude and the direction of rotational change, computing angular velocity in the tangent space (Lie algebra so(3)) of SO(3) rather than via first-order finite differences on rotation representations.
-
Experimental validation on three evaluation protocols (two AMASS-based, one real-captured), with ablation studies covering the scanning strategy, the SKFM modeling mechanism, the loss function, the flow-module components, and the input sequence length.
Main Findings
-
Best overall scores in the main comparison: On Protocol 1 (AMASS subsets CMU, BMLrub, HDM05; 90% training / 10% testing), KineST reports MPJRE 2.25, MPJPE 2.86, MPJVE 15.26, Hand PE 1.04, Upper PE 1.24, Lower PE 5.20, Root PE 2.65, Jitter 5.97, with 11M parameters. The next-best methods on those headline metrics include MMD (MPJRE 2.31, MPJPE 3.22, MPJVE 17.88, 14M parameters), HMD-Poser* (2.32, 3.15, 18.15, 17M), AvatarJLM (63M), and SAGE (137M).
-
Reported relative improvements over MMD: a 2.59% reduction in MPJRE, an 11.18% decrease in MPJPE, and a 14.65% improvement in MPJVE.
-
Trade-off with jitter-focused methods: RPM* achieves the best jitter in Protocol 1 (4.43, versus 5.97 for KineST and 4.00 for ground truth) and in Protocol 2 (3.48 versus 7.36 for KineST, with ground-truth jitter 2.93), but the paper states RPM exhibits significantly reduced pose accuracy (Protocol 1: MPJRE 3.69, MPJPE 4.88; Protocol 2: MPJRE 5.37, MPJPE 7.19).
-
Protocol 2 (larger AMASS benchmark, 12 training subsets, HumanEva and Transition for testing): KineST reports the lowest MPJRE (4.28) and MPJVE (24.08), with MPJPE 5.17 and Jitter 7.36. AvatarJLM reaches competitive MPJPE (4.93) and SAGE reports slightly lower jitter (7.13), but at 63M and 137M parameters respectively.
-
Protocol 3 (real-captured headset-and-controller data from AvatarJLM): KineST outperforms all compared methods on every metric — MPJRE 6.91, MPJPE 9.68, MPJVE 25.16, Jitter 9.49 — compared with AvatarPoser (7.28, 11.22, 31.67, 12.87) and AvatarJLM (7.01, 9.72, 27.59, 13.10).
-
Scanning strategy matters: Replacing the index-order scan in SMPL (MPJRE 2.32, MPJPE 3.11, MPJVE 17.81, Jitter 8.27) with FKS (2.28, 3.00, 16.25, 7.01) and then UKS (2.25, 2.86, 15.26, 5.97) improves all metrics. The paper attributes FKS's weaker smoothness to its branch-wise design harming body integrity.
-
STMM beats pure temporal or pure spatial modeling: Pure temporal (2.27, 2.97, 16.84, 7.83), holistic pure spatial (2.41, 3.10, 16.77, 7.72), and token-wise pure spatial (2.23, 2.93, 17.85, 9.31) all trail STMM (2.25, 2.86, 15.26, 5.97). Token-wise spatial modeling attains the lowest MPJRE but overlooks temporal continuity.
-
The geometric loss is better balanced than a finite-difference angular velocity loss: Baseline (L1 rotation + L1 orientation) gives 2.25, 2.87, 16.10, 6.75; adding the finite-difference variant gives 2.29, 3.03, 15.91, 6.44 (lower MPJVE and jitter but worse rotation and position errors); the proposed geometric loss gives 2.25, 2.86, 15.26, 5.97.
-
Every flow-module component contributes: Removing GMA (2.46, 3.16, 17.35, 6.85), removing LMA (2.27, 2.90, 16.99, 7.75), or removing the Bi-SSD block (2.37, 3.04, 20.70, 13.57) all degrade results relative to the full model.
-
Alternative losses underperform: Baseline 2.25, 2.87, 16.10, 6.75; adding a positional loss 2.26, 2.98, 16.51, 6.86; adding a velocity loss 2.29, 3.02, 17.29, 6.72; the geometric angular velocity loss 2.25, 2.86, 15.26, 5.97.
-
Sequence length 96 is the best tested setting: length 41 gives 2.32, 3.23, 18.99, 8.59; length 96 gives 2.25, 2.86, 15.26, 5.97; length 144 gives 2.47, 3.42, 18.78, 7.47; length 196 gives 2.40, 3.21, 17.46, 6.51.
-
Efficiency: Inference requires 12.9 ms to process 96 frames on an NVIDIA 4090.
Methodology in Plain English
The model takes as input a sequence of 96 frames of sparse tracking signals, where each frame carries a 3D position, 6D rotation, linear velocity, and angular velocity for each of the three tracked parts (input dimension C = 3 × (3 + 6 + 3 + 6)). A single linear layer embeds these signals, and the output is the pose parameters of the first 22 joints of the SMPL model (output dimension V = 22 × 6).
The network is built from two kinds of repeated blocks. Temporal Flow Modules (TFMs) — two of them — learn inter-frame dynamics. Each flow module contains a bidirectional SSD block (Bi-SSD) with parallel forward and backward branches; the backward branch simply processes the time-reversed input and flips it back. The two directional features are summed and then refined by a Local Motion Aggregator (LMA), a convolution-based network using a 1D convolution with kernel size 1 to capture short-range per-frame dependencies, followed by a Global Motion Aggregator (GMA), a lightweight single-layer multi-head self-attention transformer with hidden dimension constrained to 512 to capture longer-range periodic motion.
Spatiotemporal Kinematic Flow Modules (SKFMs) — also two — then inject body-structure knowledge. The core idea is to reinterpret the SSD scan as a traversal of the human kinematic tree instead of a plain index order. The joints are reordered along the kinematic chain in both forward and backward directions, using the Unified Kinematic Scan (UKS), which places the root joint centrally (forward order [21,19,17,14,15,12,20,18,16,13,9,6,3,0,1,4,7,10,2,5,8,11]) so that upper- and lower-body motion are coupled. The alternative Five-branch Kinematic Scan (FKS) strictly follows the tree branches (forward order [0,1,4,7,10,0,2,5,8,11,0,3,6,9,13,16,18,20,0,3,6,9,12,15,0,3,6,9,14,17,19,21]) but is branch-wise and less globally coherent.
The Spatiotemporal Mixing Mechanism (STMM) then reshapes the joint dimension so that sequence and joint axes are merged into a single axis before the Bi-SSD operates on them, meaning the model processes space and time together rather than one after the other. The bidirectional outputs are summed, linearly projected, and passed through LMA and GMA again.
Training uses the L1 rotation loss and L1 orientation loss plus the geometric angular velocity loss, weighted at alpha = 1, beta = 0.02, delta = 1. The geometric loss converts consecutive-frame rotations into relative rotations, maps them to axis-angle form via the matrix logarithm (using the rotation magnitude from the arccos of (Tr(V) − 1)/2 and the direction from the skew-symmetric part), and penalizes the difference between predicted and ground-truth log-rotations in L1. Training uses the Adam optimizer with learning rate 3e-4 (decayed to 3e-5 after 200,000 iterations), weight decay 1e-5, batch size 256, embedding dimension E = 256, latent joint dimension D = 64, and 22 joints. The paper reports that further details of LMA, GMA, and the geometric loss appear in the supplementary material.
Why This Matters
Impact on research: The paper shows that accurate and smooth full-body tracking can be achieved with an 11M-parameter state space model rather than the 63M- and 137M-parameter generative or Transformer architectures used by AvatarJLM and SAGE, and it demonstrates that embedding skeletal kinematic priors directly into the state space scanning order is a productive alternative to generic sequence scanning. It also argues that angular velocity supervision must respect the geometry of SO(3), since Euclidean finite differences of rotation features are not geometrically sound.
Real-world applications (as described in the paper):
- Patient rehabilitation in AR/VR.
- Realistic avatar generation and embodiment.
- Control of teleoperated humanoid robots.
- Immersive, seamless AR/VR user experiences generally.
Industry relevance: AR/VR headset and controller platforms are the intended deployment target. Because the method is described as lightweight, reports 12.9 ms inference for 96 frames on an NVIDIA 4090, and requires only the three signals (head plus two hands) that consumer headsets already provide, it targets practical on-device deployment where large generative models would be impractical.
Future Directions
- Complex motion reconstruction. The stated limitation is that reconstructing relatively complex motions such as acrobatics or gymnastics remains difficult for this approach and related methods. The authors propose incorporating richer motion priors during training while maintaining low-cost, efficient inference.
- Sequence length and temporal context. The ablation shows accuracy and smoothness degrade at input lengths of 41, 144, and 196 relative to 96, leaving the question of how to scale temporal context without loss.
- Broader real-world validation. The real-captured evaluation uses a single online headset-and-controller dataset (from AvatarJLM) under one protocol; extending to more devices, environments, and users is a natural next step.
- Sharper accuracy-smoothness trade-offs. Results show that jitter-focused methods such as RPM can reach lower jitter (4.43 on Protocol 1, 3.48 on Protocol 2) at the cost of pose accuracy, so the question of how to close the remaining gap to ground-truth jitter (4.00 and 2.93 respectively) without sacrificing accuracy stays open.
Target Audience
Researchers and engineers working on AR/VR full-body tracking, sparse-signal pose estimation, and human motion modeling; practitioners interested in state space models (Mamba/SSD) applied to structured, kinematic data; and readers studying geometrically grounded losses for rotation and motion smoothness. A reader with basic familiarity with deep learning and the SMPL body model will get the most from the paper.
Authors’ abstract
Full-body motion tracking plays an essential role in AR/VR applications, bridging physical and virtual interactions. However, it is challenging to reconstruct realistic and diverse full-body poses based on sparse signals obtained by head-mounted displays, which are the main devices in AR/VR scenarios. Existing methods for pose reconstruction often incur high computational costs or rely on separately modeling spatial and temporal dependencies, making it difficult to balance accuracy, temporal coherence, and efficiency. To address this problem, we propose KineST, a novel kinematics-guided state space model, which effectively extracts spatiotemporal dependencies while integrating local and global pose perception. The innovation comes from two core ideas. Firstly, in order to better capture intricate joint relationships, the scanning strategy within the State Space Duality framework is reformulated into kinematics-guided bidirectional scanning, which embeds kinematic priors. Secondly, a mixed spatiotemporal representation learning approach is employed to tightly couple spatial and temporal contexts, balancing accuracy and smoothness. Additionally, a geometric angular velocity loss is introduced to impose physically meaningful constraints on rotational variations for further improving motion stability. Extensive experiments demonstrate that KineST has superior performance in both accuracy and temporal consistency within a lightweight framework. Project page: https://kaka-1314.github.io/KineST/