Skip to content
AI.info

Research

CurveCodec 2: Skeleton-agnostic animation compression with a learned entropy model

Overview Research area: Computer graphics (cs.GR) — skeletal animation compression, spanning motion processing, data compression, rate–distortion optimization, and learned entropy coding. Published in

CurveCodec 2: Skeleton-agnostic animation compression with a learned entropy model
arXiv
2610.04211
Published
2026-10-03
Authors
Mingyi Shi, Huancheng Lin, Xuelin Chen, Taku Komura

AI summary

Overview

Research area: Computer graphics (cs.GR) — skeletal animation compression, spanning motion processing, data compression, rate–distortion optimization, and learned entropy coding. Published in TOG (per the paper's journal field), arXiv:2610.04211v1.

Technical level: Advanced. The paper assumes familiarity with quaternion algebra, forward kinematics, arithmetic/rANS entropy coding, and rate–distortion optimization.

Scope: One sentence — the paper measures where redundancy actually lives in production skeletal animation data, then builds a skeleton-agnostic codec (CurveCodec 2) that places a learned model only in the role its measurements show pays off, under explicitly verified error contracts.

What This Paper Is About

A skeletal animation clip stores rotations and translations for every joint at every frame, but most of that data is implied by the body's morphology and dynamics rather than by what the motion is actually doing. The core problem is that the industry-standard codec, ACL, is extremely fast and supports constant-time random access, but models neither time nor the statistics of its own symbols, while learned motion encoders couple their representation to a training skeleton and fall far short of production precision (the paper notes the learned encoders measured in its predecessor reconstruct with errors two orders of magnitude above the 0.01 cm a production codec targets). The goal is to find, by measurement, which part of a codec a learned model should take over, and to build one codec that serves arbitrary skeletons without retraining while bounding its error per joint.

Key Contributions

  1. A quantitative anatomy of motion redundancy. The paper dissects ACL stage by stage and runs controlled experiments to determine which structures a codec could exploit. Temporal predictability and per-joint sample sparsity paid; interpolation from the decoded neighbourhood, cross-joint structure, and learned transforms under a worst-case criterion did not.
  2. CurveCodec 2, a closed-loop, skeleton-agnostic rate–distortion codec under two explicit error contracts (a max gate and a mean gate), with a stated fallback for clips that fail either contract.
  3. A learned entropy model in the role where a network pays. The model is small and causal, with bit-exact integer inference across platforms, and it models the distribution of prediction residuals — described as the one place where a network consistently earned its cost in the authors' measurements.
  4. A corpus-scale benchmark against ACL, with audits of ACL's configuration and of the authors' own contracts, and same-core decode speed comparisons.

Main Findings

  • ACL's ratio comes from folding and range reduction, not from its back end. On the curated set, ACL's 17.7× ratio splits into 2.84× from folding sub-tracks that do not move (from 191 to 68 bits per joint-sample) and 3.74× from two-level range reduction with variable bit widths (to 18.1 bits). The zeroth-order entropy of ACL's own quantized samples is 97.7% of the bits it spends, so its widths are nearly tight for independent samples; first and second differences reach 83.8% and 81.1%, and real context coders about 0.80×, meaning an entropy coder on top of ACL saves at most 16 to 20%. A substantially smaller stream therefore needs its own quantizer.
  • The largest single saving is temporal prediction. Predicting each integer from its own past reduces the log-map stream by 39 to 55% against its zeroth-order cost. Second differences beat first differences by 18% at the root but only 4 to 7% at depth six and below, and an adaptive NLMS predictor wins on 37 to 44% of curves — which motivates a per-curve predictor choice. The authors' own quantizer sits close to ACL in zeroth-order entropy (0.92 to 1.07×) but its first differences cost 0.57 to 0.62× ACL's bits, versus 0.84× for ACL's own symbols: the gain appears once time is modelled.
  • The second saving is choosing which samples not to code. Under the max gate the encoder keeps 87% of the root's samples but only 57% at depth six and below at 0.01 cm, and 51% against 22% at 0.1 cm. Keys fall from 79% to 16% of samples as precision loosens, yet at p ≤ 0.1 cm more than three quarters of the remaining gaps hide at most two samples. The root keeps the most keys because its error reaches every descendant. ACL can only strip whole frames whose every joint is linearly interpolable — 0.38% of curated frames at 0.01 cm.
  • At the margin, omitting a sample is the cheapest way to spend error. An iso-mean encoder at 0.1 cm pays 0.17 step units for growing a step by one grid unit, 0.31 for a dead zone, and about 0.03 for removing an RD-selected key. Removing keys from the final codec costs 8% at 0.01 cm and 69% at 1 cm, where it matches prediction.
  • The gain from keys comes from the closed loop, not the key representation. Removing keys by a local per-curve error bound saves only 0 to 0.6%. A dyadic hierarchical codec with a learned in-betweener matches the authors' approach where keys keep 83% of the samples (the provided text is truncated at this point).
  • Interpolation has little left to find on the gaps a good encoder leaves. A nearest-neighbour oracle with millions of training samples as memory, which clearly beats cubic interpolation on random gaps, is no better than linear interpolation on the gaps the encoder actually leaves, and none of the learned in-betweeners the authors placed inside the codec paid for itself. Cubic Catmull–Rom interpolation of log-map values saves 0.5 to 2.6% over linear interpolation.
  • Log-map rotations remove a precision floor. ACL stores the (x, y, z) part of a quaternion and reconstructs w, which amplifies error near |w| ≈ 0. On the curated set ACL's maximum error is 0.071 cm at every precision from 0.01 down to 0.001 cm. Coding the rotation vector lowers the curated maximum at 0.01 cm from 0.071 to 0.024 cm. The log map is discontinuous across a full turn, which occurs on 0.33% of the test side's rotation tracks and costs below 0.2% of the bytes.
  • The end-to-end result. On a held-out test side of 4,472 clips from 33 datasets, CurveCodec 2 needs 0.37× ACL's bytes at ACL's default precision of 0.01 cm under the worst-case contract, and 0.22× at 0.1 cm under the mean contract. The estimated size of the 906-hour corpus drops from ACL's 32.6 GB to 11.9 GB at p = 0.01 cm, plus one 870 KB model shared by all clips and not counted in the stream.
  • Preprocessing is shared and accounts for a large share of the raw reduction. Dropping scale reduces ACL's raw 647 GB to 453 GB; at p = 0.01 cm folding leaves only 53% of rotation and 3.6% of translation sub-tracks animated. Both codecs start the second stage from the same animated sub-tracks.
  • ACL does not guarantee its own precision. At 0.01 cm, a clip's maximum error stays within p on only 1.5% of the test clips, which is why the paper holds both codecs to ACL's measured error on each clip rather than to the nominal p.
  • Decoding is CPU-only but whole-clip. CurveCodec 2 decodes whole clips on one CPU core, typically in a fraction of a second, and it is a storage and distribution format decoded at load time rather than a replacement for ACL's constant-time sampling.

Methodology in Plain English

The authors treat the codec as an instrument: under an explicit error contract and an exact bit count, a candidate way of structuring the motion either pays for itself or it does not. They first dissect ACL stage by stage to see where its bits go, then run controlled experiments substituting a learned component at each place one could plausibly enter — as an interpolator, as a transform, as a tokenizer, as a field.

The pipeline that survives has two parts. The lossy part runs once per clip inside the encoder. Each animated sub-track becomes a curve: rotations in the log map, translations in centimetres. A closed loop then proposes two kinds of moves — coarsening a curve's quantization step, or dropping one of its keys — decodes the candidate, recomputes the shell error of the affected subtree through forward kinematics, and accepts the move only if the error contract still holds. The skeleton's geometry is used here and nowhere else.

The lossless part only affects the number of bits, never the decoded result, so the network cannot violate the error contract. A fixed predictor (chosen per column among first and second differences, a scaled second difference, and a fourth-order integer NLMS filter) turns the quantized integers into residuals. A small learned model — a causal transformer of 108K parameters, or a 22K-parameter MLP under the max gate — reads 36 integer features of each curve's own past and outputs a mixture of three logistics whose cumulative distribution gives symbol probabilities, which drive a rANS coder. The network never sees joint identity; its only hierarchy-derived input is a four-class depth bucket. Encoder and decoder must compute identical probabilities, so the network is exported to fixed point with 16-bit weights, 64-bit accumulators, and lookup tables, with the residual stream clipped to ±2²⁸ so every downstream operand stays inside proven integer bounds.

Two contracts are verified on every decoded clip: ACL's own worst case per joint within a stated tolerance, or ACL's mean error per clip. A failing clip is re-encoded with internal limits scaled by 0.98, then 0.95, then 0.90; if none passes, or the result is larger than ACL's, ACL's own stream is stored. All reported byte counts include this fallback.

Why This Matters

Impact on research. The paper reframes animation compression as a measurement question — "what must the codec still send once the body is known?" — and delivers negative results as well as positive ones: learned interpolators, learned transforms, tokenizers, and neural fields did not pay for themselves in this setting, while the learned probability model did. That is a concrete, reproducible steer for anyone placing learned components inside a graphics codec.

Real-world applications:

  • Game and film animation libraries, where the paper's framing is strict memory and streaming budgets, and clips are decoded at load time rather than sampled randomly.
  • Motion-capture corpus storage and distribution, where corpora reach hundreds of hours; the paper estimates its own 906-hour corpus moving from 32.6 GB to 11.9 GB at p = 0.01 cm.
  • Streaming or networked delivery of character animation, where bitrate matters and the stated per-joint error bound is a quality guarantee rather than a hope.
  • Retargeting and multi-character pipelines, since one model trained once codes all 12 rigs shown in the teaser figure and transfers without retraining to a species absent from training.

Industry relevance. ACL is the production library of modern game engines, and the paper benchmarks directly against it — including a specific commit (3ee5685, the development branch after release 2.1.0) — and audits ACL's own configuration. The explicit fallback of storing ACL's stream when CurveCodec 2 loses means the codec is never worse in size than the incumbent on a per-clip basis, which is the kind of guarantee a shipping pipeline needs.

Future Directions

  • Random access. CurveCodec 2 decodes only whole clips. Reconciling its ratio with ACL's constant-time sampling would widen its applicability to runtime playback rather than load-time storage.
  • The loose-precision mean gate. The mean gate's guard is ten times ACL's clip mean rather than a multiple of p, and the paper states that at loose precision it admits large isolated errors, so it suits storage at tight precision. A contract that stays tight at both ends is left open.
  • Full-turn discontinuity. The log-map representation is continuous only short of a full turn; the jump costs a few keys and below 0.2% of bytes on 0.33% of test-side rotation tracks. The paper evaluates an unwrapping rule in supplementary Sec. C, but a cleaner treatment remains a candidate improvement.
  • The tolerance question. The max gate holds ACL's worst case within tolerance rather than identically, and the paper defers the question of what the tolerances are worth to Sec. 6.3, which is not included in the provided content. Per-dataset breakdowns of the 0.37× and 0.22× ratios, and any comparison against learned codecs other than those the authors built and measured, are likewise not reported in the content supplied.

Target Audience

Researchers and engineers working on animation compression, motion capture data pipelines, and neural data compression; graphics practitioners at studios or engine teams who need a production-grade, error-bounded alternative to ACL; and machine learning researchers interested in where a learned probability model earns its cost inside a classical codec. The measurements in Section 5 — the ACL bit budget, the key-versus-step marginal cost comparison, and the interpolation oracles — are useful on their own to anyone designing a motion codec, even without adopting CurveCodec 2.

Authors’ abstract

Skeletal motion is stored as every joint's transform at every frame, yet most of it is implied by the body rather than by what the motion is about. Compression is one way to ask what a motion must still say once the body is known, and a production codec must answer it for any skeleton with a stated error bound. Our earlier codec, CurveCodec, matched the mean error of ACL, the production library of modern game engines, with a learned prior over sparse anchors, but not ACL's worst case, and it counted its payload as floats rather than bits. Here we ask where the redundancy of skeletal motion lies and which part of a codec a learned model should take over. Measurements give three answers. At production precision the largest saving comes from predicting each quantized curve from its own past, the second from choosing per joint, in closed loop through the hierarchy, which samples not to code. On the gaps such an encoder leaves, a nearest-neighbour oracle over millions of training samples is no better than linear interpolation, and no learned in-betweener we tried paid for itself. What a network does learn is the distribution of the residuals the codec must send. CurveCodec 2 codes every sub-track as a curve in the log map, quantized in closed loop and thinned to rate-distortion-selected keys, with residuals entropy-coded under a small learned model whose integer inference is bit-exact across platforms. Two contracts are verified on every decoded clip: ACL's own worst case per joint within a stated tolerance, or ACL's mean error per clip. On a held-out test side of 4,472 clips from 33 datasets, CurveCodec 2 needs 0.37x ACL's bytes at ACL's default precision of 0.01 cm under the worst-case contract and 0.22x at 0.1 cm under the mean contract, decodes on one CPU core, and transfers without retraining to a species absent from training. Project page: https://rubbly.cn/publications/curvecodec/

Read the original paper