Learning Expressive and Compositional Motion Representation via Spectral Skills
What should a high-level behavior model send to a whole-body controller? We learn the command, a spectral skill, by predicting how motion continues. One frozen controller then tracks diverse motion, follows a language planner, and composes new behaviors by steering a base skill along spectral directions.

Try it
Steer, mix and chain skills on a simulated G1. The frozen tracker runs live in your browser.
The policy runs in ONNX Runtime Web and the physics in MuJoCo compiled to WebAssembly, 50 Hz control on 200 Hz physics. A demonstration, not a benchmark.
On the robot
One frozen controller on a Unitree G1: tracking, chaining, composition and language-directed control.
Tracking
Diverse whole-body motions from retargeted motion capture.
Skill chaining
Skills switched at run time by changing the commanded skill.
Skill composition
A base skill steered along spectral directions on the robot.
Language-directed control
A language-conditioned planner predicts skills; the controller executes them.
Method
A spectral skill is a 64-D summary of a 0.2 s motion segment, learned by predicting the motion that follows. It enters the predictor only through an affine map.

Each sample cuts a motion into a context of frames, a segment of frames and a continuation of frames; the encoder maps the segment to a skill:
A spectral model scores the continuation given the context and the skill:
It is learned by denoising with a noise predictor of the same factorized form:
The skill enters only through . At a fixed context , the response of the denoised prediction to the skill is therefore exact and does not depend on :
The eigenvectors of the response Gram over contexts are the spectral directions: orthonormal, each maximizing the mean squared predicted response among the directions orthogonal to the earlier ones. They are ranked by predicted response, not by usefulness.
A base skill is steered by adding directions with time-varying amplitudes:
A 50 Hz controller, trained with PPO on the frozen encoder, executes skills. The exactness holds for the predictor; how the robot responds is measured.
- Robot
- 29-DoF Unitree G1; Isaac Lab, 200 Hz physics, 50 Hz control
- Skill
- one 64-D continuous skill per 10-frame window, encoder frozen after pretraining
- Predictor
- bilinearly factorized conditional diffusion model
- Controller
- PPO, simulated frames for the tracking results; composition, chaining, the robot and the demo use a deployment version trained further, to frames, with hardware-oriented regularization and extra domain randomization (encoder unchanged); the planning study uses an earlier 256-D skill controller and a 38-D explicit controller from an interface study, not the affine spectral model
- Planner
- conditional flow matching (GR00T action-head architecture) over skill chunks, from language and robot history
Tracking
Against the released SONIC tracker on 4,096 motions, spectral skills cut global joint error by 62 % at comparable success.
Same setting and clips as SONIC v1.1, 4,096-motion set
Where the gap comes from
SONIC reproduces body configurations well but drifts from the commanded trajectory: root drift is 98 % of its squared global error, and its root error grows at 36 mm/s against 10 mm/s for ours. Matched-budget ablations point to predictive supervision as the main source of the gain: an encoder of the same architecture trained by reconstruction keeps local accuracy but has 2.8 times the global error.
| Method | SR ↑ | MPJPE-L ↓ | MPJPE-G ↓ |
|---|---|---|---|
| 124-motion capability set | |||
| SONIC | 100.00 | 23.79 | 173.92 |
| Ours | 100.00 | 18.22 | 65.06 |
| 4,096-motion evaluation set | |||
| SONIC | 98.88 | 26.74 | 187.86 |
| Ours | 98.27 | 20.54 | 70.51 |

On LAFAN1, trained on the same 40 clips as BFM-Zero, our controller completes 17 of the 40 clips against none for BFM-Zero, with 35.7 against 72.0 mm local error over the first 15 s. Each method runs in its own simulator.
The skill space
The skills of 129,785 motions group by behavior. Steering along spectral directions moves a walk away from the data, step by step, and the walk returns to the data when the steering is released.

Composition
Steering a base skill along one spectral direction changes one attribute while the base skill continues. The directions are computed from the trained model, without labels or retraining; what each one does is measured by running it.

Catalog of spectral directions
Of the 64 directions, the leading ones are arm gestures and later ones mostly turn the robot. Under a strict test (both signs, correlation with the prediction at least 0.7, at least 60 % of the change on the intended body part), 8 to 12 directions qualify per base and amplitude in a MuJoCo deployment simulator with sensor noise. For three tested pairs, the response to a pair agrees with the sum of the single responses (cosine 0.992 to 0.996; 0.976 for a three-direction case); additivity degrades when several directions overlap.
| Direction (walk, a = 2) | Measured effect |
|---|---|
| Left arm raise | shoulder pitch ±0.86 rad |
| Right arm raise | shoulder pitch ±0.90 rad |
| Both arms | shoulder pitch ±0.67 rad |
| Right arm yaw | shoulder yaw ±0.43 to 0.51 rad (a = 2.5 to 3.5, some turning) |
| Arms out | roll +0.45 rad per side, +0.43 (carry) |
| High steps | swing apex 15 → 36 cm (left), 14 → 32 cm (right) |
| Turn | ±130° in 4 s, stride unchanged (deployment simulator) |
| Speed | +0.33 m/s, 6° heading change (a = 3) |
The same targets, four interfaces
Ours, a SONIC token direction, a matched joint offset under our policy, and BFM-Zero’s reward prompts, each steering a walk. Kinematic replays of each method’s own rollout in one renderer, synced at the onset of steering. Ours and the joint offset run in Isaac PhysX, SONIC in a MuJoCo deployment simulator with measured sensor noise, BFM-Zero in its own MuJoCo simulator; one seed each.

Spectral directions vs joint-space forcing
Left: steering along a spectral direction. Middle and right: the same change as a joint offset under the same policy, on the action and on the servo target. The offsets can pose an arm but act outside the policy: on the action they fall in 14 of 66 matched runs over three seeds, against none of the matched steering runs (three steered high-step runs that fell are left out of the set), and they do not carry gait changes.
Chaining against SONIC
Skills switched by changing the commanded skill, with 0.04 s transitions. Same source segments, initial pose and timing for both controllers. Switches are smoother with spectral skills than with SONIC’s token: lower RMS joint acceleration in 36 of 42 chained runs and in all 13 two-step switches.
Choreographies
Four short pieces from expressive motion-capture skills, chained by switching the commanded skill and steered along spectral directions, on one frozen controller. Each shown run is open loop in Isaac PhysX and does not fall; labels mark each command (→ switch skill, + add a direction, − remove it).
Language planning
A language planner that predicts skills reaches 91.1 % closed-loop success against 77.1 % for a planner whose predicted frames are re-encoded, at matched N = 10 with the same planner architecture and budget; at N = 1 the order reverses. The controllers come from an earlier interface study, not the affine spectral model.
Skills or frames as the planner’s output
A language-conditioned planner sees the instruction and the achieved robot states and samples 30 skills; the controller executes the first , then the planner is called again. Predicting skills is compared with predicting explicit frames that the frozen encoder re-encodes, and with an explicit controller without encoder. At matched , skills lower MPJPE-L by 29 % and raise success from 77.1 % to 91.1 % relative to re-encoded frames. At the order reverses: 50.76 against 41.64 mm and 84.3 % against 94.5 %. The controllers are from an earlier interface study, not the affine spectral model. 560 episodes on the 28 training instruction-motion pairs.
| Route | MPJPE-L ↓ | SR ↑ | |
|---|---|---|---|
| Skill controller (256-D) | |||
| Oracle | – | 17.19 | 1.000 |
| Skills (140 episodes) | 1 | 50.76 | 0.843 |
| Skills | 10 | 42.83 | 0.911 |
| Skills | 30 | 38.41 | 0.914 |
| Re-encoded frames | 1 | 41.64 | 0.945 |
| Re-encoded frames | 10 | 60.18 | 0.771 |
| Re-encoded frames | 21 | 88.64 | 0.546 |
| Explicit controller (38-D frame, no encoder) | |||
| Explicit frames | 10 | 116.46 | 0.525 |
Interface ablations
What makes a skill a good command. Every arm shares one tracker recipe and 2×10⁹ simulated frames and is scored on the same 4,096 clips.
SR: success rate, % · MPJPE-L, MPJPE-G: local and global mean per-joint position error, mm · ours in bold
Predict the whole future chunk
Supervising with the motion over the full prediction interval beats compressing the future into one state or one skill.
| Prediction target | SR ↑ | MPJPE-L ↓ | MPJPE-G ↓ |
|---|---|---|---|
| Ours (future chunk) | 92.85 | 24.47 | 99.51 |
| End-point | 91.48 | 26.24 | 128.32 |
| End-point, deterministic | 91.92 | 26.25 | 110.77 |
| Next skill | 92.46 | 25.32 | 102.42 |
Pretrain offline, predictively
Offline predictive pretraining gives the best global tracking; reconstruction keeps local pose but drifts; learning the encoder jointly with the policy is weakest among the viable continuous and FSQ bottlenecks.
| Representation | Training | SR ↑ | MPJPE-L ↓ | MPJPE-G ↓ |
|---|---|---|---|---|
| Ours | offline, cont. 64 | 92.85 | 24.47 | 99.51 |
| Ours | offline, cont. 256 | 92.77 | 24.68 | 104.02 |
| Ours | offline, FSQ | 89.28 | 31.19 | 108.87 |
| Ours | offline, categorical | 68.48 | 46.97 | 273.29 |
| Reconstruction | offline, cont. | 93.53 | 22.93 | 278.98 |
| Reconstruction | offline, FSQ | 93.68 | 22.50 | 238.12 |
| Reconstruction | offline, VQ | 5.10 | 62.79 | 909.34 |
| Reconstruction | joint, cont. | 84.69 | 35.68 | 387.05 |
| Reconstruction + PG | joint, cont. | 88.92 | 30.82 | 418.26 |
| Reconstruction + PG | joint, FSQ | 86.67 | 30.69 | 420.71 |
| Reconstruction + PG | joint, VQ | 87.48 | 32.29 | 390.04 |
Continuous, dense, robot-centered
Width barely matters; learned codebooks are unstable; the encoder wants a dense, finite window of positions in the robot’s heading frame.
| Design axis | Variant | SR ↑ | MPJPE-L ↓ | MPJPE-G ↓ |
|---|---|---|---|---|
| Ours | cont. 64-D, horizon 10, robot heading | 92.85 | 24.47 | 99.51 |
| Skill width | cont. 128-D | 92.90 | 24.79 | 99.81 |
| Skill width | cont. 256-D | 92.77 | 24.68 | 104.02 |
| Bottleneck | FSQ 64×32 | 89.28 | 31.19 | 108.87 |
| Codebook | Gumbel 64×32 | 69.60 | 44.03 | 180.60 |
| Codebook | categorical 64×32 | 68.48 | 46.97 | 273.29 |
| Codebook | VQ-EMA | 0.78 | 125.21 | 678.48 |
| Encoder input | + joint velocity | 53.30 | 54.85 | 528.55 |
| Encoder input | stride 5 | 73.56 | 37.17 | 147.36 |
| Encoder input | full window | 92.63 | 24.97 | 111.85 |
| Encoder input | horizon 5 | 91.58 | 25.93 | 96.84 |
| Encoder input | horizon 20 | 56.69 | 54.93 | 635.31 |
| Anchoring | robot frame | 92.68 | 24.60 | 109.21 |
| Anchoring | expert heading | 91.82 | 31.57 | 292.63 |
BibTeX
@inproceedings{anonymous2027spectralskills,
title = {Learning Expressive and Compositional Motion Representation via Spectral Skills},
author = {Anonymous},
booktitle = {Submitted to the International Conference on Learning Representations (ICLR)},
year = {2027},
note = {Under review}
}