Under review at ICLR 2027

Learning Expressive and Compositional Motion Representation via Spectral Skills

What should a high-level behavior model send to a whole-body controller? We learn the command, a spectral skill, by predicting how motion continues. One frozen controller then tracks diverse motion, follows a language planner, and composes new behaviors by steering a base skill along spectral directions.

Paper arXiv Code Robot videos BibTeX
One simulated Unitree G1 walks a loop, switching skills (stand, walk, jog, turning walk) and adding or removing steering (right arm up, turn, fast, slow, arms down) along the way.
One frozen controller, one loop The robot switches skills (→) and adds or removes steering (+ / −) without retraining. Physics simulation.

Try it

Steer, mix and chain skills on a simulated G1. The frozen tracker runs live in your browser.

Interactive demo
Steer the tracker along spectral directions

Spectral directions, skill mixing, chaining and pushes on one frozen policy. About 45 MB on launch.

Open in a new tab

The policy runs in ONNX Runtime Web and the physics in MuJoCo compiled to WebAssembly, 50 Hz control on 200 Hz physics. A demonstration, not a benchmark.

On the robot

One frozen controller on a Unitree G1: tracking, chaining, composition and language-directed control.

01

Tracking

Diverse whole-body motions from retargeted motion capture.

15 clips
02

Skill chaining

Skills switched at run time by changing the commanded skill.

3 clips
03

Skill composition

A base skill steered along spectral directions on the robot.

8 clips
04

Language-directed control

A language-conditioned planner predicts skills; the controller executes them.

3 clips

Method

A spectral skill is a 64-D summary of a 0.2 s motion segment, learned by predicting the motion that follows. It enters the predictor only through an affine map.

Method overview. (a) Along time, three consecutive windows: a context X spanning P frames, a segment B spanning H frames and a continuation Y spanning L frames; F_theta embeds X, the encoder E_psi maps B to the skill z, mu_theta embeds Y, and the model scores the continuation with P_theta(dY | X, Z) proportional to exp of the inner product of F_theta(X) transposed times (A Z + b) with mu_theta(Y), times nu(dY), affine in Z. (b) One interface: a skill from a reference motion through the encoder, a chunk of skills from a language planner, or a base skill steered along spectral directions (arm, speed, turn) goes to one controller that outputs actions a from observations o.
Predictive spectral skills as a shared command interface (a) A context ( frames), a segment ( frames) and a continuation ( frames). The encoder maps to a skill and is trained with a factorized diffusion model that predicts from and . (b) The same frozen controller takes skills from reference tracking, from a language planner, or from a base skill steered along spectral directions.

Each sample cuts a motion into a context of frames, a segment of frames and a continuation of frames; the encoder maps the segment to a skill:

A spectral model scores the continuation given the context and the skill:

It is learned by denoising with a noise predictor of the same factorized form:

The skill enters only through . At a fixed context , the response of the denoised prediction to the skill is therefore exact and does not depend on :

The eigenvectors of the response Gram over contexts are the spectral directions: orthonormal, each maximizing the mean squared predicted response among the directions orthogonal to the earlier ones. They are ranked by predicted response, not by usefulness.

A base skill is steered by adding directions with time-varying amplitudes:

A 50 Hz controller, trained with PPO on the frozen encoder, executes skills. The exactness holds for the predictor; how the robot responds is measured.

Robot
29-DoF Unitree G1; Isaac Lab, 200 Hz physics, 50 Hz control
Skill
one 64-D continuous skill per 10-frame window, encoder frozen after pretraining
Predictor
bilinearly factorized conditional diffusion model
Controller
PPO, simulated frames for the tracking results; composition, chaining, the robot and the demo use a deployment version trained further, to frames, with hardware-oriented regularization and extra domain randomization (encoder unchanged); the planning study uses an earlier 256-D skill controller and a 38-D explicit controller from an interface study, not the affine spectral model
Planner
conditional flow matching (GR00T action-head architecture) over skill chunks, from language and robot history

Tracking

Against the released SONIC tracker on 4,096 motions, spectral skills cut global joint error by 62 % at comparable success.

Same setting and clips as SONIC v1.1, 4,096-motion set

70.5mm
global joint error, MPJPE-G
SONIC: 187.9 mm
20.5mm
local joint error, MPJPE-L
SONIC: 26.7 mm
98.3%
success rate
SONIC: 98.9 %
≈12×
fewer simulated frames
5×10¹⁰ vs 6.3×10¹¹ reported for SONIC’s largest tracker

Where the gap comes from

SONIC reproduces body configurations well but drifts from the commanded trajectory: root drift is 98 % of its squared global error, and its root error grows at 36 mm/s against 10 mm/s for ours. Matched-budget ablations point to predictive supervision as the main source of the gain: an encoder of the same architecture trained by reconstruction keeps local accuracy but has 2.8 times the global error.

MethodSR ↑MPJPE-L ↓MPJPE-G ↓
124-motion capability set
SONIC100.0023.79173.92
Ours100.0018.2265.06
4,096-motion evaluation set
SONIC98.8826.74187.86
Ours98.2720.5470.51
Root position error over the first 8 seconds for ours and SONIC, error decomposition over successful clips, and per-clip mean root error with ours below the equality line.
Global root tracking, ours vs SONIC (a, b) Root position error over the first 8 s of every clip both methods complete that lasts at least 8 s (35 of the 124-motion set, 1,057 of the 4,096-motion set): mean and interquartile range. (c) MPJPE-L, root drift and MPJPE-G. (d) Mean root error per clip; the dashed line is equality.

On LAFAN1, trained on the same 40 clips as BFM-Zero, our controller completes 17 of the 40 clips against none for BFM-Zero, with 35.7 against 72.0 mm local error over the first 15 s. Each method runs in its own simulator.

The skill space

The skills of 129,785 motions group by behavior. Steering along spectral directions moves a walk away from the data, step by step, and the walk returns to the data when the steering is released.

Four panels. (a) t-SNE map of the skills of all 129,785 clips, colored by motion kind (walk, jog, jump, dance, crawl, stand). (b) The walking corner of the map, with forward, sideways, turning and arm-action walks. (c) A walk steered along the arm, turn and high-step directions, projected onto the arm and turn directions, over contours of corpus walks that raise an arm, turn, or both. (d) Distance from the robot's motion, encoded again, to the nearest corpus skill over time; it rises with each added direction and drops back when the steering is released, while the unsteered walk (dashed) stays flat.
Steering composes skills rarely seen in the data (a) Map of all 129,785 clips (t-SNE), colored by motion kind. (b) Walking corner. (c) A walk steered along the arm, turn and high-step directions in sequence, shown on the arm and turn directions; contours: corpus walks that raise an arm or turn. (d) Distance from the robot’s motion, encoded again, to the nearest corpus skill over time; dashed: unsteered.
Walk: directions added, then released Right arm, turn and high steps are added to a 20 s walk one at a time (a = 2, 1.2, 1.2) and released together.
Jog: the same steering ladder The same steering schedule on a jog.

Composition

Steering a base skill along one spectral direction changes one attribute while the base skill continues. The directions are computed from the trained model, without labels or retraining; what each one does is measured by running it.

Tiles of rendered motion: rows walk, jog, squat and backward walk; columns base, right arm, high steps, turn, and arm plus turn.
One controller, many composed skills Four base skills, each steered along the right-arm direction, high steps, a turn, and arm and turn together. Three poses from each executed run, older ones faded. Isaac PhysX.

Catalog of spectral directions

Of the 64 directions, the leading ones are arm gestures and later ones mostly turn the robot. Under a strict test (both signs, correlation with the prediction at least 0.7, at least 60 % of the change on the intended body part), 8 to 12 directions qualify per base and amplitude in a MuJoCo deployment simulator with sensor noise. For three tested pairs, the response to a pair agrees with the sum of the single responses (cosine 0.992 to 0.996; 0.976 for a three-direction case); additivity degrades when several directions overlap.

Direction (walk, a = 2)Measured effect
Left arm raiseshoulder pitch ±0.86 rad
Right arm raiseshoulder pitch ±0.90 rad
Both armsshoulder pitch ±0.67 rad
Right arm yawshoulder yaw ±0.43 to 0.51 rad (a = 2.5 to 3.5, some turning)
Arms outroll +0.45 rad per side, +0.43 (carry)
High stepsswing apex 15 → 36 cm (left), 14 → 32 cm (right)
Turn±130° in 4 s, stride unchanged (deployment simulator)
Speed+0.33 m/s, 6° heading change (a = 3)

The same targets, four interfaces

Ours, a SONIC token direction, a matched joint offset under our policy, and BFM-Zero’s reward prompts, each steering a walk. Kinematic replays of each method’s own rollout in one renderer, synced at the onset of steering. Ours and the joint offset run in Isaac PhysX, SONIC in a MuJoCo deployment simulator with measured sensor noise, BFM-Zero in its own MuJoCo simulator; one seed each.

A walk steered four ways: our arm direction (a = 3), a SONIC token direction and the matched joint offset on the servo targets, all on the same walk clip, and a BFM-Zero reward prompt on BFM-Zero's own walk.
Ours (top) and BFM-Zero (middle) steering a walk to raise the right arm, spread the arms, speed up and turn, with executed poses, ground paths and a top view of each path.
Ours vs BFM-Zero: executed poses and paths Both reach single targets, and both arm raises also turn the walk beyond that method’s plain walk (ours +22°, BFM-Zero +59°). BFM-Zero depends on the states its reward prompt retrieves: without the 0.85 % of its motion bank that shows walking with raised arms, its walk-and-raise prompt stops walking. Our arm direction, rebuilt from a Gram without every such clip (the model itself was still trained on them), barely changes. The two ablations act at different stages, and the methods run in different simulators on different base walks.

Spectral directions vs joint-space forcing

Left: steering along a spectral direction. Middle and right: the same change as a joint offset under the same policy, on the action and on the servo target. The offsets can pose an arm but act outside the policy: on the action they fall in 14 of 66 matched runs over three seeds, against none of the matched steering runs (three steered high-step runs that fell are left out of the set), and they do not carry gait changes.

The joint offset replays the steering run's mean joint deviation. Added to the action it overshoots the pose; on the servo target the policy cancels part of it.

Chaining against SONIC

Skills switched by changing the commanded skill, with 0.04 s transitions. Same source segments, initial pose and timing for both controllers. Switches are smoother with spectral skills than with SONIC’s token: lower RMS joint acceleration in 36 of 42 chained runs and in all 13 two-step switches.

Forward, backward, left and right jogs near 2 m/s, switched in 0.04 s. In this run neither controller falls; the RMS joint acceleration during the switches is 24.6 rad/s² for ours and 77.1 for SONIC.

Choreographies

Four short pieces from expressive motion-capture skills, chained by switching the commanded skill and steered along spectral directions, on one frozen controller. Each shown run is open loop in Isaac PhysX and does not fall; labels mark each command (→ switch skill, + add a direction, − remove it).

Match day Side stretch, high knees, jog, kick, fist pump. While jogging, the right arm rises and the run turns about 80°; both directions are released before the kick.
Moving day A walk with the right arm up and high steps, a box picked up, carried around a corner along the turn direction, placed high, then the hands dusted off.
Party guest Walk on with a brief right-arm raise and a turn to face the room, YMCA, Macarena, a bow, then walk off waving with the left arm while turning.
Title bout Shadow boxing; mid-bout the turn direction pivots the boxer about 140° to a new opponent. Then a victory pose and a fist to the chest with the right arm raised.

Language planning

A language planner that predicts skills reaches 91.1 % closed-loop success against 77.1 % for a planner whose predicted frames are re-encoded, at matched N = 10 with the same planner architecture and budget; at N = 1 the order reverses. The controllers come from an earlier interface study, not the affine spectral model.

Skills or frames as the planner’s output

A language-conditioned planner sees the instruction and the achieved robot states and samples 30 skills; the controller executes the first , then the planner is called again. Predicting skills is compared with predicting explicit frames that the frozen encoder re-encodes, and with an explicit controller without encoder. At matched , skills lower MPJPE-L by 29 % and raise success from 77.1 % to 91.1 % relative to re-encoded frames. At the order reverses: 50.76 against 41.64 mm and 84.3 % against 94.5 %. The controllers are from an earlier interface study, not the affine spectral model. 560 episodes on the 28 training instruction-motion pairs.

RouteMPJPE-L ↓SR ↑
Skill controller (256-D)
Oracle–17.191.000
Skills (140 episodes)150.760.843
Skills1042.830.911
Skills3038.410.914
Re-encoded frames141.640.945
Re-encoded frames1060.180.771
Re-encoded frames2188.640.546
Explicit controller (38-D frame, no encoder)
Explicit frames10116.460.525

Interface ablations

What makes a skill a good command. Every arm shares one tracker recipe and 2×10⁹ simulated frames and is scored on the same 4,096 clips.

SR: success rate, % · MPJPE-L, MPJPE-G: local and global mean per-joint position error, mm · ours in bold

A

Predict the whole future chunk

Supervising with the motion over the full prediction interval beats compressing the future into one state or one skill.

Prediction targetSR ↑MPJPE-L ↓MPJPE-G ↓
Ours (future chunk)92.8524.4799.51
End-point91.4826.24128.32
End-point, deterministic91.9226.25110.77
Next skill92.4625.32102.42
B

Pretrain offline, predictively

Offline predictive pretraining gives the best global tracking; reconstruction keeps local pose but drifts; learning the encoder jointly with the policy is weakest among the viable continuous and FSQ bottlenecks.

RepresentationTrainingSR ↑MPJPE-L ↓MPJPE-G ↓
Oursoffline, cont. 6492.8524.4799.51
Oursoffline, cont. 25692.7724.68104.02
Oursoffline, FSQ89.2831.19108.87
Oursoffline, categorical68.4846.97273.29
Reconstructionoffline, cont.93.5322.93278.98
Reconstructionoffline, FSQ93.6822.50238.12
Reconstructionoffline, VQ5.1062.79909.34
Reconstructionjoint, cont.84.6935.68387.05
Reconstruction + PGjoint, cont.88.9230.82418.26
Reconstruction + PGjoint, FSQ86.6730.69420.71
Reconstruction + PGjoint, VQ87.4832.29390.04
C

Continuous, dense, robot-centered

Width barely matters; learned codebooks are unstable; the encoder wants a dense, finite window of positions in the robot’s heading frame.

Design axisVariantSR ↑MPJPE-L ↓MPJPE-G ↓
Ourscont. 64-D, horizon 10, robot heading92.8524.4799.51
Skill widthcont. 128-D92.9024.7999.81
Skill widthcont. 256-D92.7724.68104.02
BottleneckFSQ 64×3289.2831.19108.87
CodebookGumbel 64×3269.6044.03180.60
Codebookcategorical 64×3268.4846.97273.29
CodebookVQ-EMA0.78125.21678.48
Encoder input+ joint velocity53.3054.85528.55
Encoder inputstride 573.5637.17147.36
Encoder inputfull window92.6324.97111.85
Encoder inputhorizon 591.5825.9396.84
Encoder inputhorizon 2056.6954.93635.31
Anchoringrobot frame92.6824.60109.21
Anchoringexpert heading91.8231.57292.63

BibTeX

@inproceedings{anonymous2027spectralskills,
  title     = {Learning Expressive and Compositional Motion Representation via Spectral Skills},
  author    = {Anonymous},
  booktitle = {Submitted to the International Conference on Learning Representations (ICLR)},
  year      = {2027},
  note      = {Under review}
}