SHAC × MuJoCo Warp · PR #1535 · latest head

Ant and Humanoid now sustain audited gaits.

At MJWarp commit 02d09b1, both final policies complete uninterrupted, full-horizon audits and ViewerGL recordings. Ant runs a low-action dynamic diagonal trot for 20 seconds; Humanoid walks at about 1.20 m/s for 15 seconds. The result required fixing the learning and evaluation system around the PR—not weakening a fall into a success.

Ant audited speed
1.010–1.015 m/s

Three independent 1024-lane noisy batches, each evaluated for the complete 20 seconds.

Humanoid audited speed
1.195–1.199 m/s

Three independent 1024-lane noisy batches, each evaluated for the complete 15 seconds.

SHAC direction error
0.472% / 0.028%

Ant / Humanoid analytic directional derivatives versus central differences; both signs agree.

ViewerGL traces
20 s + 15 s

Both exact recorded trajectories pass their task gates and remain alive through the last frame.

Bottom line

The PR can support useful gait optimization

The latest analytic adjoint supplies accurate local actor directions for both systems. Reliable locomotion still needs a robust outer loop: physically meaningful support geometry, causal action conditioning, full-horizon noisy and nominal selection, and fail-closed video validation. Ant's result is an aerial dynamic trot, not a grounded walk; Humanoid's margins on lateral drift and slip are real but narrow. Those limitations are reported rather than hidden.

01 · Final ViewerGL evidence

The exact trajectories that passed are visible

MJWarp advances the policy. Newton ViewerGL rasterizes recorded states at 960×540 and 50 fps; native MuJoCo is used only for forward kinematics, never to resimulate dynamics.

Ant · exact trace passes

21.02 m in 20.0 s

1.051 m/s, full survival, 0.141 action-change RMS, 9.75% diagonal support, 1.30 diagonal switches/s, and 56.0% flight. This is a dynamic trot. Manifest + exact-trace audit

Humanoid · exact trace passes

17.94 m in 15.0 s

1.196 m/s, full survival, 0.293 m/s support-foot slip, 83.3% single support, 3.67 support alternations/s, and 6.7% flight. Manifest + exact-trace audit

Why each manifest contains two gait evaluations

Long contact rollouts are not bitwise deterministic on this GPU. Two nominal lanes can eventually enter different contact branches. The renderer therefore audits the exact qpos/qvel/action trace sent to ViewerGL and also runs an independent sibling trajectory. Both final videos pass both checks. The 3072-lane audits below remain the primary robustness evidence; a single video is qualitative evidence, not a seed statistic.

02 · Full-horizon audit

Uninterrupted, noisy, and fail-closed

Dead lanes freeze and are never reset. Speed always divides displacement by the complete requested horizon, so falling early cannot inflate performance.

01 / RESET

1024 independently perturbed states per seed.

02 / ROLLOUT

400 Ant or 600 Humanoid control decisions, no resets.

03 / PHYSICS

MJWarp PR head, exact task XML, complete action repeat.

04 / GATE

Survival, speed, posture, steering, action and contact checks.

Performance and survival

TaskEvaluationHorizonFinal aliveSpeed|lateral|Action Δ RMSGate
Ant3 × 1024 noisy lanes20.0 s98.44–99.12%1.010–1.015 m/s1.173–1.195 m.1406–.14093 / 3 pass
Ant1024 nominal lanes20.0 s100%1.034 m/s.775 m.1411Pass
Humanoid3 × 1024 noisy lanes15.0 s99.61–100%1.195–1.199 m/s.957–.983 m.30723 / 3 pass
Humanoid1024 nominal lanes15.0 s100%1.202 m/s.957 m.3060Pass

The noisy ranges are across seeds 9851/9863/9877 for Ant and 9801/9811/9829 for Humanoid. Each seed is a separate 1024-lane batch.

Ant contact signature

Support-foot slip RMS.764–.775 m/s
Two-or-more support9.46–9.52%
Diagonal support8.91–9.02%
Diagonal switches1.018–1.035 /s
Flight58.77–59.05%
Mean up / heading≥.9957 / ≥.9894
The Ant is deliberately labeled an aerial dynamic trot. Calling this a grounded walk would contradict the measured 59% flight fraction.

Humanoid contact signature

Support-foot slip RMS.296–.298 m/s
Single support83.18–83.27%
Support switches3.617–3.634 /s
Flight7.17–7.27%
Mean up / heading≥.9916 / ≥.9792
Mean survival99.65–100%
The video and contact statistics agree: alternating single support dominates, with short flight and double-support transitions rather than foot skating.

Task gates used for final selection

TaskSurvivalSpeed / postureSteering / actionContact pattern
Ant dynamic trotfinal ≥.98; mean ≥.98speed ≥.75 m/s; up ≥.90; heading ≥.75|lateral| ≤1.50 m; action Δ ≤.35slip ≤1.10 m/s; ≥5% multi-foot and diagonal support; ≥.60 diagonal switches/s; flight ≤.60
Humanoid gaitfinal ≥.95; mean ≥.98speed ≥1.0 m/s; up ≥.85; heading ≥.75|lateral| ≤1.0 m; action Δ ≤.35slip ≤.30 m/s; ≥10% single support; ≥.80 switches/s; flight ≤.25

These are task-specific success criteria, not community benchmark thresholds. Full checks and unrounded values are in the linked audits.

03 · Latest-PR SHAC result

Accurate local gradients; guarded global decisions

The frozen PPO critic supplies the one-control-interval terminal value. Every signed actor step is accepted or rejected by exact noisy and nominal full-horizon MJWarp rollouts.

TaskAnalytic slopeCentral FD slopeRelative errorLine searchIndependent holdouts
Ant11.010411.06260.4718%−1×10⁻⁴ accepted3 noisy + nominal; all pass
Humanoid121.1886121.15480.0279%anchor retained3 noisy + nominal; all pass

Humanoid's no-update result is intentional: every nonzero candidate failed the stricter safety/robust-score guard. The analytic direction is valid locally; the full-horizon contact outcome is not assumed to be smooth globally.

What was repaired in the SHAC path

  • The complete conditioned policy—not just the raw neural actor—is differentiated.
  • Ant toe positions use exact closed-form Torch FK; its maximum error versus MJWarp is 2.98×10⁻⁸ m.
  • That FK restores smooth stance-slip and gait-schedule terms that earlier code had silently detached.
  • The physics VJP uses the PR's required out-of-place step(model, data_in, data_out) form.

What remains outside the gradient

  • Hard support classification, termination, and accept/reject gates remain non-differentiable.
  • Contact topology can change under a tiny parameter step; a matching local VJP does not guarantee a good 15–20 s rollout.
  • The signed line search therefore requires noisy and nominal gates and maximizes the weaker full-horizon score.
  • Humanoid demonstrates why this guard matters: a locally valid direction was correctly refused.
04 · Why prior trainings failed

Root causes found and fixed

The investigation followed failures through reward design, geometry, conditioning, gradient plumbing, model selection, and rendering instead of attributing every collapse to the PR.

01 / LEARNER

The original baseline was myopic and ill-scaled

The old actor loss omitted division by horizon, the tanh critic saturated, and Humanoid saw only 20 ms of direct consequence. It learned a forward dive, not balance.

The September 4 diagnostics still reproduce that failure and remain linked below as historical evidence.
02 / REWARD

Ant exploited both sides of the objective

A speed-heavy PPO run reached 1.64 m/s with 1.56 m/s support slip and .739 action-rate RMS. Heavy physical penalties instead found a ≈.23 m/s standing/local-motion optimum.

The fix uses a structured dynamic-trot prior, low action rate, robust noisy replicas, and explicit contact-pattern gates.
03 / GEOMETRY

Ant “support” was measured 2 cm too high

The old 0.120 m toe cutoff counted swing feet above the model's contact envelope as planted, corrupting support and stance-slip metrics.

The final audit uses the distal capsule endpoint and 0.100 m: 0.080 m radius + 0.020 m combined geom margin.
04 / PHASE

The first Ant gait prior scrubbed its toes

An ankle-only clock could lift but not advance. A sine hip sweep reversed during mid-stance. Explicit hip/lift phase offsets and a cosine-aligned sweep produced a stable diagonal cycle.

The final controller is phase-locked joint-space PD plus a small PPO feedback residual, fitted by parallel CEM rollouts.
05 / CONDITIONING

Humanoid was stable but contact-sensitive

Previous-action low-pass conditioning removed violent control changes. A narrow hip-roll calibration corrected drift, but adjacent settings crossed discontinuously into slip or lateral failures.

Final: α=.945 and +.345 pre-tanh bias on both hip-x outputs; a 3×1024 audit was repeated after serialization.
06 / EVIDENCE

Small selection batches and sibling renders could disagree

A 256-lane SHAC holdout missed a 1.4 mm Humanoid lateral failure found at 1024 lanes. Separately replayed video audits could enter a different contact branch.

Selection now gates noisy + nominal horizons. ViewerGL now audits the exact recorded trace and refuses a failing video.

The caught Humanoid regression was not rounded away

An earlier post-SHAC checkpoint passed two noisy seeds, but seed 9611 reached 1.001411 m lateral displacement against a 1.0 m bound; its nominal rollout reached 1.099 m. That checkpoint was rejected. A calibration sweep found the final α/bias pair, then the newly serialized policy passed another full 3×1024 audit. The rejected audit and selected calibration screen are published.

05 · Training lineage

These are hybrid policies, not cold-start SHAC trophies

The components are named explicitly so the final speed is not misrepresented as pure end-to-end SHAC training.

Ant

PPO anchor → phase-PD CEM → guarded SHAC

The final conditioning uses base-actor scale .0549 and phase-PD scale .7619. CEM selected the phase, stiffness, damping, lift and bias parameters with 32 noisy replicas per candidate at the final stage. The corrected SHAC pass then accepted one signed actor update.

This trades architectural purity for an interpretable, low-action gait that survives the requested long audit.

Humanoid

Robust PPO → causal filter/calibration → guarded SHAC

Progressive PPO polishing produced the alternating gait. A causal previous-action filter and symmetric hip-roll calibration moved it into the narrow full-horizon basin. The final SHAC direction passed its derivative test, but the safety guard retained the anchor.

“No accepted SHAC update” is the correct outcome when every tested local step weakens global safety.

06 · Scope and limitations

What this result does not establish

The final policies solve the requested test under a defined reset distribution. They are not claimed as benchmark state of the art or universal controllers.

07 · Revisions and artifacts

Everything needed to inspect the claim

The summary generator fails if revisions, hashes, gates, checkpoints, posters, or videos disagree.

Pinned execution stack

MJWarp PR head02d09b139fdf091e1e859d7f41c47a8f71d30574
Newton headd37f4d3d341ccce1e06a1dff21e9a054759b4855
SoftwarePython 3.12.13 · MuJoCo 3.12.0 · MJWarp 3.10.0.1 · Warp 1.17.0 · Torch 2.11.0+cu128
GPUNVIDIA RTX PRO 6000 Blackwell Server Edition, MIG 1g.24gb
Ant XML SHA-256cd5f83ef…05e3e2be
Humanoid XML SHA-2569b29ea10…c41e61a