Three independent 1024-lane noisy batches, each evaluated for the complete 20 seconds.
Ant and Humanoid now sustain audited gaits.
At MJWarp commit 02d09b1, both final policies complete uninterrupted, full-horizon
audits and ViewerGL recordings. Ant runs a low-action dynamic diagonal trot for 20 seconds;
Humanoid walks at about 1.20 m/s for 15 seconds. The result required fixing the learning and
evaluation system around the PR—not weakening a fall into a success.
Three independent 1024-lane noisy batches, each evaluated for the complete 15 seconds.
Ant / Humanoid analytic directional derivatives versus central differences; both signs agree.
Both exact recorded trajectories pass their task gates and remain alive through the last frame.
The PR can support useful gait optimization
The latest analytic adjoint supplies accurate local actor directions for both systems. Reliable locomotion still needs a robust outer loop: physically meaningful support geometry, causal action conditioning, full-horizon noisy and nominal selection, and fail-closed video validation. Ant's result is an aerial dynamic trot, not a grounded walk; Humanoid's margins on lateral drift and slip are real but narrow. Those limitations are reported rather than hidden.
The exact trajectories that passed are visible
MJWarp advances the policy. Newton ViewerGL rasterizes recorded states at 960×540 and 50 fps; native MuJoCo is used only for forward kinematics, never to resimulate dynamics.
21.02 m in 20.0 s
1.051 m/s, full survival, 0.141 action-change RMS, 9.75% diagonal support, 1.30 diagonal switches/s, and 56.0% flight. This is a dynamic trot. Manifest + exact-trace audit
17.94 m in 15.0 s
1.196 m/s, full survival, 0.293 m/s support-foot slip, 83.3% single support, 3.67 support alternations/s, and 6.7% flight. Manifest + exact-trace audit
Why each manifest contains two gait evaluations
Long contact rollouts are not bitwise deterministic on this GPU. Two nominal lanes can eventually enter different contact branches. The renderer therefore audits the exact qpos/qvel/action trace sent to ViewerGL and also runs an independent sibling trajectory. Both final videos pass both checks. The 3072-lane audits below remain the primary robustness evidence; a single video is qualitative evidence, not a seed statistic.
Uninterrupted, noisy, and fail-closed
Dead lanes freeze and are never reset. Speed always divides displacement by the complete requested horizon, so falling early cannot inflate performance.
1024 independently perturbed states per seed.
400 Ant or 600 Humanoid control decisions, no resets.
MJWarp PR head, exact task XML, complete action repeat.
Survival, speed, posture, steering, action and contact checks.
Performance and survival
| Task | Evaluation | Horizon | Final alive | Speed | |lateral| | Action Δ RMS | Gate |
|---|---|---|---|---|---|---|---|
| Ant | 3 × 1024 noisy lanes | 20.0 s | 98.44–99.12% | 1.010–1.015 m/s | 1.173–1.195 m | .1406–.1409 | 3 / 3 pass |
| Ant | 1024 nominal lanes | 20.0 s | 100% | 1.034 m/s | .775 m | .1411 | Pass |
| Humanoid | 3 × 1024 noisy lanes | 15.0 s | 99.61–100% | 1.195–1.199 m/s | .957–.983 m | .3072 | 3 / 3 pass |
| Humanoid | 1024 nominal lanes | 15.0 s | 100% | 1.202 m/s | .957 m | .3060 | Pass |
The noisy ranges are across seeds 9851/9863/9877 for Ant and 9801/9811/9829 for Humanoid. Each seed is a separate 1024-lane batch.
Ant contact signature
| Support-foot slip RMS | .764–.775 m/s |
|---|---|
| Two-or-more support | 9.46–9.52% |
| Diagonal support | 8.91–9.02% |
| Diagonal switches | 1.018–1.035 /s |
| Flight | 58.77–59.05% |
| Mean up / heading | ≥.9957 / ≥.9894 |
Humanoid contact signature
| Support-foot slip RMS | .296–.298 m/s |
|---|---|
| Single support | 83.18–83.27% |
| Support switches | 3.617–3.634 /s |
| Flight | 7.17–7.27% |
| Mean up / heading | ≥.9916 / ≥.9792 |
| Mean survival | 99.65–100% |
Task gates used for final selection
| Task | Survival | Speed / posture | Steering / action | Contact pattern |
|---|---|---|---|---|
| Ant dynamic trot | final ≥.98; mean ≥.98 | speed ≥.75 m/s; up ≥.90; heading ≥.75 | |lateral| ≤1.50 m; action Δ ≤.35 | slip ≤1.10 m/s; ≥5% multi-foot and diagonal support; ≥.60 diagonal switches/s; flight ≤.60 |
| Humanoid gait | final ≥.95; mean ≥.98 | speed ≥1.0 m/s; up ≥.85; heading ≥.75 | |lateral| ≤1.0 m; action Δ ≤.35 | slip ≤.30 m/s; ≥10% single support; ≥.80 switches/s; flight ≤.25 |
These are task-specific success criteria, not community benchmark thresholds. Full checks and unrounded values are in the linked audits.
Accurate local gradients; guarded global decisions
The frozen PPO critic supplies the one-control-interval terminal value. Every signed actor step is accepted or rejected by exact noisy and nominal full-horizon MJWarp rollouts.
| Task | Analytic slope | Central FD slope | Relative error | Line search | Independent holdouts |
|---|---|---|---|---|---|
| Ant | 11.0104 | 11.0626 | 0.4718% | −1×10⁻⁴ accepted | 3 noisy + nominal; all pass |
| Humanoid | 121.1886 | 121.1548 | 0.0279% | anchor retained | 3 noisy + nominal; all pass |
Humanoid's no-update result is intentional: every nonzero candidate failed the stricter safety/robust-score guard. The analytic direction is valid locally; the full-horizon contact outcome is not assumed to be smooth globally.
What was repaired in the SHAC path
- The complete conditioned policy—not just the raw neural actor—is differentiated.
- Ant toe positions use exact closed-form Torch FK; its maximum error versus MJWarp is
2.98×10⁻⁸ m. - That FK restores smooth stance-slip and gait-schedule terms that earlier code had silently detached.
- The physics VJP uses the PR's required out-of-place
step(model, data_in, data_out)form.
What remains outside the gradient
- Hard support classification, termination, and accept/reject gates remain non-differentiable.
- Contact topology can change under a tiny parameter step; a matching local VJP does not guarantee a good 15–20 s rollout.
- The signed line search therefore requires noisy and nominal gates and maximizes the weaker full-horizon score.
- Humanoid demonstrates why this guard matters: a locally valid direction was correctly refused.
Root causes found and fixed
The investigation followed failures through reward design, geometry, conditioning, gradient plumbing, model selection, and rendering instead of attributing every collapse to the PR.
The original baseline was myopic and ill-scaled
The old actor loss omitted division by horizon, the tanh critic saturated, and Humanoid saw only 20 ms of direct consequence. It learned a forward dive, not balance.
Ant exploited both sides of the objective
A speed-heavy PPO run reached 1.64 m/s with 1.56 m/s support slip and .739 action-rate RMS. Heavy physical penalties instead found a ≈.23 m/s standing/local-motion optimum.
Ant “support” was measured 2 cm too high
The old 0.120 m toe cutoff counted swing feet above the model's contact envelope as planted, corrupting support and stance-slip metrics.
The first Ant gait prior scrubbed its toes
An ankle-only clock could lift but not advance. A sine hip sweep reversed during mid-stance. Explicit hip/lift phase offsets and a cosine-aligned sweep produced a stable diagonal cycle.
Humanoid was stable but contact-sensitive
Previous-action low-pass conditioning removed violent control changes. A narrow hip-roll calibration corrected drift, but adjacent settings crossed discontinuously into slip or lateral failures.
Small selection batches and sibling renders could disagree
A 256-lane SHAC holdout missed a 1.4 mm Humanoid lateral failure found at 1024 lanes. Separately replayed video audits could enter a different contact branch.
The caught Humanoid regression was not rounded away
An earlier post-SHAC checkpoint passed two noisy seeds, but seed 9611 reached 1.001411 m lateral displacement against a 1.0 m bound; its nominal rollout reached 1.099 m. That checkpoint was rejected. A calibration sweep found the final α/bias pair, then the newly serialized policy passed another full 3×1024 audit. The rejected audit and selected calibration screen are published.
These are hybrid policies, not cold-start SHAC trophies
The components are named explicitly so the final speed is not misrepresented as pure end-to-end SHAC training.
PPO anchor → phase-PD CEM → guarded SHAC
The final conditioning uses base-actor scale .0549 and phase-PD scale .7619. CEM selected the phase, stiffness, damping, lift and bias parameters with 32 noisy replicas per candidate at the final stage. The corrected SHAC pass then accepted one signed actor update.
This trades architectural purity for an interpretable, low-action gait that survives the requested long audit.
Robust PPO → causal filter/calibration → guarded SHAC
Progressive PPO polishing produced the alternating gait. A causal previous-action filter and symmetric hip-roll calibration moved it into the narrow full-horizon basin. The final SHAC direction passed its derivative test, but the safety guard retained the anchor.
“No accepted SHAC update” is the correct outcome when every tested local step weakens global safety.
What this result does not establish
The final policies solve the requested test under a defined reset distribution. They are not claimed as benchmark state of the art or universal controllers.
- Ant has ≈59% aerial control samples and ≈9% diagonal support. It is a dynamic trot, not a grounded walk.
- Humanoid's noisy lateral range (.957–.983 m against 1.0 m) and slip range (.296–.298 m/s against .300) leave limited margin.
- There are three seed batches per task, albeit 1024 independently perturbed lanes in each. No terrain, friction, mass, or actuator domain randomization was tested.
- GPU contact execution is not bitwise deterministic. Final claims rely on large ensembles; videos additionally gate their exact recorded traces.
- Training calls MJWarp's out-of-place step directly through a PyTorch bridge. It does not demonstrate gradients through Newton's public
SolverMuJoCoadapter. - The compact implementation is SHAC-style: short differentiable actor horizon, frozen terminal critic, signed full-horizon line search. It is not a canonical reproduction of every detail in the original SHAC codebase.
Everything needed to inspect the claim
The summary generator fails if revisions, hashes, gates, checkpoints, posters, or videos disagree.
Pinned execution stack
| MJWarp PR head | 02d09b139fdf091e1e859d7f41c47a8f71d30574 |
|---|---|
| Newton head | d37f4d3d341ccce1e06a1dff21e9a054759b4855 |
| Software | Python 3.12.13 · MuJoCo 3.12.0 · MJWarp 3.10.0.1 · Warp 1.17.0 · Torch 2.11.0+cu128 |
| GPU | NVIDIA RTX PRO 6000 Blackwell Server Edition, MIG 1g.24gb |
| Ant XML SHA-256 | cd5f83ef…05e3e2be |
| Humanoid XML SHA-256 | 9b29ea10…c41e61a |
Final machine-readable bundle
- Validated cross-artifact summary
- Ant: checkpoint · SHAC source · SHAC manifest · audit · CEM search
- Humanoid: checkpoint · SHAC source · SHAC manifest · audit · calibration · PPO history
- ViewerGL: Ant manifest · Humanoid manifest
Final v3 source
September 4 diagnostic archive
The earlier report correctly diagnosed the failing baseline but ended before these full-gait runs. Its artifacts remain available for comparison: