State-of-the-art perceptive humanoid locomotion is mostly built on retargeted motion capture, a generative motion prior, and a distillation stage. We report a complete system that does not require this pipeline for a humanoid to cross varied terrain. In its place, we provide a five-expert mixture routed by a terrain classifier over a sensed elevation grid, and target every foothold with a model-based planner, refined by six closed-form-gradient steps on an interpolated terrain-cost field instead of picked from a fixed grid. Driven through the same perception-to-actuation software a robot would run, inside a second, deployment-fidelity simulator, it crosses an 85 m five-segment waypoint course 18 of 20 times end-to-end (90%, Wilson 95% CI 70–97) and succeeds 20/20 on four of five segments scored alone, where an ablated configuration without the mixture and without refinement reaches 0% on both the stair and the hurdle segment. The continuous refinement moves 72.4% of foot placements across at least one grid-cell boundary, which produces real sub-grid correction. We also ran matched component studies to report what does not account for it: one expert equals five on the training distribution, four label-free routers collapse to near-uniform activation, and a self-supervised terrain embedding loses to the privileged label.
Everything below the dashed line runs unchanged in both simulators and on a robot. Above it is training-only: a privileged terrain label supervises the gate and, in one configuration, conditions the critic; a self-supervised depth embedding is the alternative critic conditioner compared against it. The deployed actor consumes neither.
A terrain classifier routes a five-expert mixture over a sensed local elevation grid. The only privileged signal anywhere in the deployed path is the terrain class used to train the gate — no motion-capture data, no generative motion prior.
Every foothold target is refined by six closed-form gradient-descent steps on an interpolated terrain-cost field, instead of picked from a fixed grid — real sub-grid correction on 72.4% of placements.
A temporal-contrastive embedding learned from the robot's own depth stream, with no terrain label. It's evaluated as an alternative critic conditioner — it does not drive the headline transfer result.
n=20 per bar, Wilson 95% intervals. The ablated arm removes the expert mixture and the continuous foothold refinement.
Privileged-label routing (left) is near-permutation; every label-free router tried collapses toward uniform activation.
Front and back camera for each terrain segment of the chained waypoint course.
videos/slope_front.mp4
videos/slope_back.mp4
videos/height_field_front.mp4
videos/height_field_back.mp4
videos/random_spread_front.mp4
videos/random_spread_back.mp4
videos/up_down_stair_front.mp4
videos/up_down_stair_back.mp4
videos/hurdle_front.mp4
videos/hurdle_back.mp4
videos/chained_course_front.mp4
videos/chained_course_back.mp4
@inproceedings{anonymous2027terrain,
title = {Terrain Traversal Without Motion Priors: Expert-Routed
Humanoid Locomotion at Deployment Fidelity},
author = {Anonymous Authors},
booktitle = {IEEE International Conference on Robotics and
Automation (ICRA)},
year = {2027},
note = {Under review}
}