Abstract
Vision-Language-Action (VLA) models have advanced robotic navigation by unifying perception and planning, yet most systems operate in an open-loop manner, generating plans without awareness of locomotion feasibility or the ability to adapt when execution becomes unsafe. This disconnect leads to cascading failures on challenging terrains and surfaces where pre-planned trajectories become infeasible. We propose Thinking on Its Feet (TOIF), a three-tier hierarchical control framework that closes the loop between low-level proprioceptive sensing and high-level semantic path planning for terrain-aware humanoid robot navigation.
At the low level, an RL locomotion policy is augmented with a multi-detector anomaly system that evaluates terrain-specific feasibility and safety constraints using compact proprioceptive features. Detected anomalies are accumulated through leaky integrators and translated into structured natural-language feedback, enabling the high-level planner to replan over a semantic scene graph. At the mid level, a VLA model generates velocity commands that bridge high-level waypoints and low-level execution. We further introduce a risk-aware velocity modulation mechanism that proactively reduces speed based on real-time anomaly confidence.
We validate TOIF in Isaac Lab simulation and on a physical humanoid robot across a range of challenging terrains. Results demonstrate that TOIF substantially improves navigation success rate and safety over open-loop baselines, and that the proprioceptive anomaly detectors transfer effectively to real-world deployment.
System Overview
Figure 2. TOIF system architecture. Commands flow downward through scene-graph planning, VLA navigation, and RL locomotion. Proprioceptive feedback flows upward through anomaly detection, leaky integration, and semantic translation, enabling the planner to replan with terrain-aware context.
TOIF operates as a three-tier cascade:
- High level (on-demand): An LLM planner receives a semantic scene graph and navigation task, producing waypoints with natural-language instructions. When proprioceptive feedback triggers a replan, the blocked edge is added to a persistent list and the LLM selects an alternative route.
- Mid level (~1 Hz): A VLA model (NaVILA) converts the current instruction and egocentric image sequence into metric velocity commands (vx, vy, ω). Risk-aware velocity modulation continuously scales speed based on real-time anomaly confidence.
- Low level (50 Hz): A PPO-trained RL locomotion policy executes joint-position targets on the Unitree G1 (29 DoF). A multi-detector anomaly system monitors proprioceptive signals and accumulates evidence through terrain-specific leaky integrators to filter transient noise.
Key Contributions
- A three-tier hierarchical architecture with bidirectional information flow that couples downward command propagation with upward proprioceptive feedback via natural-language terrain context.
- A risk-aware terrain adaptation mechanism combining sigmoid-based velocity modulation with leaky-integrator replan triggers using terrain-specific time constants to balance responsiveness against false-positive suppression.
- A multi-detector proprioceptive anomaly system that learns compact, task-specific feature subsets for each terrain type (slippery, uneven, stepped), validated both in Isaac Lab simulation and on a real Unitree G1 with different sensor configurations.
Results
Feature Selection
Using simulation slip detection as a representative case, the full 107D proprioceptive feature set achieves only AUC = 0.66, because joint-state features introduce noise arising from the compensation dynamics of the RL policy. A compact 11D contact-feature set (foot forces, air times, projected gravity) achieves AUC = 0.94, confirming that contact signals encode terrain interaction before policy compensation attenuates the signal.
Navigation Performance
We evaluate over 20 trials per condition on two scenarios: S1 Warehouse Exit (detection-triggered semantic rerouting around a blocked doorway) and S2 Rocky Beach (continuous terrain adaptation on 18 m of uneven coastal terrain).
| Scenario | Method | SR (%) ↑ | Time (s) ↓ | Falls ↓ | Replans |
|---|---|---|---|---|---|
| S1: Warehouse Exit (T-Corridor) | |||||
| RRT* + Heuristic | 40 | 71.77 ± 12.98 | 0.5 ± 0.6 | — | |
| RRT* + DWA | 45 | 26.73 ± 8.58 | 0.0 ± 0.0 | — | |
| NaVILA E2E | 40 | 49.04 ± 7.17 | 0.2 ± 0.4 | — | |
| TOIF (Ours) | 65 | 37.34 ± 12.10 | 0.0 ± 0.0 | 0.80 ± 0.40 | |
| S2: Rocky Beach (Mixed Terrain) | |||||
| RRT* + Heuristic | 20 | 42.88 ± 12.76 | 0.8 ± 0.4 | — | |
| RRT* + DWA | 35 | 35.70 ± 2.08 | 0.60 ± 0.52 | — | |
| NaVILA E2E | 35 | 68.09 ± 10.20 | 0.65 ± 0.49 | — | |
| TOIF (Ours) | 60 | 84.81 ± 16.63 | 0.40 ± 0.52 | 0.50 ± 0.53 | |
Table 1. Navigation performance comparison (20 trials per condition, mean ± std). Best results in bold.
Ablation Study
Each ablated variant removes exactly one module, isolating its marginal contribution to overall system performance.
| Scenario | Variant | SR (%) ↑ | Time (s) ↓ | Falls ↓ | Replans |
|---|---|---|---|---|---|
| S1: Warehouse Exit (T-Corridor) | |||||
| TOIF (Full) | 65 | 37.0 ± 12.0 | 0.0 ± 0.0 | 0.8 ± 0.4 | |
| w/o Velocity Modulation | 50 | 39.0 ± 6.9 | 0.0 ± 0.0 | 0.9 ± 0.3 | |
| w/o Anomaly Detection | 20 | 20.5 ± 1.5 | 0.3 ± 0.5 | — | |
| w/o High-Level Planning | 15 | 44.2 ± 9.8 | 0.1 ± 0.3 | — | |
| S2: Rocky Beach (Mixed Terrain) | |||||
| TOIF (Full) | 60 | 84.81 ± 16.63 | 0.40 ± 0.52 | 0.50 ± 0.53 | |
| w/o Velocity Modulation | 50 | 67.81 ± 10.23 | 0.50 ± 0.50 | 0.45 ± 0.51 | |
| w/o Anomaly Detection | 25 | 64.62 ± 11.53 | 0.75 ± 0.44 | — | |
| w/o High-Level Planning | 35 | 63.20 ± 8.85 | 0.65 ± 0.49 | — | |
Table 2. Ablation study (20 trials per condition). Each row removes one module from the full system.
Real-World Validation
We deploy the anomaly detection component on a physical Unitree G1 and evaluate it across three laboratory terrain types: uneven terrain (speed bumps), stepped terrain (40 cm-high toolboxes), and slippery terrain (wooden boards inclined at 30–45°). In contrast to simulation, the real-world detectors are trained directly on manually annotated physical robot data, with an 80/20 temporal split to prevent data leakage. Supervised classifiers (AdaBoost, XGBoost) outperform the one-class methods used in simulation, as real-world nominal data exhibits higher variance. We then deploy the complete three-tier pipeline on the physical G1 in a laboratory navigation task.
Experiment Videos with Live Anomaly Scores
Uneven Terrain (Rough): Unitree G1 traversing speed bumps. AdaBoost detector (AUC = 0.874), threshold τ = −0.250.
Figure 3. Real-world anomaly detection on a physical Unitree G1 across three terrain types. Each row shows robot snapshots (top) and temporal anomaly scores (bottom). (a) Uneven terrain (τ = −0.250); (b) slippery terrain (τ = 0.080); (c) stepped terrain (τ = 0.040). Bold curves: smoothed scores; light traces: raw output; dashed lines: F1-optimal thresholds.
Figure 4. ROC curves of candidate models for real-world uneven-terrain anomaly detection, trained on physical Unitree G1 data (best: AdaBoost, AUC = 0.874).
| Terrain | Model | Dim | AUC | Key Features |
|---|---|---|---|---|
| Uneven | AdaBoost | 4D | 0.874 | torso sway, hip torque |
| Stepped | XGBoost | 1D | 0.942 | IMU yaw orientation |
| Slippery | LOF | 4D | 0.813 | ankle pos, trunk incl. |
Table 3. Real-world anomaly detectors (Unitree G1).
Full-Pipeline Deployment
The robot is instructed to reach an exit sign, while a wooden pallet on a hand truck blocks the direct route. Upon encountering the obstacle, the proprioceptive anomaly detectors trigger replanning, and the planner reroutes the robot along an alternative path to the exit. The robot reached the exit in 3 out of 5 trials, in line with the simulation success rates (60–65%); the complete runs are shown in the video at the top of this page.
Figure 5. Full-pipeline deployment on the physical Unitree G1. The robot walks toward the exit sign (1); a wooden pallet blocks the direct route and an anomaly is flagged (2–3); replanning turns the robot toward the alternative route (4), which it follows to the exit (5–6). Green/red borders: normal traversal / detection–replanning phase; yellow/green arrows: initial/replanned route.
Gallery
Motivation: TOIF closes the open-loop gap in VLA-based humanoid navigation.
Three-tier architecture with bidirectional information flow.
S2 Rocky Beach: 18 m coastal terrain requiring continuous terrain adaptation.
Real-world anomaly detection on a physical Unitree G1 across three terrain types.
BibTeX
@inproceedings{toif2026,
title = {Thinking on Its Feet: Terrain-Aware Hierarchical Control
with Proprioceptive Feedback for Humanoid Navigation},
author = {Yuhai Wang and Chenghao Wang and Xianyao Li and Zening Teng and Xiao Hu and Rongxuan Zhou and Alireza Ramezani and Jing Du and Yang Ye},
booktitle = {IEEE-RAS International Conference on Humanoid Robots (Humanoids)},
year = {2026},
url = {https://yuhaiw.github.io/TOIF/}
}