Asimov Flex: base walk v0.7
Velocity-commanded walking policy for Asimov 1, trained in Isaac Lab with PPO. It is a base walk meant for later skills to build on: forward, backward and sideways walking, turning while walking or on the spot, and standing still.
Version 0.7 brings back the top speed. Walk v0.6 walked slower than commanded at the top speed because its steps were shorter: the gait clock sets how many steps it takes per second, and the speed-tracking reward barely changed close to the command. Version 0.7 adds a reward for forward speed relative to the command. It is only active on even ground, which is judged from the ground height during training only. It is walk v0.6 trained further for 1,000 iterations in one stage, with the full push rate from the start and a little more practice on stairs going up. The policy itself never sees the ground: it uses only the robot's joint and body sensors, so no camera or lidar is needed. The inputs, the outputs, the gait clock and the timing are the same as in v0.6. Earlier training for rough ground and stairs, for balance (pushes from every direction, steady forces), for differences between simulation and hardware (IMU tilt bias, motor strength and gain changes, extra mass, joint friction) and for the robot's timing (about 35 ms command delay at 50 Hz) is kept.
Tested in simulation only (Isaac Lab, and MuJoCo with the public Asimov 1 model). It has not been run on the real robot.
policy.onnxis the exported policy for inference, producing joint-position actions.agent.yamlrecords the actor-critic and AMP training settings, including the motion data configuration.env.yamlrecords the simulation and locomotion task settings, including observations, actions, commands, and rewards.- Training and evaluation code: menloresearch/isaac_asimov.
Inputs and outputs
- Control rate 50 Hz. Actions: 23 joint-position offsets,
target = default_pose + 0.25 * action, no filter. - 80 inputs: the same 78 as the standard Asimov 1 walking policy, in the same order, followed by a 2-value gait clock.
- Gait clock:
(sin(2 pi phi), cos(2 pi phi)), withphiin [0, 1).- Each policy step while walking:
phi += elapsed / period, whereelapsedis the measured real time since the previous policy step (0.02 s at 50 Hz). Use the measured time, not a fixed 0.02 s, if the loop runs slower. periodis 0.8 s at command speeds up to 0.2 m/s and 0.7 s at 0.8 m/s and above, linear in between (planar command speed).- When
|v_xy| + |w_z| <= 0.1the robot stands: the clock outputs(0, 0)andphistays at 0. - Walking starts from
phi = 0(both feet down); the right foot lifts first.
- Each policy step while walking:
Results in simulation
Measured at the robot's timing (50 Hz, 35 ms command delay). Terrain results in Isaac are from new random seeds with both versions on the same seeds, 192 robots per case. MuJoCo uses 128 seeds per case. Falls are the share of test episodes; pushes are velocity impulses, worst of 4 directions.
| Test | v0.7 | v0.6 |
|---|---|---|
| Forward speed at a 0.8 m/s command (Isaac / MuJoCo) | 0.702 / 0.699 m/s | 0.665 / 0.663 m/s |
| Forward speed tracking error at 0.8 m/s (Isaac) | 12.2% | 16.9% |
| Falls on flat ground, all walking tests (256 robots each) | 0% | 0% |
| Stairs up, 6 / 8 / 10 / 12 cm steps (Isaac) | 0.0% / 1.0% / 13.0% / 97.4% falls | 1.0% / 0.5% / 12.0% / 96.4% |
| 10 cm steps on 0.35 m treads, up / down (Isaac) | 31.2% / 1.6% falls | 37.5% / 7.8% |
| Rough ground, ±6 / ±8 cm (Isaac) | 7.3% / 96.4% falls | 6.8% / 95.3% |
| Rough ground, ±6 cm (MuJoCo) | 26.6% falls | 24.2% |
| Stairs up 10 cm (MuJoCo) | 56.2% falls | 84.4% |
| Stairs up 6 cm, MuJoCo, 5 contact settings pooled (0.8 m / 0.35 m treads) | 1.7% / 8.8% falls | 4.7% / 13.3% |
| Standing push 0.6 / 0.8 m/s (MuJoCo) | 0% / 0% falls | 0% / 0% |
| Standing push 0.6 / 0.8 / 1.0 m/s (Isaac) | 1.6% / 0.8% / 3.9% falls | 0.8% / 7.0% / not run |
| Action jitter, mean of 11 tests / walking at 0.5 m/s | 0.093 / 0.084 | 0.091 / 0.085 |
Starting to walk and turn from a standstill: no falls in 1,024 robots.
Known limits:
- Walking sideways has more action jitter than v0.6 (0.102 vs 0.085).
- In MuJoCo it walks at 0.699 m/s at a 0.8 m/s command, just at the 0.70 m/s target.
- 12 cm steps going up still fail almost every time, and rough ground with ±8 cm bumps is still too much.
- It has not been run on the real robot yet.