MMMMoK9922's picture
Record public archive status
4d6a0cd verified
|
Raw History Blame Contribute Delete
1.04 kB
---
license: other
library_name: pytorch
tags:
- robotics
- reinforcement-learning
- rainbow-dqn
- speed-control
- metarobobench
---
# MetaRoboBench Speed-Control Heads
Offboarding archive of Rainbow DQN, CQL, and candidate-value heads trained for chunk-level Fast/Normal/Slow selection on top of Pi05 checkpoint `21650`.
The primary model is `rainbow/shared_all50_final/models/shared_all50.pt`:
- 50-task shared Rainbow head
- 100,491 decisions and 99,962 updates
- 1,941/2,500 successes = 77.64%
- Fast/Normal/Slow chunk shares = 38.6/32.3/29.1%
- Mean speed switches = about 4.23 per episode
The accelerated shared checkpoint is a distinct earlier checkpoint, not a duplicate. Per-task Rainbow, CQL, and candidate-value models are retained as baselines and negative-result artifacts.
Code reference:
- Repository: https://github.com/MMMM3202/MetaRoboBench
- Branch: `codex/vlm-rainbow-shared-20260821`
- Commit: `5b36567f0c6174ad82a3d4c669134f7b28acc945`
See `MODEL_MANIFEST.md` for the full inventory and result caveats.