Bench2Drive-RLInfra v0.0.1 Checkpoints
This repository provides the v0.0.1 checkpoints and training configurations for Bench2Drive-RLInfra.
Please refer to Bench2Drive-RLInfra for project details, setup/training instructions, and how to use these checkpoints for closed-loop evaluation in CARLA.
Models
| Model | Evaluation weights | Training steps / updates | Rollout / collection size |
|---|---|---|---|
| PPO | PPO/policy.pth |
100M environment steps | 8,192 transitions per rollout |
| A2C | A2C/policy.pth |
100M environment steps | 8,192 transitions per rollout |
| SAC | SAC/policy.pth |
10M environment steps | 256 transitions per collection interval |
| TD3 | TD3/policy.pth |
10M environment steps | 256 transitions per collection interval |
| DrivePi0 | DrivePi0/rl_finetune_update_000008_export.pt |
8 post-training updates | 32 rollouts per update |
| MindDrive | MindDrive/rl_finetune_update_000004_export.pth |
4 post-training updates | 32 rollouts per update |
PPO and A2C benefit noticeably from scaling up training beyond the original setup. SAC and TD3 show more modest gains from longer training in this setup, so we train them for 10M environment steps.
For SAC and TD3, we release the checkpoint with the highest success rate on the training routes. This metric often declines later in training, so the selected checkpoint is usually earlier than the final training step. The 10M steps reported above refer to the full training run.
See each model's config.yaml for the full training configuration.