Bench2Drive-RLInfra v0.0.1 Checkpoints

This repository provides the v0.0.1 checkpoints and training configurations for Bench2Drive-RLInfra.

Please refer to Bench2Drive-RLInfra for project details, setup/training instructions, and how to use these checkpoints for closed-loop evaluation in CARLA.

Models

Model Evaluation weights Training steps / updates Rollout / collection size
PPO PPO/policy.pth 100M environment steps 8,192 transitions per rollout
A2C A2C/policy.pth 100M environment steps 8,192 transitions per rollout
SAC SAC/policy.pth 10M environment steps 256 transitions per collection interval
TD3 TD3/policy.pth 10M environment steps 256 transitions per collection interval
DrivePi0 DrivePi0/rl_finetune_update_000008_export.pt 8 post-training updates 32 rollouts per update
MindDrive MindDrive/rl_finetune_update_000004_export.pth 4 post-training updates 32 rollouts per update

PPO and A2C benefit noticeably from scaling up training beyond the original setup. SAC and TD3 show more modest gains from longer training in this setup, so we train them for 10M environment steps.

For SAC and TD3, we release the checkpoint with the highest success rate on the training routes. This metric often declines later in training, so the selected checkpoint is usually earlier than the final training step. The 10M steps reported above refer to the full training run.

See each model's config.yaml for the full training configuration.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support