DriveJev-4B
DriveJev 1.0: an open-source, efficient System I decision model for autonomous driving
Code | Live demo | Evaluation
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
DriveJev 1.0 at the wheel in JevPilot with full traffic: an oncoming left-turner, a 4-way stop, a red-light runner after our green, a pedestrian behind a parked car, pedestrians at a turn exit and cut-ins on the interstate. Every drive reached its destination with no collision and no violation.
DriveJev is a System I driving model. At every decision it looks at three camera frames and a short description of the situation, scores the behaviours the car can carry out at that instant (cruise, stop at the line, hold, start, turn, yield, brake) and returns a probability for each, all in one forward pass. A local executor turns the chosen behaviour into steering and speed. This repository is the complete model in one folder: the vision-language backbone and the DriveJev decision head.
- Driving ability. Rule-following, steady driving on city streets, small-town roads and the interstate, through junctions with traffic lights and stop signs, among cars and pedestrians. On held-out closed-loop tests DriveJev 1.0 drives 28 of 35 interaction episodes and 25 of 28 base episodes cleanly and meets 241 of 242 scripted hazards without contact.
- Inference speed. 111 ms per decision (9 Hz) on one RTX 5090, with no text or trajectory decoding.
- Extensibility. Fully open and local: the code, these weights, the JevPilot closed-loop harness and the demo can be inspected, retrained and fine-tuned on new scenarios.
Model
| Inputs | wide front camera 640Γ384 at t-0.5 s and t, tele camera 384Γ224 (15Β° vertical field of view) at t, a whitelisted JSON state (ego, navigation, current behaviour, perception summary) and 2 to 8 offered behaviours; about 1,000 prompt tokens |
| Backbone | 4.5 B-parameter vision-language model, BF16, one forward pass per decision |
| Decision head | pointer_mlp: q = MLP([LN(h_dec); LN(h_state)]), k = MLP(LN(h_cand)), logit = qΒ·k/β256; 4.2 M parameters, FP32 |
| Output | a probability for every offered behaviour; the executed behaviour is the argmax |
| Behaviours | keep_route_cruise, stop_at_line, hold_stop, proceed_route, route_turn, yield_agent, continue_current, emergency_brake |
| Speed | 111 ms per decision at batch size 1 on one RTX 5090 (about 11 GB of GPU memory); decisions at 4 Hz in closed loop |
The head reads the hidden states at the Decision: token, at the end of the state block and at the last token of each behaviour. The prompt compiler accepts only whitelisted state fields, so signal colours, traffic rules and other agents' plans never reach the model: it reads the light from the camera. Behaviour IDs never appear in the prompt and the behaviours are put in a canonical order.
Training (summary). Soft-target cross-entropy against an observable privileged teacher, a policy that reads the traffic rules and rolls the world forward for every behaviour using only the road users the student's perception has seen, so every label can be explained from the model's inputs. Data: teacher-driven episodes in the whole JevPilot world (5 % of decision slots perturbed to visit recoverable mistakes), hazard-dense episodes, interaction episodes covering all nine hazard and interaction kinds, and DAgger rounds in which DriveJev drives and the teacher relabels the states it visits; decisions changed by an interaction conflict are up-weighted. 119 k labels in total; model selection used validation seeds only.
Results
| Test suite | Policy | Success β | Collisions β | Violations β | Route completed β |
|---|---|---|---|---|---|
| Interaction (35 episodes) | Privileged teacher (upper bound) | 97% [91, 100] (34/35) | 0 | 0 | 99% |
| DriveJev 1.0 | 80% [66, 91] (28/35) | 1 | 4 | 97% | |
| Base (28 episodes) | Privileged teacher (upper bound) | 96% [89, 100] (27/28) | 0 | 0 | 99% |
| DriveJev 1.0 | 89% [79, 100] (25/28) | 0 | 2 | 100% |
Closed-loop drives in JevPilot on seeds never used for training or model selection: the world waits for each 4 Hz decision, AEB is off, and success means reaching the destination with no collision and no violation (red light, amber that could still be stopped for, unserved stop sign); 95 % bootstrap intervals in brackets. None of the 16 oncoming platoons and left-turners, 25 pedestrians hidden behind parked cars, 24 cut-ins, 6 late red-light runners or 9 four-way-stop contentions of the interaction suite ends in contact. Full tables and metric definitions: docs/evaluation.md.
Files
DriveJev-4B/
βββ drivejev_config.json prompt, camera and behaviour contract, decision-head configuration
βββ head.safetensors decision head (4.2 M parameters, FP32)
βββ backbone/ vision-language backbone: config, tokenizer, weights, NOTICE.md and Apache-2.0 LICENSE
Usage
The model is loaded with the DriveJev code. A GPU with 16 GB or more is recommended.
git clone --recursive https://github.com/benmagnifico/DriveJev.git && cd DriveJev
(cd third_party/jevpilot && npm ci) # the JevPilot driving world
pip install -r requirements.txt
hf download benmagnifico/DriveJev-4B --local-dir checkpoints/DriveJev-4B
python examples/predict.py --model checkpoints/DriveJev-4B # six bundled JevPilot decisions
from drivejev import DrivingPolicy
policy = DrivingPolicy.from_pretrained("benmagnifico/DriveJev-4B", device="cuda:0")
record = {
"student_obs": {...}, # ego / nav / maneuver / recent_actions / traffic
"candidates": [{"candidate_id": "stop_at_line", "action_type": "stop_at_line", "target_id": "j2-1",
"speed_profile": "line_stop", "description": "Approach and stop ..."}, ...],
"images": [{"camera": "front", "relative_time": -0.5, "path": "front_t-0.5.png"},
{"camera": "front", "relative_time": 0, "path": "front_t0.png"},
{"camera": "front_tele", "relative_time": 0, "path": "tele_t0.png"}],
}
out = policy.predict(record)
print(out["candidate_id"], out["probabilities"]) # chosen behaviour + a probability for every offered one
serve/serve.py --model checkpoints/DriveJev-4B exposes the same call over HTTP for the closed-loop harness and the live JevPilot demo (DRIVEJEV_MODEL=checkpoints/DriveJev-4B bash demo/start.sh). The state schema, behaviours and executor are documented in docs/simulator.md.
Intended use
- Research only. Not for real vehicles. DriveJev was trained and evaluated in the JevPilot simulator; its inputs (rendered cameras, a perception summary computed by the simulator adapter) and its executor are part of that setup.
- It chooses between behaviours offered by the executor; the executor steers and controls speed.
License
The repository is released under the Apache License 2.0. backbone/ keeps the Apache 2.0 license and notice of the Qwen-Drive-1.0-4B checkpoint it comes from (backbone/LICENSE, backbone/NOTICE.md). The DriveJev code on GitHub is MIT-licensed.
Acknowledgements
DriveJev builds on Qwen-Drive-1.0 (the vision-language backbone and its image processing), JevPilot (the driving world, physics, traffic and rules), and the decision-model idea of TypeSafe's Jev and its open reconstruction Kev.
Citation
If you find DriveJev helpful, please cite it. The code is archived on Zenodo; the DOI 10.5281/zenodo.23118991 covers all versions and always resolves to the latest one.
@misc{li2026drivejev,
title = {DriveJev: Real-Time Driving Behaviour Selection with a Vision-Language Decision Model},
author = {Jingguang Li and Kailang Ma and Zuyi Guo and Yebo Wu and Heye Huang},
year = {2026},
doi = {10.5281/zenodo.23118991},
howpublished = {\url{https://github.com/benmagnifico/DriveJev}}
}





