DriveJev-4B

DriveJev 1.0: an open-source, efficient System I decision model for autonomous driving

Code  |  Live demo  |  Evaluation

DOI 10.5281/zenodo.23118991

DriveJev holding at the line on green while an oncoming car turns left across its path DriveJev waiting at a 4-way stop while three cars take their turn DriveJev braking at the line for a car that runs its red just after the ego's green
DriveJev yielding to a pedestrian stepping out from behind a parked car DriveJev stopping mid-turn for two pedestrians on the exit crosswalk DriveJev yielding to cars that cut in on the interstate

DriveJev 1.0 at the wheel in JevPilot with full traffic: an oncoming left-turner, a 4-way stop, a red-light runner after our green, a pedestrian behind a parked car, pedestrians at a turn exit and cut-ins on the interstate. Every drive reached its destination with no collision and no violation.

DriveJev is a System I driving model. At every decision it looks at three camera frames and a short description of the situation, scores the behaviours the car can carry out at that instant (cruise, stop at the line, hold, start, turn, yield, brake) and returns a probability for each, all in one forward pass. A local executor turns the chosen behaviour into steering and speed. This repository is the complete model in one folder: the vision-language backbone and the DriveJev decision head.

  • Driving ability. Rule-following, steady driving on city streets, small-town roads and the interstate, through junctions with traffic lights and stop signs, among cars and pedestrians. On held-out closed-loop tests DriveJev 1.0 drives 28 of 35 interaction episodes and 25 of 28 base episodes cleanly and meets 241 of 242 scripted hazards without contact.
  • Inference speed. 111 ms per decision (9 Hz) on one RTX 5090, with no text or trajectory decoding.
  • Extensibility. Fully open and local: the code, these weights, the JevPilot closed-loop harness and the demo can be inspected, retrained and fine-tuned on new scenarios.

Three DriveJev 1.0 decisions: front camera with tele inset and the probability of every offered behaviour

DriveJev overview: inputs, vision-language model, decision head, semantic executor and the JevPilot closed loop

Model

Inputs wide front camera 640Γ—384 at t-0.5 s and t, tele camera 384Γ—224 (15Β° vertical field of view) at t, a whitelisted JSON state (ego, navigation, current behaviour, perception summary) and 2 to 8 offered behaviours; about 1,000 prompt tokens
Backbone 4.5 B-parameter vision-language model, BF16, one forward pass per decision
Decision head pointer_mlp: q = MLP([LN(h_dec); LN(h_state)]), k = MLP(LN(h_cand)), logit = q·k/√256; 4.2 M parameters, FP32
Output a probability for every offered behaviour; the executed behaviour is the argmax
Behaviours keep_route_cruise, stop_at_line, hold_stop, proceed_route, route_turn, yield_agent, continue_current, emergency_brake
Speed 111 ms per decision at batch size 1 on one RTX 5090 (about 11 GB of GPU memory); decisions at 4 Hz in closed loop

The head reads the hidden states at the Decision: token, at the end of the state block and at the last token of each behaviour. The prompt compiler accepts only whitelisted state fields, so signal colours, traffic rules and other agents' plans never reach the model: it reads the light from the camera. Behaviour IDs never appear in the prompt and the behaviours are put in a canonical order.

Training (summary). Soft-target cross-entropy against an observable privileged teacher, a policy that reads the traffic rules and rolls the world forward for every behaviour using only the road users the student's perception has seen, so every label can be explained from the model's inputs. Data: teacher-driven episodes in the whole JevPilot world (5 % of decision slots perturbed to visit recoverable mistakes), hazard-dense episodes, interaction episodes covering all nine hazard and interaction kinds, and DAgger rounds in which DriveJev drives and the teacher relabels the states it visits; decisions changed by an interaction conflict are up-weighted. 119 k labels in total; model selection used validation seeds only.

Results

DriveJev 1.0 results: closed-loop success, hazards met without contact and decision speed

Test suite Policy Success ↑ Collisions ↓ Violations ↓ Route completed ↑
Interaction (35 episodes) Privileged teacher (upper bound) 97% [91, 100] (34/35) 0 0 99%
DriveJev 1.0 80% [66, 91] (28/35) 1 4 97%
Base (28 episodes) Privileged teacher (upper bound) 96% [89, 100] (27/28) 0 0 99%
DriveJev 1.0 89% [79, 100] (25/28) 0 2 100%

Closed-loop drives in JevPilot on seeds never used for training or model selection: the world waits for each 4 Hz decision, AEB is off, and success means reaching the destination with no collision and no violation (red light, amber that could still be stopped for, unserved stop sign); 95 % bootstrap intervals in brackets. None of the 16 oncoming platoons and left-turners, 25 pedestrians hidden behind parked cars, 24 cut-ins, 6 late red-light runners or 9 four-way-stop contentions of the interaction suite ends in contact. Full tables and metric definitions: docs/evaluation.md.

Files

DriveJev-4B/
β”œβ”€β”€ drivejev_config.json   prompt, camera and behaviour contract, decision-head configuration
β”œβ”€β”€ head.safetensors       decision head (4.2 M parameters, FP32)
└── backbone/              vision-language backbone: config, tokenizer, weights, NOTICE.md and Apache-2.0 LICENSE

Usage

The model is loaded with the DriveJev code. A GPU with 16 GB or more is recommended.

git clone --recursive https://github.com/benmagnifico/DriveJev.git && cd DriveJev
(cd third_party/jevpilot && npm ci)                          # the JevPilot driving world
pip install -r requirements.txt
hf download benmagnifico/DriveJev-4B --local-dir checkpoints/DriveJev-4B
python examples/predict.py --model checkpoints/DriveJev-4B   # six bundled JevPilot decisions
from drivejev import DrivingPolicy

policy = DrivingPolicy.from_pretrained("benmagnifico/DriveJev-4B", device="cuda:0")
record = {
    "student_obs": {...},                     # ego / nav / maneuver / recent_actions / traffic
    "candidates": [{"candidate_id": "stop_at_line", "action_type": "stop_at_line", "target_id": "j2-1",
                    "speed_profile": "line_stop", "description": "Approach and stop ..."}, ...],
    "images": [{"camera": "front", "relative_time": -0.5, "path": "front_t-0.5.png"},
               {"camera": "front", "relative_time": 0, "path": "front_t0.png"},
               {"camera": "front_tele", "relative_time": 0, "path": "tele_t0.png"}],
}
out = policy.predict(record)
print(out["candidate_id"], out["probabilities"])   # chosen behaviour + a probability for every offered one

serve/serve.py --model checkpoints/DriveJev-4B exposes the same call over HTTP for the closed-loop harness and the live JevPilot demo (DRIVEJEV_MODEL=checkpoints/DriveJev-4B bash demo/start.sh). The state schema, behaviours and executor are documented in docs/simulator.md.

Intended use

  • Research only. Not for real vehicles. DriveJev was trained and evaluated in the JevPilot simulator; its inputs (rendered cameras, a perception summary computed by the simulator adapter) and its executor are part of that setup.
  • It chooses between behaviours offered by the executor; the executor steers and controls speed.

License

The repository is released under the Apache License 2.0. backbone/ keeps the Apache 2.0 license and notice of the Qwen-Drive-1.0-4B checkpoint it comes from (backbone/LICENSE, backbone/NOTICE.md). The DriveJev code on GitHub is MIT-licensed.

Acknowledgements

DriveJev builds on Qwen-Drive-1.0 (the vision-language backbone and its image processing), JevPilot (the driving world, physics, traffic and rules), and the decision-model idea of TypeSafe's Jev and its open reconstruction Kev.

Citation

If you find DriveJev helpful, please cite it. The code is archived on Zenodo; the DOI 10.5281/zenodo.23118991 covers all versions and always resolves to the latest one.

@misc{li2026drivejev,
  title  = {DriveJev: Real-Time Driving Behaviour Selection with a Vision-Language Decision Model},
  author = {Jingguang Li and Kailang Ma and Zuyi Guo and Yebo Wu and Heye Huang},
  year   = {2026},
  doi    = {10.5281/zenodo.23118991},
  howpublished = {\url{https://github.com/benmagnifico/DriveJev}}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for benmagnifico/DriveJev-4B

Finetuned
Qwen/Qwen3.5-4B
Finetuned
(5)
this model