JevFlash: Doom Basic Decision Model (0.6B)
A from-scratch replica of NanoJev's
non-generative "decision model" architecture β instead of generating text,
it scores a fixed set of candidate actions (left / right / shoot /
noop) given a game state, in a single forward pass β fine-tuned on
Qwen3-0.6B-Base to play ViZDoom's basic scenario (aim and shoot a
stationary monster).
Closed-loop result: 65% episode success rate, beating a uniform-random baseline (45%) β the first of 10 training attempts to do so. Full, honestly-documented history of everything that didn't work along the way (degenerate collapses, a checkpoint that looked great on paper but turned out to be a memorized shortcut, and the eventual discovery that this specific failure mode was a bad random seed) is in the project repo:
Code, data, and full write-up: https://github.com/themaker00001/JevFlash
What this model does
Given a text description of the game state (health, ammo, position, and visible target bounding box) and a question ("pick the best action"), it outputs a probability score for each of up to 4 candidate actions in one non-autoregressive forward pass. No chain-of-thought, no token generation at inference time.
Training data
~2,300 decision examples collected via epsilon-greedy exploration (10% of executed actions were random, always labeled with the correct/heuristic action) against a simple proportional aim-and-shoot heuristic β the same technique the original NanoJev project used with a real pretrained RL policy, adapted here to a much simpler, self-authored expert.
Files
best.safetensorsβ model weights (backbone + decision head)tokenizer/,backbone_config/β tokenizer and Qwen3 backbone configsummary.json,train_log.json,timing.json,target_audit.jsonβ training run artifacts
License
Apache-2.0 (inherited from the Qwen3-0.6B-Base backbone).
- Downloads last month
- 177
Model tree for SUPER321/jevflash-doom-basic-0.6b
Base model
Qwen/Qwen3-0.6B-Base