Download README.md from SUPER321/jevflash-doom-basic-0.6b: direct link, hf CLI and curl.
- Browser
- Download file 2.16 kB
-
https://huggingface.co/SUPER321/jevflash-doom-basic-0.6b/resolve/main/README.md
- Command line
-
hf download hf://SUPER321/jevflash-doom-basic-0.6b/README.md
-
curl -L -o README.md https://huggingface.co/SUPER321/jevflash-doom-basic-0.6b/resolve/main/README.md
license: apache-2.0
base_model: Qwen/Qwen3-0.6B-Base
tags:
- vizdoom
- reinforcement-learning-adjacent
- decision-model
- behavior-cloning
pipeline_tag: text-generation
JevFlash: Doom Basic Decision Model (0.6B)
A from-scratch replica of NanoJev's
non-generative "decision model" architecture — instead of generating text,
it scores a fixed set of candidate actions (left / right / shoot /
noop) given a game state, in a single forward pass — fine-tuned on
Qwen3-0.6B-Base to play ViZDoom's basic scenario (aim and shoot a
stationary monster).
Closed-loop result: 65% episode success rate, beating a uniform-random baseline (45%) — the first of 10 training attempts to do so. Full, honestly-documented history of everything that didn't work along the way (degenerate collapses, a checkpoint that looked great on paper but turned out to be a memorized shortcut, and the eventual discovery that this specific failure mode was a bad random seed) is in the project repo:
Code, data, and full write-up: https://github.com/themaker00001/JevFlash
What this model does
Given a text description of the game state (health, ammo, position, and visible target bounding box) and a question ("pick the best action"), it outputs a probability score for each of up to 4 candidate actions in one non-autoregressive forward pass. No chain-of-thought, no token generation at inference time.
Training data
~2,300 decision examples collected via epsilon-greedy exploration (10% of executed actions were random, always labeled with the correct/heuristic action) against a simple proportional aim-and-shoot heuristic — the same technique the original NanoJev project used with a real pretrained RL policy, adapted here to a much simpler, self-authored expert.
Files
best.safetensors— model weights (backbone + decision head)tokenizer/,backbone_config/— tokenizer and Qwen3 backbone configsummary.json,train_log.json,timing.json,target_audit.json— training run artifacts
License
Apache-2.0 (inherited from the Qwen3-0.6B-Base backbone).