|
Download README.md from SUPER321/jevflash-doom-basic-0.6b: direct link, hf CLI and curl.
- Browser
- Download file 2.16 kB
-
https://huggingface.co/SUPER321/jevflash-doom-basic-0.6b/resolve/main/README.md
- Command line
-
hf download hf://SUPER321/jevflash-doom-basic-0.6b/README.md
-
curl -L -o README.md https://huggingface.co/SUPER321/jevflash-doom-basic-0.6b/resolve/main/README.md
2.16 kB
| license: apache-2.0 | |
| base_model: Qwen/Qwen3-0.6B-Base | |
| tags: | |
| - vizdoom | |
| - reinforcement-learning-adjacent | |
| - decision-model | |
| - behavior-cloning | |
| pipeline_tag: text-generation | |
| # JevFlash: Doom Basic Decision Model (0.6B) | |
| A from-scratch replica of [NanoJev](https://github.com/TianyuCodings/NanoJev)'s | |
| non-generative "decision model" architecture β instead of generating text, | |
| it scores a fixed set of candidate actions (`left` / `right` / `shoot` / | |
| `noop`) given a game state, in a single forward pass β fine-tuned on | |
| Qwen3-0.6B-Base to play ViZDoom's **basic** scenario (aim and shoot a | |
| stationary monster). | |
| **Closed-loop result: 65% episode success rate**, beating a uniform-random | |
| baseline (45%) β the first of 10 training attempts to do so. Full, | |
| honestly-documented history of everything that didn't work along the way | |
| (degenerate collapses, a checkpoint that looked great on paper but turned | |
| out to be a memorized shortcut, and the eventual discovery that this | |
| specific failure mode was a bad random seed) is in the project repo: | |
| **Code, data, and full write-up**: https://github.com/themaker00001/JevFlash | |
| ## What this model does | |
| Given a text description of the game state (health, ammo, position, and | |
| visible target bounding box) and a question ("pick the best action"), it | |
| outputs a probability score for each of up to 4 candidate actions in one | |
| non-autoregressive forward pass. No chain-of-thought, no token generation | |
| at inference time. | |
| ## Training data | |
| ~2,300 decision examples collected via epsilon-greedy exploration (10% of | |
| executed actions were random, always labeled with the correct/heuristic | |
| action) against a simple proportional aim-and-shoot heuristic β the same | |
| technique the original NanoJev project used with a real pretrained RL | |
| policy, adapted here to a much simpler, self-authored expert. | |
| ## Files | |
| - `best.safetensors` β model weights (backbone + decision head) | |
| - `tokenizer/`, `backbone_config/` β tokenizer and Qwen3 backbone config | |
| - `summary.json`, `train_log.json`, `timing.json`, `target_audit.json` β training run artifacts | |
| ## License | |
| Apache-2.0 (inherited from the Qwen3-0.6B-Base backbone). | |