VibeGame-8B

I wanted to describe a little game and get a project I could open, play, and change. VibeGame is that experiment: a request goes in; project.godot, a scene, and GDScript come out together as JSON.

Measured results: 6/500 imported (1.2%); 5/500 ran cleanly for 60 seconds (1.0%). This is a small research prototype. A clean run does not prove the requested mechanics work, and selected successful clips do not represent the model's usual output. The results below include every failure.

This first version focuses on small 2D games with shapes drawn in code. “8B” is the project name; the base is Qwen2.5-Coder-7B-Instruct, with 7.61 billion parameters. Training used Unsloth QLoRA on a free Colab Tesla T4. Evaluation began on the free tier; after its GPU allowance ran out, the project owner purchased Colab Pro. The earlier T4 experiment stopped at 72 recorded results: 64 generations, eight interrupted attempts, and one compile/clean-run success. A separate full A100 pass was predeclared to reduce completion time. Its hardware and versions must be reported with the final measurements; results from the two runs are never pooled.

The one-minute idea

“Prompt → playable game in 60 seconds” is the target. The measurements below show where this release lands.

Watch the 60-second walkthrough. Sixty seconds is the edited video's duration, not a generation-speed claim. This video is a success-only showcase: exactly the first three clean results in frozen prompt order. It is not representative of the model. All three recordings use untouched outputs, and the full 500-prompt results and actual times stay visible. The original benchmark sample slots 0, 200, and 400 remain unchanged in the evidence. Gameplay recordings use fixed timesteps without player input, so they cannot prove responsive controls or real-time generation speed.

Try it locally

Install Godot 4.5.1 and Ollama. Download the Q4_K_M file and Modelfile from the GGUF repository, then use the project source:

# Point FROM in Modelfile at the downloaded GGUF.
ollama create vibegame -f Modelfile
python scripts/generate.py --backend ollama --model vibegame \
  --prompt "Make a tiny game where I catch falling stars and avoid meteors" \
  --output my-game.json
python scripts/extract_project.py my-game.json --output my-game

The generation command writes JSON with a files mapping. The extractor checks that format and writes into an empty project directory. Inspect the scripts, open project.godot in Godot, and press Play. Extraction alone does not check compilation; the included Linux validation harness tests import and runtime behavior in isolation.

LoRA adapter · Merged FP16 weights · Training data

What it learned from

The Codex assistant authored 100 distinct complete-project candidates during this build, plus nine earlier seeds, with original game logic and geometric visuals. Linux admission accepted 108 of 109, producing 97 training examples and 11 development examples. One failed and was discarded. Every admitted project imported and ran cleanly for 60 seconds. Those are data checks, separate from the model scores below.

The 180 MIT GDScript reference files are stored separately and were not loaded directly into training. Collection accepts only MIT, Apache-2.0, or CC0 code, pins repository commits, and keeps file origins and license notices. It excludes assets, copyleft code, and conflicting or ambiguous notices. Third-party source keeps its original license.

The Stack v2 contains GDScript but was excluded because its gated terms and bulk-download agreement need separate attention. Data notes and provenance also document the optional local 14B teacher pilot and its rejected attempts.

Results, including the misses

These scores evaluate the LoRA adapter on its 4-bit base, on NVIDIA A100-SXM4-40GB. The merged FP16 and GGUF files have not been separately benchmarked. Values are filled only from a completed run. This table reports the separate full A100 pass on Colab Pro. The partial T4 experiment is preserved under evaluation-t4-partial/; its results are not pooled with this pass.

Check Result
Frozen prompts 500
Strict project format and zero-error import 6/500 (1.2%)
Clean 60-second run 5/500 (1.0%)
Warm prompt-to-launch median, successful launches only 282.75 seconds
Successful launches measured within 60 seconds 0/500

Both success rates include all 500 prompts. Each gets one generation, with no repairs or retries. Invalid JSON, parse errors, rejected APIs, early exits, and script errors count as failures; infrastructure failures leave the evaluation incomplete. All outputs and logs accompany the summary.

The warm timer excludes model loading. Each prompt receives the full batch generation time, plus its queue, project import, and startup time. We never divide latency by batch size. Interrupted or resumed timings are excluded; the clean-run check then lasts another 60 seconds. In the separate partial T4 experiment, the median among measured successful launches was 1,394.69 seconds, with none within 60 seconds. CPU export overlapped parts of that T4 experiment and its timings include that contention. The A100 protocol, training-prompt capacity checks and budget receipts accompany its evaluation; those measurements are separate from model quality scores.

The benchmark crosses 25 game families with 20 rule variants. Its 500 frozen prompts never enter training, development loss, or checkpoint selection. These are compositional variants, and the base model's pretraining data is outside our contamination audit.

Where it breaks

A clean run does not establish fun, correct controls, prompt fidelity, or even a useful visible scene. Expect broken collision rules, unreachable goals, blank screens, and wrong Godot APIs. The small dataset and 4096-token context limit complex games, 3D, multiplayer, custom assets, and long inventories. Human playtesting matters.

Inference runs locally. Publication creates no paid endpoint. Provider mappings are checked during release; account provider switches and spending limits are verified separately. inference: false only hides the widget.

Help make it better

Try a prompt and share the exact output, Godot version, and what happened when you played. Distinct permissively licensed examples, playtest notes, and checks for blank-but-running scenes would help. Keep blind prompts out of training and include source licenses with contributions.

Downloads last month
-
Safetensors
Model size
8B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for khursheed/VibeGame-8B

Base model

Qwen/Qwen2.5-7B
Finetuned
(471)
this model