Pac-Man Decision Tiny · 17,601 parameters

A small, newly trained PyTorch policy that chooses among legal Pac-Man directions. It learns from a five-second simulation/search teacher and runs locally on CPU. The safe tensor weight file is about 72 KB.

Interactive experiment explorer · Training data · Qwen LoRA comparison

What the first experiment found

Same 4,096 training examples and 512 validation examples for both learned policies:

Policy Validation top-choice agreement with search teacher
Original Qwen3-0.6B, option-token scoring 25.39%
Qwen3-0.6B, one-epoch LoRA pilot 32.23%
Uniform random legal choice (expected) 31.67%
Always first listed option 38.87%
This structured policy 50.20%

The Qwen pilot remains near chance and its 32 inspected decisions always selected the second option. This tiny baseline was added after that diagnosis, before its own gameplay test. It is a separate supervised model trained from scratch, not a Qwen adapter, a Jev model, or a multimodal model.

Independent gameplay pilot

Three held-out seeds (100, 101, 102), one fixed maze, speed 1, realtime clock, 180 simulated seconds maximum per game. Actual request latency advances the game. Training was stopped during inference. All tests ran on the same DGX host; the tiny policy runs on CPU, while Qwen runs on the GPU.

Policy Mean pellets / game Mean per-game request p50
base 0.0 21 ms
greedy 173.3 0 ms
random 33.0 0 ms
structured 304.3 1 ms
trained 49.0 21 ms

Zero ms denotes integer-ms reporting resolution, not zero computation.

Every per-seed score is in gameplay_results.json. This is a three-game exploratory comparison, not a claim about unseen maps or games. The tiny policy was selected using validation cross-entropy, never gameplay test scores. Jev was not measured.

Architecture and training

Each legal option is encoded by 15 numeric facts: current lives, remaining pellets, distance to junction, frightened time, ghost distances and presence, approaching ghost flag, edible-ghost distance, nearby food, power-pellet distance, and whether the option means turning back. Feature scales are fixed constants in the source.

A shared 15 → 64 → 64 option network is followed by mean/max set pooling and a 192 → 64 → 1 scoring network. Softmax is restricted to legal choices. Shared weights and set pooling make it permutation-equivariant: reordering options reorders the scores. The measured permutation check differed by at most 2.1e-07.

Training: NVIDIA GB10, FP32, seed 42, AdamW 0.001, weight decay 0.0001, batch 128, 100 epochs. Check validation every 10 epochs; epoch 90 has the lowest cross-entropy (1.0931) and is released. Training plus validation took 8.6 seconds, excluding data generation and model setup.

The label is a soft distribution over a search oracle's values (temperature 50). No future state or oracle score is supplied at inference. See the data card for the collection recipe, seed ranges, duplicate removal and exact file hashes.

Run it on your computer

Download structured_policy.py, predict_tiny.py, and example.json from this repo, then run:

pip install torch safetensors huggingface_hub
python predict_tiny.py --request example.json

The script downloads this repository's tiny weight file and prints a legal direction and candidate distribution. It executes on CPU. The model uses custom PyTorch code; it is not an AutoModelForCausalLM checkpoint.

To reproduce training on a CUDA GPU:

python download_data.py
python train_structured.py

Scope and attribution

The policy uses structured text-derived game facts, not pixels. It has not been trained or evaluated for general language reasoning, real-world inspection or multimodal perception. Candidate scores are relative preferences, not calibrated probabilities of survival. Three seeds are a small pilot; wider evaluation remains necessary before drawing a general performance conclusion.

Game engine, feature encoder and search oracle: grapeot/decision-pacman, commit 593ed1f59f4f97ad0bda7287ff304e9667360089, MIT (LICENSE.upstream). This project supplies the fresh data collection, shared-option policy architecture, new GPU training run and artifacts. Weights are Apache-2.0. Personal research project.

Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
17.6k params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train guilindev/pacman-decision-tiny

Space using guilindev/pacman-decision-tiny 1