LunarLander-v3 in 32 parameters (max-safety)
A 32-parameter single-layer neural network that solves Gymnasium
LunarLander-v3:
231.2 mean return, 99.7% landings, 60/60 seeds, crashing in only ~0.25% of episodes.
logits = Linear(8 -> 4, bias=False)(obs) # 32 parameters
action = argmax(logits)
It is not a distilled big model and there is no teacher: the architecture was found by a
zero-backprop search over "C2 centroids" (KMeans over s โ onehot(a) followed by
evolution). The same object has two readings โ a single linear layer, or a
single-centroid nearest-neighbour policy; they are provably identical:
argmin_a ||c - C2(o,a)||^2 == argmax_a (o @ W), W[:, a] = c[a*8 : (a+1)*8]
The weights shipped here are the max-safety variant (see Choice of objective below).
Evaluation
60 evaluation seeds x 100 episodes = 6000 episodes, deterministic greedy policy:
| metric | max-safety (this model) | crash-averse | mean-optimal teacher |
|---|---|---|---|
| mean return | 231.23 | 239.46 | 247.20 |
| std over seeds | 3.39 | 2.87 | 5.89 |
| min / max seed mean | 223.30 / 238.74 | 232.21 / 244.62 | 232.78 / 260.70 |
| seeds with mean >= 200 | 60 / 60 | 60 / 60 | 60 / 60 |
landing rate (env +100) |
99.67% | 99.42% | 94.10% |
crash rate (env -100) |
0.25% | 0.57% | 5.88% |
| parameters | 32 | 32 | 32 |
Reference random policy on this env is about -100..-500; the "solved" threshold is 200. Reproduce with:
pip install -r requirements.txt
python verify.py --episodes 100 --seeds 60
What "landing rate" means. An episode ends with the environment's own ยฑ100 signal:
+100 when the lander comes to rest (both legs grounded, low speed, small angle) โ a safe
landing; -100 when the lander's body hits the ground or it flies off screen โ a crash
(see LunarLander._get_reward). "Landing rate" is the fraction of episodes that received the
+100 verdict. Measured exactly on the 6000 episodes:
| verdict | episodes | share |
|---|---|---|
landed (env +100) |
5980 | 99.67% |
crashed (env -100) |
15 | 0.25% |
| timeout (1000 steps) | 5 | 0.08% |
The return >= 100 proxy (99.67%) agrees with the environment's own verdict on 6000 / 6000
episodes. python landing_check.py reproduces this cross-check.
Because a safe landing is worth roughly +100..+320 (the terminal +100 plus shaping and leg contacts), the "solved" criterion (mean >= 200) is equivalent to "lands almost every time".
Choice of objective (why 231 and not 247)
In this environment average score and reliability trade off, and the shipped weights pick
maximum safety. Starting from the mean-optimal teacher (247.2 mean, 5.9% crashes) we refined the
weights evolutionarily with a crash penalty (mean - 500 * crash_rate) on a separate train
seed set, evaluating on held-out seeds. On the 60 canonical evaluation seeds:
| mean | landings | crashes | |
|---|---|---|---|
| mean-optimal teacher | 247.2 | 94.1% | 5.9% |
| crash-averse | 239.5 | 99.4% | 0.6% |
| max-safety (this model) | 231.2 | 99.7% | 0.25% |
The crashes are not caused by a few "bad seeds" โ all 60 seeds pass, and failures are spread over episodes (2-11 crashes out of 100 per seed); they start from faster, more tilted spawns. They are also policy-addressable: the same 32-parameter class contains a policy with ~20x fewer crashes, so 5.9% is a compromise of one point, not a hard limit of the class. Reliability is bought with mean return โ within the same 32 parameters.
Earlier revisions (crash-averse 239.5 / 0.6%; mean-optimal 247.2 / 5.9%) are preserved in this repository's git history.
Reproducibility (why not to over-trust the crash rate)
The max-safety weights above are start 1 of an experiment with 6 evolutionary runs of the
same objective (mean - 500 * crash_rate). The six runs land at very different points, and
several barely move from the teacher at all:
| start | TEST mean | TEST crash rate |
|---|---|---|
| 0 | 247.7 | 5.7% |
| 1 | 231.6 | 0.22% |
| 2 | 250.1 | 2.0% |
| 3 | 246.8 | 6.7% |
| 4 | 258.1 | 3.5% |
| 5 | 237.7 | 4.0% |
| (teacher) | 248.3 | 5.5% |
All six still solve every seed (>= 230 mean, 60/60), so the regime is safe โ but the crash rate is noisy: only 2 of 6 starts reached <= 2%. Reproducing a specific crash rate is therefore not guaranteed at this small evolution budget; run more starts (or a larger budget) if you need a target reliability. Two starts Pareto-dominate the teacher on both axes (start 2: 250.1 / 2.0%; start 4: 258.1 / 3.5%), i.e. higher score and fewer crashes than the mean-optimal policy.
Out-of-sample check (fresh seeds)
Because the shipped weights were selected as the best of those 6 runs, the numbers above were re-checked on entirely fresh seeds โ three disjoint blocks of 60 seeds x 100 episodes (18,000 episodes per policy) that were used neither for training nor for selection:
| policy | mean | landings | crashes | seeds >= 200 |
|---|---|---|---|---|
| max-safety (this model) | 231.4 | 99.83% | 0.17% | 60/60 in every block |
| mean-optimal teacher | 246.0 | 93.9% | 6.0% | 60/60 |
| start 2 (kept in history) | 248.6 | 97.7% | 2.3% | 60/60 |
| start 4 (kept in history) | 258.0 | 96.3% | 3.5% | 60/60 |
The shipped model reproduces its canonical numbers on fresh seeds (231.4 vs 231.2 mean; 0.17% vs 0.25% crashes), so its low crash rate is a property of the weights, not of the seed set used to pick them. The two variants kept in history also hold up out-of-sample and Pareto-dominate the mean-optimal teacher on both axes (higher score and fewer crashes).
Usage
Standalone (only torch + safetensors):
from model import TinyLunarPolicy
policy = TinyLunarPolicy.from_pretrained("Dimitrius174/lunarlander-lin32")
# or: TinyLunarPolicy.from_safetensors("model.safetensors")
action = policy.act([0.0, 1.4, 0.0, -0.6, 0.0, 0.0, 0.0, 0.0]) # int in 0..3
With transformers (custom code via auto_map):
import torch
from transformers import AutoModel
model = AutoModel.from_pretrained("Dimitrius174/lunarlander-lin32", trust_remote_code=True).eval()
logits = model(torch.tensor([0.0, 1.4, 0.0, -0.6, 0.0, 0.0, 0.0, 0.0])) # shape (4,)
action = int(logits.argmax())
Gymnasium loop:
import gymnasium as gym
from model import TinyLunarPolicy
env = gym.make("LunarLander-v3")
policy = TinyLunarPolicy.from_pretrained("Dimitrius174/lunarlander-lin32")
obs, _ = env.reset(seed=0)
done = False
while not done:
obs, _, term, trunc, _ = env.step(policy.act(obs))
done = term or trunc
Observation and action space
obs (8,): x, y, vx, vy, angle, angular_velocity, leg_left_contact, leg_right_contact.
action (4): do nothing, fire left engine, fire main engine, fire right engine.
Is 32 really minimal?
In a capacity study of small networks distilled from this policy (same data, optimiser and steps; verified on the same 60 seeds), the findings were:
- the linear 32-parameter model was the only one that reproduced the policy under every regime tried (activation, initialisation, training budget);
- a 1-hidden-layer
tanhMLP needed ~43-56 parameters (43 borderline, 56 reliable); ReLU-family MLPs at the same budget were unreliable (some seeds solve, most do not), so extra non-linearity costs parameters and training fragility;- beyond ~56 parameters nothing improves (the environment, not the network, sets the ceiling).
Buying a linear layer instead of an MLP is therefore the cheap and robust choice.
Files
| file | what |
|---|---|
model.py |
standalone TinyLunarPolicy (nn.Module), no extra deps |
modeling_tiny.py, configuration_tiny.py |
transformers wrappers (auto_map, trust_remote_code) |
config.json |
architecture + env metadata |
model.safetensors |
the 32 weights (linear.weight, shape 4x8) โ max-safety variant |
centroid_safe.npy |
the shipped weights in centroid form (source of model.safetensors) |
verify.py |
structural checks + live Gymnasium evaluation |
landing_check.py |
cross-checks the "landing rate" against the environment's own ยฑ100 verdict |
export_weights.py |
rebuild model.safetensors from the source centroid vector |
License
MIT. See LICENSE.
- Downloads last month
- 42