LunarLander-v3 in 32 parameters (max-safety)

A 32-parameter single-layer neural network that solves Gymnasium LunarLander-v3: 231.2 mean return, 99.7% landings, 60/60 seeds, crashing in only ~0.25% of episodes.

logits = Linear(8 -> 4, bias=False)(obs)     # 32 parameters
action = argmax(logits)

It is not a distilled big model and there is no teacher: the architecture was found by a zero-backprop search over "C2 centroids" (KMeans over s โŠ— onehot(a) followed by evolution). The same object has two readings โ€” a single linear layer, or a single-centroid nearest-neighbour policy; they are provably identical:

argmin_a ||c - C2(o,a)||^2  ==  argmax_a (o @ W),   W[:, a] = c[a*8 : (a+1)*8]

The weights shipped here are the max-safety variant (see Choice of objective below).

Evaluation

60 evaluation seeds x 100 episodes = 6000 episodes, deterministic greedy policy:

metric max-safety (this model) crash-averse mean-optimal teacher
mean return 231.23 239.46 247.20
std over seeds 3.39 2.87 5.89
min / max seed mean 223.30 / 238.74 232.21 / 244.62 232.78 / 260.70
seeds with mean >= 200 60 / 60 60 / 60 60 / 60
landing rate (env +100) 99.67% 99.42% 94.10%
crash rate (env -100) 0.25% 0.57% 5.88%
parameters 32 32 32

Reference random policy on this env is about -100..-500; the "solved" threshold is 200. Reproduce with:

pip install -r requirements.txt
python verify.py --episodes 100 --seeds 60

What "landing rate" means. An episode ends with the environment's own ยฑ100 signal: +100 when the lander comes to rest (both legs grounded, low speed, small angle) โ€” a safe landing; -100 when the lander's body hits the ground or it flies off screen โ€” a crash (see LunarLander._get_reward). "Landing rate" is the fraction of episodes that received the +100 verdict. Measured exactly on the 6000 episodes:

verdict episodes share
landed (env +100) 5980 99.67%
crashed (env -100) 15 0.25%
timeout (1000 steps) 5 0.08%

The return >= 100 proxy (99.67%) agrees with the environment's own verdict on 6000 / 6000 episodes. python landing_check.py reproduces this cross-check.

Because a safe landing is worth roughly +100..+320 (the terminal +100 plus shaping and leg contacts), the "solved" criterion (mean >= 200) is equivalent to "lands almost every time".

Choice of objective (why 231 and not 247)

In this environment average score and reliability trade off, and the shipped weights pick maximum safety. Starting from the mean-optimal teacher (247.2 mean, 5.9% crashes) we refined the weights evolutionarily with a crash penalty (mean - 500 * crash_rate) on a separate train seed set, evaluating on held-out seeds. On the 60 canonical evaluation seeds:

mean landings crashes
mean-optimal teacher 247.2 94.1% 5.9%
crash-averse 239.5 99.4% 0.6%
max-safety (this model) 231.2 99.7% 0.25%

The crashes are not caused by a few "bad seeds" โ€” all 60 seeds pass, and failures are spread over episodes (2-11 crashes out of 100 per seed); they start from faster, more tilted spawns. They are also policy-addressable: the same 32-parameter class contains a policy with ~20x fewer crashes, so 5.9% is a compromise of one point, not a hard limit of the class. Reliability is bought with mean return โ€” within the same 32 parameters.

Earlier revisions (crash-averse 239.5 / 0.6%; mean-optimal 247.2 / 5.9%) are preserved in this repository's git history.

Reproducibility (why not to over-trust the crash rate)

The max-safety weights above are start 1 of an experiment with 6 evolutionary runs of the same objective (mean - 500 * crash_rate). The six runs land at very different points, and several barely move from the teacher at all:

start TEST mean TEST crash rate
0 247.7 5.7%
1 231.6 0.22%
2 250.1 2.0%
3 246.8 6.7%
4 258.1 3.5%
5 237.7 4.0%
(teacher) 248.3 5.5%

All six still solve every seed (>= 230 mean, 60/60), so the regime is safe โ€” but the crash rate is noisy: only 2 of 6 starts reached <= 2%. Reproducing a specific crash rate is therefore not guaranteed at this small evolution budget; run more starts (or a larger budget) if you need a target reliability. Two starts Pareto-dominate the teacher on both axes (start 2: 250.1 / 2.0%; start 4: 258.1 / 3.5%), i.e. higher score and fewer crashes than the mean-optimal policy.

Out-of-sample check (fresh seeds)

Because the shipped weights were selected as the best of those 6 runs, the numbers above were re-checked on entirely fresh seeds โ€” three disjoint blocks of 60 seeds x 100 episodes (18,000 episodes per policy) that were used neither for training nor for selection:

policy mean landings crashes seeds >= 200
max-safety (this model) 231.4 99.83% 0.17% 60/60 in every block
mean-optimal teacher 246.0 93.9% 6.0% 60/60
start 2 (kept in history) 248.6 97.7% 2.3% 60/60
start 4 (kept in history) 258.0 96.3% 3.5% 60/60

The shipped model reproduces its canonical numbers on fresh seeds (231.4 vs 231.2 mean; 0.17% vs 0.25% crashes), so its low crash rate is a property of the weights, not of the seed set used to pick them. The two variants kept in history also hold up out-of-sample and Pareto-dominate the mean-optimal teacher on both axes (higher score and fewer crashes).

Usage

Standalone (only torch + safetensors):

from model import TinyLunarPolicy

policy = TinyLunarPolicy.from_pretrained("Dimitrius174/lunarlander-lin32")
# or: TinyLunarPolicy.from_safetensors("model.safetensors")

action = policy.act([0.0, 1.4, 0.0, -0.6, 0.0, 0.0, 0.0, 0.0])   # int in 0..3

With transformers (custom code via auto_map):

import torch
from transformers import AutoModel

model = AutoModel.from_pretrained("Dimitrius174/lunarlander-lin32", trust_remote_code=True).eval()
logits = model(torch.tensor([0.0, 1.4, 0.0, -0.6, 0.0, 0.0, 0.0, 0.0]))   # shape (4,)
action = int(logits.argmax())

Gymnasium loop:

import gymnasium as gym
from model import TinyLunarPolicy

env = gym.make("LunarLander-v3")
policy = TinyLunarPolicy.from_pretrained("Dimitrius174/lunarlander-lin32")
obs, _ = env.reset(seed=0)
done = False
while not done:
    obs, _, term, trunc, _ = env.step(policy.act(obs))
    done = term or trunc

Observation and action space

obs (8,): x, y, vx, vy, angle, angular_velocity, leg_left_contact, leg_right_contact. action (4): do nothing, fire left engine, fire main engine, fire right engine.

Is 32 really minimal?

In a capacity study of small networks distilled from this policy (same data, optimiser and steps; verified on the same 60 seeds), the findings were:

  • the linear 32-parameter model was the only one that reproduced the policy under every regime tried (activation, initialisation, training budget);
  • a 1-hidden-layer tanh MLP needed ~43-56 parameters (43 borderline, 56 reliable);
  • ReLU-family MLPs at the same budget were unreliable (some seeds solve, most do not), so extra non-linearity costs parameters and training fragility;
  • beyond ~56 parameters nothing improves (the environment, not the network, sets the ceiling).

Buying a linear layer instead of an MLP is therefore the cheap and robust choice.

Files

file what
model.py standalone TinyLunarPolicy (nn.Module), no extra deps
modeling_tiny.py, configuration_tiny.py transformers wrappers (auto_map, trust_remote_code)
config.json architecture + env metadata
model.safetensors the 32 weights (linear.weight, shape 4x8) โ€” max-safety variant
centroid_safe.npy the shipped weights in centroid form (source of model.safetensors)
verify.py structural checks + live Gymnasium evaluation
landing_check.py cross-checks the "landing rate" against the environment's own ยฑ100 verdict
export_weights.py rebuild model.safetensors from the source centroid vector

License

MIT. See LICENSE.

Downloads last month
42
Safetensors
Model size
32 params
Tensor type
F32
ยท
Video Preview
loading