PvZ Survival Endless champion (recurrent PPO)
The best-playing policy from pvz-rl-env, a Gymnasium environment built on a pybind11 bridge into the real Plants vs. Zombies engine.
- Architecture: convolutional encoder + Mamba (SSM) block, factorized actor head over a
masked
Discrete(496)action space. Trained with PPO (sparse reward, KL-anchored) from a behavior-cloning initialization of a hand-written heuristic teacher. - Observation v1:
(5, 9, 36)spatial + 24 globals. - Training deck (seed types): Sunflower, Twin Sunflower, Melon-pult, Winter Melon,
Gloom-shroom, Snow Pea, Pumpkin, Garlic, Squash, Jalapeno —
[1, 41, 39, 44, 42, 5, 30, 36, 17, 20].
Results (deterministic, wave 1 start, 20-seed paired panels)
| config | mean abs wave | median |
|---|---|---|
| endless, 3000 sun, training deck | 49.7 | 50.5 |
| endless, 3000 sun, coffee deck (slot 7 -> Coffee Bean) | 72.1 | 61.5 |
| stage-1 completion suites (3000 sun) | 21.0, 100% completions | 21 |
Coffee-deck note: this policy was never trained on Coffee Bean. Its slot-7 actions get funneled by the env's legality mask into legal Coffee Bean plays on sleeping mushrooms, and it wakes them usefully — a free skill from masking, not from training.
Usage
git clone https://github.com/awtrisk/pvz-rl-env
cd pvz-rl-env
# build the engine and drop in your own main.pak + properties/ (see README)
import torch
from network import PvZActorCritic
agent = PvZActorCritic()
agent.load_state_dict(torch.load("pvz_endless_champion.pt", weights_only=True)["agent_state"])
agent.eval()
To watch it play: python scripts/record.py --checkpoint pvz_endless_champion.pt --start-sun 3000
(recordings end at the stage-1 boundary; the endless chain flag is env-side).
Weights are one file, a single {"agent_state": state_dict} payload — load with
weights_only=True.