PvZ Survival Endless champion (recurrent PPO)

The best-playing policy from pvz-rl-env, a Gymnasium environment built on a pybind11 bridge into the real Plants vs. Zombies engine.

  • Architecture: convolutional encoder + Mamba (SSM) block, factorized actor head over a masked Discrete(496) action space. Trained with PPO (sparse reward, KL-anchored) from a behavior-cloning initialization of a hand-written heuristic teacher.
  • Observation v1: (5, 9, 36) spatial + 24 globals.
  • Training deck (seed types): Sunflower, Twin Sunflower, Melon-pult, Winter Melon, Gloom-shroom, Snow Pea, Pumpkin, Garlic, Squash, Jalapeno — [1, 41, 39, 44, 42, 5, 30, 36, 17, 20].

Results (deterministic, wave 1 start, 20-seed paired panels)

config mean abs wave median
endless, 3000 sun, training deck 49.7 50.5
endless, 3000 sun, coffee deck (slot 7 -> Coffee Bean) 72.1 61.5
stage-1 completion suites (3000 sun) 21.0, 100% completions 21

Coffee-deck note: this policy was never trained on Coffee Bean. Its slot-7 actions get funneled by the env's legality mask into legal Coffee Bean plays on sleeping mushrooms, and it wakes them usefully — a free skill from masking, not from training.

Usage

git clone https://github.com/awtrisk/pvz-rl-env
cd pvz-rl-env
# build the engine and drop in your own main.pak + properties/ (see README)
import torch
from network import PvZActorCritic

agent = PvZActorCritic()
agent.load_state_dict(torch.load("pvz_endless_champion.pt", weights_only=True)["agent_state"])
agent.eval()

To watch it play: python scripts/record.py --checkpoint pvz_endless_champion.pt --start-sun 3000 (recordings end at the stage-1 boundary; the endless chain flag is env-side).

Weights are one file, a single {"agent_state": state_dict} payload — load with weights_only=True.

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading