rush-soldier

A soldier network for rush, a 2D team shooter for 500 v 500 battles with control points (code, world, viewer and training: https://github.com/MedDevSystems/rush-arena).

One network plays every soldier of an army, all 8 classes (the class is part of the observation). The army is commanded by the heuristic army commander of the rush package (respawn points, which control point each squad of 8 takes); the network observes its order and fights.

Use

pip install -e git+https://github.com/MedDevSystems/rush-arena#egg=rush-arena
rush-battle --blue hf:koskokos/rush-soldier --red random --map Front50 --view
rush-battle --blue hf:koskokos/rush-soldier --red hf:koskokos/rush-soldier --map Warfront500 --device cuda
from rush.battle import play
play("hf:koskokos/rush-soldier", "random", map_name="Front50", out="battle.arena.bin.gz")

Model

Per soldier and decision (every 4 frames at 60 fps): a transformer over entity tokens (nearest allies and enemies, bullets, control points, last-seen enemies, wall grids), an LSTM, and a state-conditioned attention readout over further tokens (up to 56 soldiers, the 32 most dangerous bullets, a 17 × 17 coarse map of the battle); four action heads — move (9), fire (2), turn (7, fine and coarse), dash (2) — masked by the rules. 1.7 M parameters. config.json carries the full observation layout; the weights are in model.safetensors.

Training

Self-play by generations: the network plays a pool of its past generations and learns from the battle outcome with PPO (per-soldier rewards for damage, kills, captures, score and the win, weighted per class), with a KL anchor to its starting point that decays over training. A candidate becomes the next generation only when a sequential one-sided paired t-test (5 %) over side-swapped battle pairs on the same seed shows it is stronger. This checkpoint is the latest accepted generation. The same procedure ships with the code (rush-selfplay, rush-gate).

Limitations

  • A soldier, not a commander: the strategy (respawns, squad orders) comes from the heuristic army commander.
  • Trained on the rules of rush as published with this model (world version 15, observation layout 4); other rules or layouts need retraining.
  • Greedy actions are used for play; sampling is used in training.

License

MIT.

Downloads last month
15
Safetensors
Model size
1.68M params
Tensor type
F32
·
Video Preview
loading