MazeGPT: Reactive 2D Maze Navigation via Pure RL (Zero Pretraining)

Toolkit for Training, Evaluating, and Quantizing Reactive Maze Navigation Agents

Open In Colab License: MIT PyTorch Hugging Face


Task Formalization & Environment Specifications

MazeGPT investigates how compact autoregressive Transformers can develop reactive spatial navigation and obstacle avoidance in partially observable 2D maze environments (POMDP) under pure reinforcement learning (Zero Pretraining).

Environment Dynamics & Observation Model:

  • Observation Space (Observation): At each step, the environment injects a 4-cell local field of view U, D, L, R (where . denotes an open path, and # denotes a wall; out-of-bounds cells are treated as walls).
  • Action Space: Discrete actions U, D, L, R with step delimiters.
  • Transition Dynamics: If the agent attempts to collide into a wall or out of bounds, the action is marked illegal, and the agent remains in its current cell.
  • Sparse Reward: A sparse reward (+1.0) is granted solely upon reaching the terminal goal G; intermediate rewards are 0. No shortest-path BFS/A supervision is provided during training*.

Quickstart & Navigation


Step 1: Quickstart on Google Colab

Train and evaluate pure RL agents using free Colab GPUs/CPUs:

  1. Download Colab_Run_Maze_Transformer.ipynb.
  2. Open Google Colab -> Click Upload -> Select the .ipynb file.
  3. Click Runtime -> Run All.

Step 2: Local Installation & Environment Setup

# 1. Clone repository
git clone https://huggingface.co/Hana-ame/maze-transformer
cd maze-transformer

# 2. Install dependencies (Python >= 3.8, PyTorch >= 2.0)
pip install torch openpyxl huggingface_hub matplotlib pandas
pip install -e .

Step 3: Interactive Maze Simulator

# Test dynamic agent navigation on a randomized 5x7 maze
./use_model.sh --checkpoint checkpoints/l2_d64_maze_forced_obs_rl_final.pt

Step 4: Training via maze_config.json

1. Configure maze_config.json

{
  "layers": 2,
  "d": 64,
  "heads": 4,
  "steps": 120,
  "batch_size": 6,
  "lr": 3e-4,
  "min_size": 5,
  "max_size": 9,
  "single": true,
  "datasource": {
    "type": "random_perfect_maze",
    "observation": "forced_obs_4cell",
    "reward": "sparse_goal_reach"
  }
}

2. Launch Trainer

# Pure RL (GRPO) training
python -m maze_transformer.train --config maze_config.json

# 50-step smoke verification
python -m maze_transformer.train --quick

Step 5: Multi-Scale Maze Navigation Benchmark (5x5 - 9x9)

Evaluate goal reaching rates and collision counts across random maze topologies:

python -m maze_transformer.evaluate     --checkpoint checkpoints/l2_d64_maze_forced_obs_rl_final.pt     --sizes 5x5,5x7,6x6,6x9     --n_trials 24

Step 6: Dynamic INT8 Quantization Benchmark

Quantize linear weights via PyTorch dynamic INT8 to assess policy stability:

python -m maze_transformer.quantize --checkpoint checkpoints/l2_d64_maze_forced_obs_rl_final.pt

Pretrained Model Checkpoints

Checkpoint Identifier Architecture Protocol Notes & Performance
l2_d64_maze_forced_obs_rl_final.pt 2L 路 64d 路 4 heads Pure RL (GRPO 120 steps) Primary model (5x7 maze reach rate 83.3%)
l2_d64_maze_rnn_rl_ref.pt 1L 路 128d 路 GRU-RNN Pure RL (120 steps) Recurrent baseline control model
l2_d64_maze_sft_ref.pt 2L 路 64d 路 Transformer SFT (BFS Shortest Path) Supervised navigation upper bound baseline

Maze Experiment Excel Workbooks

All 13 maze navigation experiments (GRPO vs. REINFORCE vs. RNN across 60/120/300 steps, context compression, and INT8 quantization) are recorded in:
馃憠 EXPERIMENTS_ALL.xlsx


License

This repository is distributed under the MIT License.

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading