- MazeGPT: Reactive 2D Maze Navigation via Pure RL (Zero Pretraining)
- Task Formalization & Environment Specifications
- Quickstart & Navigation
- Step 1: Quickstart on Google Colab
- Step 2: Local Installation & Environment Setup
- Step 3: Interactive Maze Simulator
- Step 4: Training via
maze_config.json - Step 5: Multi-Scale Maze Navigation Benchmark (5x5 - 9x9)
- Step 6: Dynamic INT8 Quantization Benchmark
- Pretrained Model Checkpoints
- Maze Experiment Excel Workbooks
- License
MazeGPT: Reactive 2D Maze Navigation via Pure RL (Zero Pretraining)
Toolkit for Training, Evaluating, and Quantizing Reactive Maze Navigation Agents
Task Formalization & Environment Specifications
MazeGPT investigates how compact autoregressive Transformers can develop reactive spatial navigation and obstacle avoidance in partially observable 2D maze environments (POMDP) under pure reinforcement learning (Zero Pretraining).
Environment Dynamics & Observation Model:
- Observation Space (Observation): At each step, the environment injects a 4-cell local field of view
U, D, L, R(where.denotes an open path, and#denotes a wall; out-of-bounds cells are treated as walls). - Action Space: Discrete actions
U, D, L, Rwith step delimiters. - Transition Dynamics: If the agent attempts to collide into a wall or out of bounds, the action is marked illegal, and the agent remains in its current cell.
- Sparse Reward: A sparse reward (+1.0) is granted solely upon reaching the terminal goal
G; intermediate rewards are 0. No shortest-path BFS/A supervision is provided during training*.
Quickstart & Navigation
- Step 1: Quickstart on Google Colab
- Step 2: Local Installation & Environment Setup
- Step 3: Interactive Maze Simulator
- Step 4: Training via maze_config.json
- Step 5: Multi-Scale Maze Navigation Benchmark (5x5 - 9x9)
- Step 6: Dynamic INT8 Quantization Benchmark
- Pretrained Model Checkpoints
- Maze Experiment Excel Workbooks
Step 1: Quickstart on Google Colab
Train and evaluate pure RL agents using free Colab GPUs/CPUs:
- Download
Colab_Run_Maze_Transformer.ipynb. - Open Google Colab -> Click Upload -> Select the
.ipynbfile. - Click Runtime -> Run All.
Step 2: Local Installation & Environment Setup
# 1. Clone repository
git clone https://huggingface.co/Hana-ame/maze-transformer
cd maze-transformer
# 2. Install dependencies (Python >= 3.8, PyTorch >= 2.0)
pip install torch openpyxl huggingface_hub matplotlib pandas
pip install -e .
Step 3: Interactive Maze Simulator
# Test dynamic agent navigation on a randomized 5x7 maze
./use_model.sh --checkpoint checkpoints/l2_d64_maze_forced_obs_rl_final.pt
Step 4: Training via maze_config.json
1. Configure maze_config.json
{
"layers": 2,
"d": 64,
"heads": 4,
"steps": 120,
"batch_size": 6,
"lr": 3e-4,
"min_size": 5,
"max_size": 9,
"single": true,
"datasource": {
"type": "random_perfect_maze",
"observation": "forced_obs_4cell",
"reward": "sparse_goal_reach"
}
}
2. Launch Trainer
# Pure RL (GRPO) training
python -m maze_transformer.train --config maze_config.json
# 50-step smoke verification
python -m maze_transformer.train --quick
Step 5: Multi-Scale Maze Navigation Benchmark (5x5 - 9x9)
Evaluate goal reaching rates and collision counts across random maze topologies:
python -m maze_transformer.evaluate --checkpoint checkpoints/l2_d64_maze_forced_obs_rl_final.pt --sizes 5x5,5x7,6x6,6x9 --n_trials 24
Step 6: Dynamic INT8 Quantization Benchmark
Quantize linear weights via PyTorch dynamic INT8 to assess policy stability:
python -m maze_transformer.quantize --checkpoint checkpoints/l2_d64_maze_forced_obs_rl_final.pt
Pretrained Model Checkpoints
| Checkpoint Identifier | Architecture | Protocol | Notes & Performance |
|---|---|---|---|
l2_d64_maze_forced_obs_rl_final.pt |
2L 路 64d 路 4 heads | Pure RL (GRPO 120 steps) | Primary model (5x7 maze reach rate 83.3%) |
l2_d64_maze_rnn_rl_ref.pt |
1L 路 128d 路 GRU-RNN | Pure RL (120 steps) | Recurrent baseline control model |
l2_d64_maze_sft_ref.pt |
2L 路 64d 路 Transformer | SFT (BFS Shortest Path) | Supervised navigation upper bound baseline |
Maze Experiment Excel Workbooks
All 13 maze navigation experiments (GRPO vs. REINFORCE vs. RNN across 60/120/300 steps, context compression, and INT8 quantization) are recorded in:
馃憠 EXPERIMENTS_ALL.xlsx
License
This repository is distributed under the MIT License.