docs: drop R4-η» retrospective paragraph β it reintroduced the wrong shaping=0.5 value 8eeeb67 Lee93whut Lee93whut commited on Jun 1
docs: polish pass β README wording, experiment_log structure + honest retrospective, LaTeX micro-style e1ecae1 Lee93whut Lee93whut commited on Jun 1
chore: codebase hygiene pass β untrack weights, migrate to logging, tidy comments 17bc537 Lee93whut Lee93whut commited on Jun 1
docs: clean up R3/R4 record and consolidate technical narrative 92423f0 Lee93whut Lee93whut commited on Jun 1
docs: finalize R4 documentation β Dueling 84% Holdout, full ablation record acbd4c5 Lee93whut commited on Jun 1
refactor(model): update architecture docs and set dueling as default algorithm 34ad2cc Lee93whut commited on Jun 1
feat(round4): four-algorithm ablation β Dueling best at 84% Holdout 44cfe4c Lee93whut commited on Jun 1
docs(round4): finalize R4 Double DQN results β 78% Holdout, Grid-SPL clarification b14b412 Lee93whut commited on Jun 1
chore(weights): upgrade vanilla/dueling/double_dueling to 4-channel R4 weights 3379ed4 Lee93whut commited on Jun 1
fix(demo): strengthen anti-loop by penalizing moves toward high-frequency cells a888a00 Lee93whut commited on May 31
chore(weights): update double DQN weight to R4-A3 4-channel (78% holdout) f3ed6b3 Lee93whut commited on May 31
fix(demo): auto-infer input_channels from checkpoint weight shape 4f4fb4a Lee93whut commited on May 31
docs: README β results table, architecture, quickstart, references 385cc9f Lee93whut commited on May 31
feat(demo): Streamlit web demo β Plotly heatmap, anti-loop inference a264030 Lee93whut commited on May 31
docs(round4): complete experiment record β A1/A2/A3 full EVAL data and conclusions a91b194 Lee93whut commited on May 31
style(train): remove forward-reference quotes from type hints (Python 3.10+) 274376b Lee93whut commited on May 31
fix(train): use terminated-only mask for TD bootstrap (Gymnasium v0.26) 670449d Lee93whut commited on May 31
fix(train): guarantee BFS-connected start/goal, bounded retry with fallback 92a3812 Lee93whut commited on May 31
feat(round4): upgrade obs 3->4 channels (visited_map) + EVAL-based checkpoint 062d629 Lee93whut commited on May 31
feat(round3): buffer=80k + target_freq=1500 + shaping=0.5 β 74% holdout, SPL=0.735 c1b9ba8 Lee93whut commited on May 31
feat(round2): extended training, Double DQN 64% holdout, SPL=0.633 ff1b1b8 Lee93whut commited on May 31
feat(round1): baseline DQN variants β Vanilla/Double/Dueling/Double+Dueling bf17b0c Lee93whut commited on May 31