Commit History
Fix API_BASE_URL to correct HF inference /v1 endpoint 892a583
Fix HF Inference API URL to include model path /v1 6f64e17
Fix API_BASE_URL - remove duplicate /v1 4588d74
Switch to HF Inference API + zephyr-7b (no provider needed) 0445d40
Switch to Phi-3.5-mini-instruct (free HF router) 9cbe0b9
Switch to Llama-3.2-3B-Instruct (supported by HF router) ccce53a
Add timeout+stderr logging to LLM calls, reduce max_tokens to 64 3e1d372
Log to stderr + gunicorn capture-output for HF diagnosis f9c6cbc
Cache bust + debug env print for HF token diagnosis b23f827
Fix Dockerfile MODEL_NAME to Qwen2.5-0.5B 5e93eed
Fix checkpoint path + default model to Qwen2.5-0.5B for HF API fc9640e
Default to multi-agent mode on load 10f7e3d
Track PNGs with LFS db9ddf0
Run 3: Qwen2.5-0.5B GRPO 500 steps | reward 0.96→1.75 | conflict 0.375→0.1875 | dashboard chart | agents wired 0a0a4fc
add /game route for game mode 0729b9c
add game mode files from remote 46cf054
final submission: real reward curve from training logs 05806fd
fix: clean up auto_multi, set mx_obs=None on done 1acf997
remove filter: let LLM produce natural conflicts for LLM vs LoRA comparison be71a02
fix: use correct steps/seed per task on mode switch 12d5534
debug max_steps + safety counter in runAllMulti 29f3ccc
fix: await reset before runAllMulti loop 5416bf7
fix: auto-reset multi env when switching to multi-agent mode 3fe7e1a
debug: log filtered proposals before step_multiagent d7e15ea
fix: enforce sensor ownership server-side, drop cross-agent proposals 4101a93
fix: restrict LLM proposals to agent-type sensors, prevent S1 conflicts 2dfcf50
auto-load easy task on page load a185249
add lora checkpoints and training updates 099deaf
exclude checkpoints from hf space 3345275
track png with lfs 9144a2b
fix: rewards, conflict detection, training comparison panel f114961
Merge branch 'main' of https://github.com/smritis21/arya b95ebe2
Merge pull request #12 from smritis21/chathu01 e493bfd unverified
Merge branch 'main' into chathu01 d5c45b7 unverified
feat: improve UI visualization and training evaluation with smooth transitions and before/after metrics c334ad7
Remove AUDIT_UPDATED.md and reward_curve.png 43c73bc
Fix: task difficulty counts, dashboard legend modal, training metrics, reward curve, /metrics/history endpoint cfa1b2b
Merge pull request #11 from smritis21/chathu01 8a9a8ed unverified
fix: resolve training stagnation by randomizing curriculum seeds and enabling real-time metrics logging 989636e
fix: evaluate() uses trained agents, grader uses agent instances in eval loop e8071ae
feat: add format_state() + llm_propose() — structured ISR prompts with temperature=0.3 generation and JSON validation 95ab542
feat: cap GRPO at 30 steps via GRPOConfig(max_steps=30) — correct TrainingArguments approach b11269a
fix: remove max_steps from train() call — TRL GRPOTrainer does not support it as a runtime arg ff227e8
feat: fast demo mode — 20 episodes, 10 steps, GRPO capped at 30 steps, model saved to checkpoints/ c2a6d75
fix: strip unsupported GRPOConfig args — Unsloth patched config only accepts minimal args e1a5fd6
refactor: modularize GRPOConfig definition and implement robust keyword argument handling for GRPOTrainer initialization 85ace5d
refactor: improve model loading robustness and simplify simulation environment logic in train_colab.py 70c2524
train collab changes 0412fdb
Vicky commited on