Commit History

Keep mx_obs after episode done, prevent reset_multi first error
53a8778

smiit commited on

Fix API_BASE_URL to correct HF inference /v1 endpoint
892a583

smiit commited on

Fix HF Inference API URL to include model path /v1
6f64e17

smiit commited on

Fix API_BASE_URL - remove duplicate /v1
4588d74

smiit commited on

Switch to HF Inference API + zephyr-7b (no provider needed)
0445d40

smiit commited on

Switch to Phi-3.5-mini-instruct (free HF router)
9cbe0b9

smiit commited on

Switch to Llama-3.2-3B-Instruct (supported by HF router)
ccce53a

smiit commited on

Add timeout+stderr logging to LLM calls, reduce max_tokens to 64
3e1d372

smiit commited on

Log to stderr + gunicorn capture-output for HF diagnosis
f9c6cbc

smiit commited on

Cache bust + debug env print for HF token diagnosis
b23f827

smiit commited on

Fix Dockerfile MODEL_NAME to Qwen2.5-0.5B
5e93eed

smiit commited on

Fix checkpoint path + default model to Qwen2.5-0.5B for HF API
fc9640e

smiit commited on

Default to multi-agent mode on load
10f7e3d

smiit commited on

Track PNGs with LFS
db9ddf0

smiit commited on

Run 3: Qwen2.5-0.5B GRPO 500 steps | reward 0.96→1.75 | conflict 0.375→0.1875 | dashboard chart | agents wired
0a0a4fc

smiit commited on

add /game route for game mode
0729b9c

smiit commited on

add game mode files from remote
46cf054

smiit commited on

final submission: real reward curve from training logs
05806fd

smiit commited on

fix: clean up auto_multi, set mx_obs=None on done
1acf997

smiit commited on

remove filter: let LLM produce natural conflicts for LLM vs LoRA comparison
be71a02

smiit commited on

fix: use correct steps/seed per task on mode switch
12d5534

smiit commited on

debug max_steps + safety counter in runAllMulti
29f3ccc

smiit commited on

fix: await reset before runAllMulti loop
5416bf7

smiit commited on

fix: auto-reset multi env when switching to multi-agent mode
3fe7e1a

smiit commited on

debug: log filtered proposals before step_multiagent
d7e15ea

smiit commited on

fix: enforce sensor ownership server-side, drop cross-agent proposals
4101a93

smiit commited on

fix: restrict LLM proposals to agent-type sensors, prevent S1 conflicts
2dfcf50

smiit commited on

auto-load easy task on page load
a185249

smiit commited on

add lora checkpoints and training updates
099deaf

smiit commited on

exclude checkpoints from hf space
3345275

smiit commited on

track png with lfs
9144a2b

smiit commited on

fix: rewards, conflict detection, training comparison panel
f114961

smiit commited on

Merge branch 'main' of https://github.com/smritis21/arya
b95ebe2

smiit commited on

Merge pull request #12 from smritis21/chathu01
e493bfd
unverified

chathurvitha commited on

Merge branch 'main' into chathu01
d5c45b7
unverified

chathurvitha commited on

feat: improve UI visualization and training evaluation with smooth transitions and before/after metrics
c334ad7

chathurvitha commited on

Remove AUDIT_UPDATED.md and reward_curve.png
43c73bc

smiit commited on

Fix: task difficulty counts, dashboard legend modal, training metrics, reward curve, /metrics/history endpoint
cfa1b2b

smiit commited on

Merge pull request #11 from smritis21/chathu01
8a9a8ed
unverified

chathurvitha commited on

fix: resolve training stagnation by randomizing curriculum seeds and enabling real-time metrics logging
989636e

chathurvitha commited on

fix: evaluate() uses trained agents, grader uses agent instances in eval loop
e8071ae

smiit commited on

feat: add format_state() + llm_propose() — structured ISR prompts with temperature=0.3 generation and JSON validation
95ab542

Vizxal commited on

feat: cap GRPO at 30 steps via GRPOConfig(max_steps=30) — correct TrainingArguments approach
b11269a

Vizxal commited on

fix: remove max_steps from train() call — TRL GRPOTrainer does not support it as a runtime arg
ff227e8

Vizxal commited on

feat: fast demo mode — 20 episodes, 10 steps, GRPO capped at 30 steps, model saved to checkpoints/
c2a6d75

Vizxal commited on

fix: strip unsupported GRPOConfig args — Unsloth patched config only accepts minimal args
e1a5fd6

Vizxal commited on

refactor: modularize GRPOConfig definition and implement robust keyword argument handling for GRPOTrainer initialization
85ace5d

Vizxal commited on

refactor: improve model loading robustness and simplify simulation environment logic in train_colab.py
70c2524

Vizxal commited on

train collab changes
0412fdb

Vicky commited on

Person 3: unify AryaXEnv mode param + fix evaluate() to use trained agents
3771eaf

smiit commited on