Log generation time and token count so the GPU duration can be measured, not guessed f2d515b Running verified adelelsayed1991 commited on 8 days ago
Raise ZeroGPU duration 180->300s: 180 was killing longer completions mid-generation 7bbc0fc verified adelelsayed1991 commited on 8 days ago
Match sft_train.ipynb eval config: max_new_tokens 768->2048, GPU duration 120->180 7b67c79 verified adelelsayed1991 commited on 8 days ago
Increase GPU task duration budget to 120s (45s was too tight for 4-bit tensor packing + generation, causing GPU task aborted) 39dadef verified adelelsayed1991 commited on 9 days ago
Fix: use sft/seed_42/best (not rl/), correct plan-then-SQL prompt, 4-bit NF4 quantization, and SQL extraction from fenced block, per MODEL_CARD.md 19f0c5b verified adelelsayed1991 commited on 9 days ago
Fix: use PEFT torch_device=cpu (not device_map, which was silently ignored) to force CPU-side adapter load ecd77d9 verified adelelsayed1991 commited on 9 days ago
Fix: force adapter safetensors load onto CPU (device_map=cpu) to avoid ZeroGPU emulation gap a7cc8ca verified adelelsayed1991 commited on 9 days ago
Fix: point PeftModel.from_pretrained at the correct adapter subfolder (rl/seed_42/best) 64c1aa0 verified adelelsayed1991 commited on 9 days ago
Add app.py, requirements.txt, schema.sql, README.md to serve the fhirsql model via ZeroGPU 3434a32 verified adelelsayed1991 commited on 9 days ago