really clean comparison, and publishing the adapters is great. one column i'd add next time: format adherence, not just eval loss. on coding agent traces the failure that bites in practice is malformed tool calls and missed stop tokens, and with quantized and abliterated models (i build grunz, which runs them) we see those degrade while the model otherwise looks fine. a simple "percent of generated tool calls that actually parse" on a held out set would show whether the 4-bit runs are really equivalent where it matters
Ben
GRUNZAI
AI & ML interests
None yet
Recent Activity
repliedto harshitkgupta's post about 3 hours ago
Fine-tuned Qwen 2.5 (0.5B → 3B) on real coding-agent traces, 10 controlled runs, one 16GB Mac. Compared PyTorch MPS vs. Apple MLX for local LoRA SFT — and the honest answer is "it depends on what you're optimizing for":
• PyTorch MPS: 2.2x–5.7x faster raw throughput, but hits a hard memory wall — can't load a 3B model in FP16 on 16GB.
• Apple MLX: 4-bit QLoRA fits 3B+ models with almost flat memory scaling as context grows (+109 MB going from 1k→4k tokens).
• 4-bit quantization doesn't cost you convergence — eval loss tracks closely across backends.
• The bigger surprise: most of MLX's slowdown isn't the 4-bit dequant tax. Two of the 10 runs went unquantized to isolate it — dequant only explains 1.07x–1.4x of the gap. A ~4.1–4.6x framework-level gap remains either way.
All 10 LoRA adapters + Trackio logs are public so the numbers are checkable, not just claimed.
Full writeup: https://huggingface.co/blog/harshitkgupta/fine-tuning-coding-agents-on-mac-pytorch-mps-mlx
new activity 2 days ago
DontPlanToEnd/UGI-Leaderboard:Test all of the new qwen3.5'sOrganizations
None yet