Anthony Mikinka
anthonymikinka
AI & ML interests
ANE Focused
Recent Activity
liked a model about 20 hours ago
abenzerps/Qwen-Image-2.1-Uncensored-GGUF reacted to harshitkgupta's post with ๐ about 20 hours ago
Fine-tuned Qwen 2.5 (0.5B โ 3B) on real coding-agent traces, 10 controlled runs, one 16GB Mac. Compared PyTorch MPS vs. Apple MLX for local LoRA SFT โ and the honest answer is "it depends on what you're optimizing for":
โข PyTorch MPS: 2.2xโ5.7x faster raw throughput, but hits a hard memory wall โ can't load a 3B model in FP16 on 16GB.
โข Apple MLX: 4-bit QLoRA fits 3B+ models with almost flat memory scaling as context grows (+109 MB going from 1kโ4k tokens).
โข 4-bit quantization doesn't cost you convergence โ eval loss tracks closely across backends.
โข The bigger surprise: most of MLX's slowdown isn't the 4-bit dequant tax. Two of the 10 runs went unquantized to isolate it โ dequant only explains 1.07xโ1.4x of the gap. A ~4.1โ4.6x framework-level gap remains either way.
All 10 LoRA adapters + Trackio logs are public so the numbers are checkable, not just claimed.
Full writeup: https://huggingface.co/blog/harshitkgupta/fine-tuning-coding-agents-on-mac-pytorch-mps-mlx
liked a model about 21 hours ago
huihui-ai/Huihui-Qwen3.8-27B-abliterated-GGUFOrganizations
None yet