Hugging Face's logo Hugging Face
  • Models
  • Datasets
  • Spaces
  • Buckets new
  • Docs
  • Enterprise
  • Pricing
    • Website
      • Tasks
      • HuggingChat
      • Collections
      • Languages
      • Organizations
    • Community
      • Blog
      • Posts
      • Daily Papers
      • Hardware
      • Learn
      • Discord
      • Forum
      • GitHub
    • Solutions
      • Team & Enterprise
      • Hugging Face PRO
      • Enterprise Support
      • Inference Providers
      • Inference Endpoints
      • Storage Buckets

  • Log In
  • Sign Up
Nitya Gupta's picture

Nitya Gupta

nitya-gupta
1
·

AI & ML interests

None yet

Recent Activity

reacted to harshitkgupta's post with 🚀 about 22 hours ago
Fine-tuned Qwen 2.5 (0.5B → 3B) on real coding-agent traces, 10 controlled runs, one 16GB Mac. Compared PyTorch MPS vs. Apple MLX for local LoRA SFT — and the honest answer is "it depends on what you're optimizing for": • PyTorch MPS: 2.2x–5.7x faster raw throughput, but hits a hard memory wall — can't load a 3B model in FP16 on 16GB. • Apple MLX: 4-bit QLoRA fits 3B+ models with almost flat memory scaling as context grows (+109 MB going from 1k→4k tokens). • 4-bit quantization doesn't cost you convergence — eval loss tracks closely across backends. • The bigger surprise: most of MLX's slowdown isn't the 4-bit dequant tax. Two of the 10 runs went unquantized to isolate it — dequant only explains 1.07x–1.4x of the gap. A ~4.1–4.6x framework-level gap remains either way. All 10 LoRA adapters + Trackio logs are public so the numbers are checkable, not just claimed. Full writeup: https://huggingface.co/blog/harshitkgupta/fine-tuning-coding-agents-on-mac-pytorch-mps-mlx
upvoted an article about 22 hours ago
Fine-Tuning Coding Agents on a 16GB Mac: PyTorch MPS vs. Apple MLX
View all activity

Organizations

Solvren AI's profile picture

nitya-gupta 's datasets

None public yet
Company
TOS Privacy About Careers
Website
Models Datasets Spaces Pricing Docs