Rider Jones
mazuj2
ยท
AI & ML interests
None yet
Recent Activity
reacted to FredyRivera-dev's post with โค๏ธ about 10 hours ago
We wrote a full technical guide on how to train a bilingual (ES/EN) LLM from scratch: TinyQwen.
Covers:
- Hybrid architecture based on Qwen3.5
- Pre-training with 15B tokens
- Cost benchmark between H200 and B200
- Post-training with SFT + LoRA
- Full code and data, open source
With ~$11 of compute on an H200 we ran an initial training run, enough to validate the full architecture and pipeline.
Blog post: https://aquiles-ai.vercel.app/blog/tinyqwen-from-scratch
Technical feedback welcome, especially from anyone looking to replicate the pipeline with more compute. liked a model 3 days ago
mindlab-research/Macaron-V1-Tall new activity 22 days ago
unsloth/Qwen3.6-27B-MTP-GGUF:FAST!!!! 39tps!Organizations
None yet