view post Post 4345 🚀 Introducing Halo 1.0Today, we are open-sourcing Halo, the training framework we use to train every model at White Circle.It comes with: 🧠 Full post-training stack: SFT, DPO/KTO/SMPO, reward modeling, GRPO, distillation🤖 Async multi-turn RL with vLLM/SGLang rollouts and sandboxed tool use⚡ ~2.8× TRL throughput on 8× B300 (EP+FSDPv2, FA4, fp8/fp4)🤗 Dense HF models + 15 MoE families (Qwen, GLM, Mistral, DeepSeek-V4…)🛠️ One halo command, prebuilt Docker images, and docs for humans and agents💻 https://github.com/whitecircle/haloTry it and tell us what you're training See translation 1 reply · 🚀 9 9 ❤️ 4 4 + Reply
view post Post 73 43% less answer time in a looped transformer by changing answer supervisioncheck out https://x.com/advprop/status/2098087083470373010?s=20 + full writeup linked there See translation 🔥 1 1 + Reply
Vikhr: The Family of Open-Source Instruction-Tuned Large Language Models for Russian Paper • 2405.13929 • Published May 22, 2024 • 55