SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD Paper • 2607.20145 • Published 6 days ago • 63
Janus: Disaggregating Attention and Experts for Scalable MoE Inference Paper • 2512.13525 • Published Dec 15, 2025 • 6
Janus: Disaggregating Attention and Experts for Scalable MoE Inference Paper • 2512.13525 • Published Dec 15, 2025 • 6