DFlash Collection Block Diffusion for Flash Speculative Decoding • 23 items • Updated 28 days ago • 151
view article Article Intel XPU Kernel Skill: LLM-driven Triton kernel optimization for the Hugging Face Kernel Hub danf • Jun 17 • 11
view article Article Getting More from Your Test-Time Compute Budget with Portfolio Beam Search danelbaz • Feb 24 • 8
Prune Once for All: Sparse Pre-Trained Language Models Paper • 2111.05754 • Published Nov 10, 2021 • 2
view article Article DeepMath: A lightweight math reasoning Agent with smolagents +1 danf, mber, moshew • Dec 4, 2025 • 40
view article Article Accelerating Qwen3-8B Agent on Intel® Core™ Ultra with Depth-Pruned Draft Models +3 imargulis, ofirzaf, sguskin, guybd, pcuenq • Sep 29, 2025 • 25
view article Article Accelerating Qwen3-8B Agent on Intel® Core™ Ultra with Depth-Pruned Draft Models +3 imargulis, ofirzaf, sguskin, guybd, pcuenq • Sep 29, 2025 • 25
view article Article Breaking Language Barriers in Mathematical AI: Introducing Hebrew Math Tutor danf • Sep 7, 2025 • 3