TreeFlash Parallel AR-Approximation for Faster Speculative Decoding (https://arxiv.org/abs/2606.03819) peerrh/treeflash-qwen3-8b Text Generation • 1B • Updated Jun 18 • 16 peerrh/treeflash-qwen3-4b Text Generation • 0.7B • Updated Jun 18 • 20 peerrh/treeflash-qwen3-coder-30b-a3b Text Generation • 0.5B • Updated Jun 18 • 14
SD-Square peerrh/qwen3-14b-distilled Updated Nov 13, 2025 • 8 peerrh/qwen3-14b-steered Updated Nov 13, 2025 • 17 peerrh/qwen3-8b-distilled Updated Nov 13, 2025 • 8 peerrh/qwen3-8b-steered Updated Nov 13, 2025 • 18
TreeFlash Parallel AR-Approximation for Faster Speculative Decoding (https://arxiv.org/abs/2606.03819) peerrh/treeflash-qwen3-8b Text Generation • 1B • Updated Jun 18 • 16 peerrh/treeflash-qwen3-4b Text Generation • 0.7B • Updated Jun 18 • 20 peerrh/treeflash-qwen3-coder-30b-a3b Text Generation • 0.5B • Updated Jun 18 • 14
SD-Square peerrh/qwen3-14b-distilled Updated Nov 13, 2025 • 8 peerrh/qwen3-14b-steered Updated Nov 13, 2025 • 17 peerrh/qwen3-8b-distilled Updated Nov 13, 2025 • 8 peerrh/qwen3-8b-steered Updated Nov 13, 2025 • 18