bottlecapai/ThinkingCap-Qwen3.8-27B-NVFP4A4-AWQ Image-Text-to-Text • 20B • Updated about 3 hours ago • 110 • 9
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression Paper • 2609.19969 • Published 9 days ago • 179
ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search Paper • 2609.13356 • Published 15 days ago • 262
Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation Paper • 2609.11638 • Published 16 days ago • 700
Unlocking Lossless Speedups in LLMs via Discrete Diffusion Paper • 2609.04010 • Published 23 days ago • 113
Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States Paper • 2609.04196 • Published 23 days ago • 70
Rethinking On-Policy Distillation of Large Language Models II: One Training Example Paper • 2609.04172 • Published 23 days ago • 101
Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM Paper • 2609.04098 • Published 23 days ago • 84
Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning Paper • 2609.03430 • Published 23 days ago • 187
LatentPress: Context Compression Beyond Text and Vision Paper • 2609.01507 • Published 25 days ago • 120
Running on CPU Upgrade Featured 143 H3 Acceleration Arena 🥇 143 Blind A/B ranking of MiniMax-H3 acceleration variants
H3-World: Turning Language Understanding into World Control Paper • 2609.01560 • Published 25 days ago • 52