Running 2 RIS-Kernel Long Context Inference Demo 🔬 2 Run long-context LLM inference with O(N log N) sparse attention
unsloth/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-GGUF Text Generation • 33B • Updated Aug 13 • 68.5k • 102
Running Featured 245 Gemma 4 WebGPU 🚀 245 Run Gemma 4 locally in-browser on WebGPU w/ Transformers.js