view article Article BenchMIRT: What are LLM benchmarks actually measuring? allenai • 25 days ago • 26
view article Article Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps +1 iamleonie, burtenshaw, sergiopaniego • 24 days ago • 137
view article Article NeoMME: an efficient Multimodal-native and Multilingual Encoder Hcompany • 23 days ago • 111
view article Article Introducing @huggingface/kernels: 200+ WebGPU Kernels for Local AI nico-martin, Xenova • 26 days ago • 82
view article Article State of Open Models: Summer 2026 Observations +1 AdinaY, multimodalart, irenesolaiman • Aug 14 • 214