view article Article Holo4: powering generalist computer-use agents Hcompany • about 15 hours ago • 34
view article Article Transformers now runs llama.cpp quants +1 marcsun13, ArthurZ, lysandre • 7 days ago • 79
view article Article Jun Kim, oMLX creator and maintainer, joins Hugging Face to support the MLX community +3 pcuenq, lysandre, victor, julien-c, Jundot • 7 days ago • 75
DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence Paper • 2606.19348 • Published Apr 26 • 46
GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models Paper • 2508.06471 • Published Aug 8, 2025 • 215
Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation Paper • 2609.11115 • Published 19 days ago • 172
HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness? Paper • 2609.01437 • Published 28 days ago • 221
view article Article Measuring benchmark optimization in speech recognition +5 tlebryk02, bezzam, aliceebaird, dayllon, jpc, jens-hume-ai, tzirakis • Aug 21 • 67
view article Article We changed one line and the benchmark score moved 0.21 AUROC FINAL-Bench • Aug 22 • 15
view article Article Multi-Vector (Late Interaction) Embedding Models with Sentence Transformers +1 tomaarsen, NohTow, raphaelsty • Aug 18 • 117
SWE-bench Collection SWE-bench (Lite, Verified, Multimodal, Multilingual) all in one place! • 5 items • Updated Dec 14, 2025 • 17
view article Article State of Open Models: Summer 2026 Observations +1 AdinaY, multimodalart, irenesolaiman • Aug 14 • 216
view article Article LFM2.5-VL-3B for Better and Faster Vision Capabilities for the Edge LiquidAI • Aug 12 • 51