VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks Paper • 2610.00972 • Published 5 days ago • 7
Typed decision models (System One) Collection Open models that answer yes/no, choice and score questions with probabilities, plus typed-decisions, a public benchmark to compare them. • 17 items • Updated 6 days ago • 3
Devstral 2 Collection A couple of agentic LLMs for software engineering tasks, excelling at using tools to explore codebases, edit multiple files, and power SWE Agents. • 2 items • Updated Mar 2 • 61
Mistral Large 3 Collection A state-of-the-art, open-weight, general-purpose multimodal model with a granular Mixture-of-Experts architecture. • 4 items • Updated Dec 2, 2025 • 104
Mistral Small 4 Collection A state-of-the-art model, open-weight, with a granular Mixture-of-Experts architecture that fuses instruct, reasoning and agentic skills. • 3 items • Updated Mar 16 • 82
Mistral Medium 3.5 Collection Our first flaship models handling instruction-following, reasoning, and coding in a single set of opened-weights. • 2 items • Updated Apr 29 • 24
Decision models Collection GGUF decision models for the /v1/systemone API in llama.cpp • 8 items • Updated about 15 hours ago • 15
LOCI: Spatial Linear Memory for Streaming World Models Paper • 2609.40222 • Published 5 days ago • 11
ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization Paper • 2610.00906 • Published 5 days ago • 66
Physis-Lang: Self-Evolving Language as a Physical Representation for Video World Model Paper • 2609.40358 • Published 6 days ago • 16
Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents Paper • 2609.39982 • Published 6 days ago • 113
PivotOPD: Learning to Recover from Pivotal Mistakes in Multi-Turn Agents Paper • 2609.40285 • Published 6 days ago • 21