Five Raters, One Rule, Five Different Answers: What Happened When We Measured LLM Annotation Agreement leventbulut • 4 days ago • 1
Fast memory, slow weights: how a production agent system improves around the model, and the model itself pavle-scalably • 4 days ago • 1
Prefix cache, not throughput: 14 days serving Qwen3.8-27B NVFP4 to production agents on two RTX 5090s pavle-scalably • 5 days ago
Governed Agent Memory with Structured Judgment: An AtMem–Jev Retrieval Study javadtaghia • 5 days ago
From Coding Agents to Physical Agents: Code as a New Interface Between AI and the Physical World kkakkkka • 6 days ago • 3
BananaMind 2 Pro: We've (almost) matched SmolLM2 at 20x fewer tokens... Trained On a 5070 Ti Banaxi-Tech • 6 days ago • 6