A Lie Detector Test for Language Models: Reading Knowledge a Model Won't Reveal Paper • 2609.21996 • Published 7 days ago • 6
CADWorld: Computer-Use Benchmark for Long-Horizon Computer-Aided Design Paper • 2609.16251 • Published 11 days ago • 12
When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models Paper • 2609.19671 • Published 8 days ago • 49
SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness Paper • 2609.20519 • Published 8 days ago • 130
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression Paper • 2609.19969 • Published 8 days ago • 177
PACT: Can Enterprise AI Assistants Be Trusted Under Pressure? Paper • 2609.18605 • Published 9 days ago • 40
VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control Paper • 2609.19554 • Published 8 days ago • 41
LimiX-2: A Contextual Mechanism Network Towards General Structured-Data Intelligence Paper • 2609.17488 • Published 10 days ago • 728