AutomatosX/AX-Qwen3.6-35B-A3B-MLX-OptiQ-4bit-MTP Image-Text-to-Text • 35B • Updated 8 days ago • 926 • 2
MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Paper • 2607.11562 • Published 19 days ago • 77
Does VLA Even Know the Basics? Measuring Commonsense and World Knowledge Retention in Vision-Language-Action Models Paper • 2606.19297 • Published Jun 17 • 80
AnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language Models Paper • 2607.02269 • Published 30 days ago • 9
Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding Paper • 2605.29707 • Published May 28 • 152
Crafter: A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs Paper • 2605.30611 • Published May 28 • 253
stefanocarrera/autophagycode_D_he_train-mercury_Qwen3-4B_strategy_surplexity_t1_g5_run1_metrics Viewer • Updated Jun 2 • 164 • 79 • 1