-
CUGA Agent
🤖105Configurable Generalist Agent, leader in AppWorld Benchmark
-
MiroEval: Benchmarking Multimodal Deep Research Agents in Process and Outcome
Paper • 2603.28407 • Published • 72 -
How Well Do Agentic Skills Work in the Wild: Benchmarking LLM Skill Usage in Realistic Settings
Paper • 2604.04323 • Published • 41
David PRO
AustinOS
·
AI & ML interests
yes
Organizations
None yet
good
- RunningFeatured105
CUGA Agent
🤖105Configurable Generalist Agent, leader in AppWorld Benchmark
-
MiroEval: Benchmarking Multimodal Deep Research Agents in Process and Outcome
Paper • 2603.28407 • Published • 72 -
How Well Do Agentic Skills Work in the Wild: Benchmarking LLM Skill Usage in Realistic Settings
Paper • 2604.04323 • Published • 41