SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Paper • 2608.10692 • Published Aug 11 • 13
SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Paper • 2608.10692 • Published Aug 11 • 13
SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Paper • 2608.10692 • Published Aug 11 • 13
AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World Environments Paper • 2607.05174 • Published Jul 6 • 1
LLMEval-Logic: A Solver-Verified Chinese Benchmark for Logical Reasoning of LLMs with Adversarial Hardening Paper • 2605.19597 • Published May 19 • 20