--- title: Pev Leaderboard emoji: 📊 colorFrom: blue colorTo: indigo sdk: static app_file: index.html pinned: false license: cc-by-4.0 short_description: Personal-agent decision benchmark results and comparisons tags: - leaderboard - personal-agent - decision-making - calibration - benchmark --- # Pev Leaderboard Results on Pev-Bench, a pre-registered benchmark of seven families of short decision questions for a personal agent with long-term memory (approvals, sharing, forgetting, memory relevance, notification level, routing and option choice). The page shows family-macro accuracy, calibration and automation coverage for six models on the public TEST set (two renderer halves, every model run once), paired comparisons with 95% confidence intervals and Holm-adjusted p-values, per-family results, the one-shot HIDDEN gate and the DEV references, with the known caveats (`pick_option` remains unsolved). The page content is licensed under CC-BY-4.0; the model and the dataset are gated and released for non-commercial research only. [Paper (PDF)](https://github.com/EnvLoop/Pev/blob/main/docs/tech-report/pev.pdf) · [Model](https://huggingface.co/EnvLoop/Pev-27B-LoRA) · [Dataset](https://huggingface.co/datasets/EnvLoop/Pev-Bench) · [Code](https://github.com/EnvLoop/Pev) · [Collection](https://huggingface.co/collections/EnvLoop/pev-6abf32b6b1940477ad4c52c3) EnvLoop Research, research@envloop.ai