Spaces:
Running
Download README.md from EnvLoop/Pev-Leaderboard: direct link, hf CLI and curl.
- Browser
- Download file 1.43 kB
-
https://huggingface.co/spaces/EnvLoop/Pev-Leaderboard/resolve/main/README.md
- Command line
-
hf download hf://spaces/EnvLoop/Pev-Leaderboard/README.md
-
curl -L -o README.md https://huggingface.co/spaces/EnvLoop/Pev-Leaderboard/resolve/main/README.md
title: Pev Leaderboard
emoji: 馃搳
colorFrom: blue
colorTo: indigo
sdk: static
app_file: index.html
pinned: false
license: cc-by-4.0
short_description: Personal-agent decision benchmark results and comparisons
tags:
- leaderboard
- personal-agent
- decision-making
- calibration
- benchmark
Pev Leaderboard
Results on Pev-Bench, a pre-registered benchmark of seven families of short decision questions for a personal
agent with long-term memory (approvals, sharing, forgetting, memory relevance, notification level, routing and option
choice). The page shows family-macro accuracy, calibration and automation coverage for six models on the public TEST
set (two renderer halves, every model run once), paired comparisons with 95% confidence intervals and Holm-adjusted
p-values, per-family results, the one-shot HIDDEN gate and the DEV references, with the known caveats
(pick_option remains unsolved). The page content is licensed under CC-BY-4.0; the model and the dataset are gated
and released for non-commercial research only.
Paper (PDF) 路 Model 路 Dataset 路 Code 路 Collection
EnvLoop Research, research@envloop.ai