Pev-Leaderboard / README.md
EnvLoopBot's picture
Upload 2 files
177849c verified
|
Raw History Blame Contribute Delete
1.43 kB
metadata
title: Pev Leaderboard
emoji: 馃搳
colorFrom: blue
colorTo: indigo
sdk: static
app_file: index.html
pinned: false
license: cc-by-4.0
short_description: Personal-agent decision benchmark results and comparisons
tags:
  - leaderboard
  - personal-agent
  - decision-making
  - calibration
  - benchmark

Pev Leaderboard

Results on Pev-Bench, a pre-registered benchmark of seven families of short decision questions for a personal agent with long-term memory (approvals, sharing, forgetting, memory relevance, notification level, routing and option choice). The page shows family-macro accuracy, calibration and automation coverage for six models on the public TEST set (two renderer halves, every model run once), paired comparisons with 95% confidence intervals and Holm-adjusted p-values, per-family results, the one-shot HIDDEN gate and the DEV references, with the known caveats (pick_option remains unsolved). The page content is licensed under CC-BY-4.0; the model and the dataset are gated and released for non-commercial research only.

Paper (PDF) 路 Model 路 Dataset 路 Code 路 Collection

EnvLoop Research, research@envloop.ai