Spaces:
Running
Download docs/description.md from ai-forever/WebPageBench: direct link, hf CLI and curl.
- Browser
- Download file 3.83 kB
-
https://huggingface.co/spaces/ai-forever/WebPageBench/resolve/main/docs/description.md
- Command line
-
hf download hf://spaces/ai-forever/WebPageBench/docs/description.md
-
curl -L -o description.md https://huggingface.co/spaces/ai-forever/WebPageBench/resolve/main/docs/description.md
A newer version of the Gradio SDK is available: 6.30.0
About
Evaluation platform for web agents on an anonymized benchmark (152 tasks: 65 base tasks + 87 clones of UI patterns and the dark theme; six hub domains). The hub home is navigation, not a separate task set. Mock sites and UI wording are not tied to real brands. Details: README.md.
Runs in results/ are the 152-task canon (passed tasks / 152; tasks that were not run count as 0). Pilot 51-task runs have been removed.
Benchmark
Base scenarios (clones inherit the source section):
| Section | Key | Tasks |
|---|---|---|
| Market | shop |
18 |
| Books | books |
11 |
| Grocery | grocery |
8 |
| Rail | rail |
9 |
| Hotels | hotels |
8 |
| Files | files |
11 |
| Base | 65 | |
+ UI clones and theme:dark |
87 | |
| Canon | 152 |
The agent runs through a harness with a chosen LLM on anonymized mock sites. DOM values of AGENT_HARNESS:
browser-use, ouroboros-cut, ouroboros-full-isolated, ouroboros-full-evolving, hermes-ouroboros, openmanus, openhands, deepagents
Screenshot GUI harnesses: qwen3-vl, uitars, jedi, opencua, evocua, fara.
Input the harness sends to the model in a run:
| Input | When |
|---|---|
text |
openhands; text-only models on Ouroboros (a screenshot is not injected into the LLM); DeepSeek on browser-use / openmanus (DOM only) |
text+image |
browser-use, openmanus, ouroboros-cut, ouroboros-full-* when the model sees a screenshot |
image |
GUI harnesses: qwen3-vl, uitars, opencua, evocua, fara |
Leaderboard metrics
Primary: Success Rate
| Field | Source | Description |
|---|---|---|
success_rate |
run.success_rate |
passed_tasks / total_tasks |
| Task success | tests[].success |
Every condition passed (dab_check.all_passed) |
Do not confuse this with Agent Completion. The agent may have called done while the task conditions are not satisfied.
Speed and efficiency
| Field | Description |
|---|---|
total_duration_seconds |
Sum of per-task durations. The Speed view sorts by this value (lower is better). |
avg_duration_seconds |
Mean time per task (create_track + agent + check) |
avg_agent_steps |
Mean agent steps (browser-use and compatible harnesses) |
avg_tokens_per_task |
Mean tokens per task |
total_cost_usd |
Estimated LLM cost (bench_eval.llm_cost) |
Reliability
| Field | Description |
|---|---|
agent_completion_rate |
Share of tasks with agent_is_done |
agent_dab_agreement_rate |
Agreement between completion and the condition check |
| Done & pass | Tasks where agent_is_done and every condition passed (tests[].success). Stricter than Success Rate: a pass without done, or done without a pass, does not count. |
UI taxonomy
Aggregates in ui_taxonomy_stats inside results.json:
by_checked_classβ classes named inconditions(BASKET, DATE, NAV, β¦)by_primaryβ the task's primary UI classby_ui_patternβ widget patterns (date:popup_grid, β¦)
The leaderboard shows the top UI badges from by_primary.
Leaderboard views
- Success Rate β primary ranking
- Speed β
total_duration_seconds(lower is better). Average time is the mean per-task duration. - Cost β cost and value (success / $)
Input (text / text+image / image) and Harness filters rebuild the main table, sections, UI taxonomy, and Metrics.