WebPageBench / docs /description.md
ai-forever's picture
add fixes ui
cdb54bf
|
Raw History Blame Contribute Delete
3.83 kB

A newer version of the Gradio SDK is available: 6.30.0

Upgrade

About

Evaluation platform for web agents on an anonymized benchmark (152 tasks: 65 base tasks + 87 clones of UI patterns and the dark theme; six hub domains). The hub home is navigation, not a separate task set. Mock sites and UI wording are not tied to real brands. Details: README.md.

Runs in results/ are the 152-task canon (passed tasks / 152; tasks that were not run count as 0). Pilot 51-task runs have been removed.

Benchmark

Base scenarios (clones inherit the source section):

Section Key Tasks
Market shop 18
Books books 11
Grocery grocery 8
Rail rail 9
Hotels hotels 8
Files files 11
Base 65
+ UI clones and theme:dark 87
Canon 152

The agent runs through a harness with a chosen LLM on anonymized mock sites. DOM values of AGENT_HARNESS:

browser-use, ouroboros-cut, ouroboros-full-isolated, ouroboros-full-evolving, hermes-ouroboros, openmanus, openhands, deepagents

Screenshot GUI harnesses: qwen3-vl, uitars, jedi, opencua, evocua, fara.

Input the harness sends to the model in a run:

Input When
text openhands; text-only models on Ouroboros (a screenshot is not injected into the LLM); DeepSeek on browser-use / openmanus (DOM only)
text+image browser-use, openmanus, ouroboros-cut, ouroboros-full-* when the model sees a screenshot
image GUI harnesses: qwen3-vl, uitars, opencua, evocua, fara

Leaderboard metrics

Primary: Success Rate

Field Source Description
success_rate run.success_rate passed_tasks / total_tasks
Task success tests[].success Every condition passed (dab_check.all_passed)

Do not confuse this with Agent Completion. The agent may have called done while the task conditions are not satisfied.

Speed and efficiency

Field Description
total_duration_seconds Sum of per-task durations. The Speed view sorts by this value (lower is better).
avg_duration_seconds Mean time per task (create_track + agent + check)
avg_agent_steps Mean agent steps (browser-use and compatible harnesses)
avg_tokens_per_task Mean tokens per task
total_cost_usd Estimated LLM cost (bench_eval.llm_cost)

Reliability

Field Description
agent_completion_rate Share of tasks with agent_is_done
agent_dab_agreement_rate Agreement between completion and the condition check
Done & pass Tasks where agent_is_done and every condition passed (tests[].success). Stricter than Success Rate: a pass without done, or done without a pass, does not count.

UI taxonomy

Aggregates in ui_taxonomy_stats inside results.json:

  • by_checked_class β€” classes named in conditions (BASKET, DATE, NAV, …)
  • by_primary β€” the task's primary UI class
  • by_ui_pattern β€” widget patterns (date:popup_grid, …)

The leaderboard shows the top UI badges from by_primary.

Leaderboard views

  1. Success Rate β€” primary ranking
  2. Speed β€” total_duration_seconds (lower is better). Average time is the mean per-task duration.
  3. Cost β€” cost and value (success / $)

Input (text / text+image / image) and Harness filters rebuild the main table, sections, UI taxonomy, and Metrics.

Links