# About Evaluation platform for web agents on an **anonymized** benchmark (**152** tasks: 65 base tasks + 87 clones of UI patterns and the dark theme; **six** hub domains). The hub home is navigation, not a separate task set. Mock sites and UI wording are not tied to real brands. Details: [README.md](https://github.com/ai-forever/WebPageBench/blob/main/README.md). Runs in `results/` are the **152**-task canon (passed tasks / 152; tasks that were not run count as 0). Pilot 51-task runs have been removed. ## Benchmark Base scenarios (clones inherit the source section): | Section | Key | Tasks | |---------|-----|------:| | Market | `shop` | 18 | | Books | `books` | 11 | | Grocery | `grocery` | 8 | | Rail | `rail` | 9 | | Hotels | `hotels` | 8 | | Files | `files` | 11 | | **Base** | | **65** | | + UI clones and `theme:dark` | | 87 | | **Canon** | | **152** | The agent runs through a harness with a chosen LLM on anonymized mock sites. DOM values of `AGENT_HARNESS`: `browser-use`, `ouroboros-cut`, `ouroboros-full-isolated`, `ouroboros-full-evolving`, `hermes-ouroboros`, `openmanus`, `openhands`, `deepagents` Screenshot GUI harnesses: `qwen3-vl`, `uitars`, `jedi`, `opencua`, `evocua`, `fara`. Input the harness sends to the model in a run: | Input | When | |-------|------| | `text` | `openhands`; text-only models on Ouroboros (a screenshot is not injected into the LLM); DeepSeek on `browser-use` / `openmanus` (DOM only) | | `text+image` | `browser-use`, `openmanus`, `ouroboros-cut`, `ouroboros-full-*` when the model sees a screenshot | | `image` | GUI harnesses: `qwen3-vl`, `uitars`, `opencua`, `evocua`, `fara` | ## Leaderboard metrics ### Primary: Success Rate | Field | Source | Description | |-------|--------|-------------| | `success_rate` | `run.success_rate` | `passed_tasks / total_tasks` | | Task success | `tests[].success` | Every condition passed (`dab_check.all_passed`) | Do not confuse this with **Agent Completion**. The agent may have called `done` while the task conditions are not satisfied. ### Speed and efficiency | Field | Description | |-------|-------------| | `total_duration_seconds` | Sum of per-task durations. The Speed view sorts by this value (lower is better). | | `avg_duration_seconds` | Mean time per task (create_track + agent + check) | | `avg_agent_steps` | Mean agent steps (browser-use and compatible harnesses) | | `avg_tokens_per_task` | Mean tokens per task | | `total_cost_usd` | Estimated LLM cost (`bench_eval.llm_cost`) | ### Reliability | Field | Description | |-------|-------------| | `agent_completion_rate` | Share of tasks with `agent_is_done` | | `agent_dab_agreement_rate` | Agreement between completion and the condition check | | Done & pass | Tasks where `agent_is_done` and every condition passed (`tests[].success`). Stricter than Success Rate: a pass without `done`, or `done` without a pass, does not count. | ### UI taxonomy Aggregates in `ui_taxonomy_stats` inside `results.json`: - `by_checked_class` — classes named in `conditions` (BASKET, DATE, NAV, …) - `by_primary` — the task's primary UI class - `by_ui_pattern` — widget patterns (`date:popup_grid`, …) The leaderboard shows the top UI badges from `by_primary`. ## Leaderboard views 1. **Success Rate** — primary ranking 2. **Speed** — `total_duration_seconds` (lower is better). Average time is the mean per-task duration. 3. **Cost** — cost and value (success / $) **Input** (`text` / `text+image` / `image`) and **Harness** filters rebuild the main table, sections, UI taxonomy, and Metrics. ## Links - [Repository](https://github.com/ai-forever/WebPageBench) - [Tasks](https://github.com/ai-forever/WebPageBench/blob/main/bench_config.json) - [About this benchmark](https://github.com/ai-forever/WebPageBench/blob/main/docs/description.md)