Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
11 changes: 6 additions & 5 deletions .github/workflows/benchmark.yml
Original file line number Diff line number Diff line change
Expand Up @@ -18,16 +18,16 @@ on:
default: 4.0.0-dev.64
playwright-mcp:
description: '@playwright/mcp version'
default: latest
default: 0.0.83
agent-browser:
description: 'agent-browser version'
default: latest
default: 0.38.2
playwright-cli:
description: '@playwright/cli version'
default: latest
default: 0.1.22
stagehand:
description: 'browserbase/stagehand git ref (branch, tag or sha) to build its Claude Code MCP server from; latest = newest release tag'
default: latest
default: cd7b230778cf92269e4cb90e80d97f5113781c51
model:
description: 'Model for every agent: Claude through Anthropic, others through OpenRouter (needs the OPENROUTER_API_KEY secret)'
type: choice
Expand Down Expand Up @@ -87,7 +87,8 @@ jobs:
name: Run and judge
needs: plan
runs-on: ubuntu-latest
# GitHub's maximum: 50 tasks × 7 setups take four to five hours
# GitHub's maximum: 50 tasks × 7 setups take four to five hours, so a
# full run of the 100 tasks is two workflow runs, 50 task ids each
timeout-minutes: 360
steps:
- uses: actions/checkout@v5
Expand Down
22 changes: 11 additions & 11 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@

How often does a coding agent get a browser task right, and what does it cost?

This repository runs a fixed sample of 50 tasks from [Online-Mind2Web](https://github.com/OSU-NLP-Group/Online-Mind2Web), tasks on live websites written by the benchmark's authors, against seven browser tool setups. The model, the agent harness and the prompt are the same for every setup. Results: [benchmark.webdriver.io](https://benchmark.webdriver.io).
This repository runs a fixed sample of 100 tasks from [Online-Mind2Web](https://github.com/OSU-NLP-Group/Online-Mind2Web), tasks on live websites written by the benchmark's authors, against seven browser tool setups. The model, the agent harness and the prompt are the same for every setup. Results: [benchmark.webdriver.io](https://benchmark.webdriver.io).

| Setup | What the agent gets | npm package |
|---|---|---|
Expand Down Expand Up @@ -70,12 +70,12 @@ The report prints medians per setup, plus a pass matrix per task.

[Online-Mind2Web](https://github.com/OSU-NLP-Group/Online-Mind2Web) (COLM 2025) is 300 tasks on 136 live websites, written by researchers at Ohio State and judged by the benchmark's own judge. Neither we nor any tool vendor chose or tuned for these tasks.

- **The sample:** 50 tasks, split across easy, medium and hard in the dataset's own proportions, drawn with a fixed seed from a fixed dataset revision ([`src/mind2web.ts`](src/mind2web.ts)). [`tasks/online-mind2web.json`](tasks/online-mind2web.json) lists their ids; anyone with the dataset can recompute it. The dataset is gated on Hugging Face, so its task texts stay out of this repository: the runner downloads them with `HF_TOKEN` (accept the [dataset terms](https://huggingface.co/datasets/osunlp/Online-Mind2Web) first).
- **The sample:** 100 tasks, split across easy, medium and hard in the dataset's own proportions, drawn with a fixed seed from a fixed dataset revision ([`src/mind2web.ts`](src/mind2web.ts)). [`tasks/online-mind2web.json`](tasks/online-mind2web.json) lists their ids; anyone with the dataset can recompute it. The dataset is gated on Hugging Face, so its task texts stay out of this repository: the runner downloads them with `HF_TOKEN` (accept the [dataset terms](https://huggingface.co/datasets/osunlp/Online-Mind2Web) first).
- **The prompt:** the task and its start page, plus one rule: don't sign in, create accounts, pay or enter personal data; stop right before that. Same for every setup.
- **Screenshots:** after every tool call the harness, not the agent, screenshots the page the agent is on, over the Chrome DevTools Protocol of the browser the run started ([`src/screenshots.ts`](src/screenshots.ts)). No agent pays tokens for them. Playwright (MCP and CLI) starts Chrome without a debugging port, so for this suite its config adds one; nothing else about any setup changes. We checked that capturing doesn't change tool behaviour (same tokens and results with and without, for Stagehand and agent-browser).
- **The judge:** [WebJudge](https://github.com/OSU-NLP-Group/Online-Mind2Web#-webjudge) with `o4-mini`, as its authors recommend (86% agreement with human reviewers), at a pinned commit ([`src/judge.ts`](src/judge.ts)). It sees the task, the agent's actions and the screenshots, not the agent's final answer, and comes from another model family than the agents, so it can't favour its own. The action history is exactly what the agent issued (the shell command or the tool call), never a tool's reply, so a tool with chattier output gains nothing. Two changes to running it, both mechanical: WebJudge sends `max_tokens=512` and `temperature=0`, which OpenAI's reasoning models reject (and 512 tokens would go to reasoning), so the one API call sends `max_completion_tokens` instead; and its worker processes need the `fork` start method, which a wrapper sets. The judge's reasoning for every run is published in `judgments-*.jsonl`, for spot checks.
- **Unjudged runs don't count:** a run the judge couldn't decide stays pending and is left out of every number, with a note in the report.
- **Uncertainty is shown:** 50 tasks can't separate close results. The website shows a 95% confidence interval under every success rate.
- **Uncertainty is shown:** 100 tasks can't separate close results either. The website shows a 95% confidence interval under every success rate.

Limitations to keep in mind:

Expand Down Expand Up @@ -108,18 +108,18 @@ Start the [Benchmark workflow](../../actions/workflows/benchmark.yml) with **Run
|---|---|---|
| `webdriverio` | `10.0.0-alpha.175` | `@wdio/cli` version for `wdio-session` (v10 and up; `latest` is still v9, which has no `wdio session`) |
| `wdio-mcp` | `4.0.0-dev.64` | `@wdio/mcp` version |
| `playwright-mcp` | `latest` | `@playwright/mcp` version, for both Playwright setups |
| `agent-browser` | `latest` | `agent-browser` version |
| `playwright-cli` | `latest` | `@playwright/cli` version |
| `stagehand` | `latest` | git ref of `browserbase/stagehand` to build (branch, tag or sha); `latest` is the newest `@browserbasehq/stagehand@x.y.z` release tag |
| `playwright-mcp` | `0.0.83` | `@playwright/mcp` version, for both Playwright setups |
| `agent-browser` | `0.38.2` | `agent-browser` version |
| `playwright-cli` | `0.1.22` | `@playwright/cli` version |
| `stagehand` | `cd7b230…` | git ref of `browserbase/stagehand` to build (branch, tag or sha); `latest` is the newest `@browserbasehq/stagehand@x.y.z` release tag |
| `model` | `claude-sonnet-5` | model for every agent: `claude-sonnet-5`, or `deepseek-flash-4-1` through OpenRouter (needs the `OPENROUTER_API_KEY` repository secret) |
| `runs` | `1` | runs per task and setup |
| `setups`, `tasks` | `all` | comma-separated ids to run a subset |
| `seed` | `1` | seed for the run order |
| `concurrency` | `4` | runs at the same time |
| `publish` | on | commit the results to this repository |

npm versions accept an exact version, a dist-tag (`latest`, `next`) or a range. A setup whose tool can't be installed is skipped, and the report says why.
Every default is pinned to the version of the published results, so a new run adds to their rows instead of starting new ones. npm versions accept an exact version, a dist-tag (`latest`, `next`) or a range. A setup whose tool can't be installed is skipped, and the report says why.

The workflow runs every setup in one job, has WebJudge judge every run, then publishes:

Expand All @@ -138,12 +138,12 @@ Requires Node.js 24, Chrome and `python3` (for the judge). The Agent SDK picks u
npm install
npm run bench -- --dry-run # print the shuffled plan
npm run bench -- --setups wdio-session,playwright-mcp --tasks om2w-180ed2ec
npm run bench # everything: 7 setups × 50 tasks
npm run bench # everything: 7 setups × 100 tasks
node src/judge.ts results/<id> # WebJudge decides every run, writes judgments-*.jsonl
node src/publish.ts results/<id> # report.md, results index, README section
```

`node src/mind2web.ts sample` picks the 50 tasks again (same seed, same sample) and writes `tasks/online-mind2web.json`.
`node src/mind2web.ts sample` picks the 100 tasks again (same seed, same sample; a larger `--size` keeps every task of a smaller one) and writes `tasks/online-mind2web.json`.

Pick versions with `WDIO_VERSION`, `WDIO_MCP_VERSION`, `PLAYWRIGHT_MCP_VERSION`, `PLAYWRIGHT_CLI_VERSION`, `AGENT_BROWSER_VERSION` and `STAGEHAND_REF` (default `latest`, except `WDIO_VERSION` and `WDIO_MCP_VERSION`: `wdio session` ships with v10, and `@wdio/mcp` 4 is still a dev build; see the workflow inputs above for the current defaults). To test an unreleased WebdriverIO:

Expand All @@ -153,7 +153,7 @@ export WDIO_LOCAL=/path/to/webdriverio # a built checkout of webdriverio/webdr

Runner options: `--setups`, `--tasks`, `--runs` (default 1), `--model` (default `claude-sonnet-5`), `--seed`, `--out-dir`, `--max-turns` (default 80), `--timeout-min` (default 10), `--concurrency` (default 1; parallel browsers compete for CPU, which shows in the time per task), `--screenshots` (always on: WebJudge decides from them).

A full run is 350 agent runs and takes four to five hours at concurrency 4.
A full run is 700 agent runs. 350 take four to five hours at concurrency 4, and a GitHub job stops after six, so in the workflow a full run is two runs with 50 task ids each (`tasks`). Runs of the same tool versions and model count toward one row on the website.

### Stagehand

Expand Down
10 changes: 8 additions & 2 deletions site/app.js
Original file line number Diff line number Diff line change
Expand Up @@ -93,8 +93,14 @@ const swatch = (id) => `<i class="swatch" style="background:${colorOf(id)}" aria

const state = { suite: '', model: '', axis: 'cost', sort: { key: 'rank', dir: 1 } }

/** a version that ran every task of its suite, not a pilot on a few of them */
const isComplete = (g) => Object.keys(g.perTask).length >= (data.tasks[g.suite] ?? []).length
/**
* A version that ran as many tasks as any version did with this model, not a
* pilot on a few of them. Per model, so a larger sample that one model has
* already run doesn't turn the other model's results into pilots.
*/
const isComplete = (g) => Object.keys(g.perTask).length >= Math.max(...data.groups
.filter((o) => o.suite === g.suite && o.model === g.model)
.map((o) => Object.keys(o.perTask).length))

/** the newest fully tested version of each tool (or its newest, if none is), ranked by success, then cost */
function latestGroups () {
Expand Down
6 changes: 3 additions & 3 deletions site/index.html
Original file line number Diff line number Diff line change
Expand Up @@ -11,7 +11,7 @@
<meta property="og:site_name" content="WebdriverIO">
<meta property="og:url" content="https://benchmark.webdriver.io/">
<meta property="og:title" content="Browser Agent Benchmark">
<meta property="og:description" content="How often do AI agents get a browser task right, and what does it cost? WebdriverIO, Playwright, Stagehand and agent-browser on 50 live-web tasks from Online-Mind2Web.">
<meta property="og:description" content="How often do AI agents get a browser task right, and what does it cost? WebdriverIO, Playwright, Stagehand and agent-browser on 100 live-web tasks from Online-Mind2Web.">
<meta name="twitter:card" content="summary">
<meta name="twitter:title" content="Browser Agent Benchmark">
<meta name="twitter:description" content="Success rate, cost, tokens and time of browser tools for AI agents on live-web tasks. Open data, every run published.">
Expand Down Expand Up @@ -110,7 +110,7 @@ <h2 id="method-title">Methodology</h2>
<div class="method-grid">
<article class="card">
<h3>Tasks we didn't write</h3>
<p>The tasks are a fixed random sample of 50 tasks from <a href="https://github.com/OSU-NLP-Group/Online-Mind2Web">Online-Mind2Web</a> (COLM 2025): 300 tasks on 136 live websites, written by researchers at Ohio State. The sample keeps the dataset's mix of easy, medium and hard tasks and uses a fixed seed, so anyone can recompute it. No tool vendor, including us, picked these tasks.</p>
<p>The tasks are a fixed random sample of 100 tasks from <a href="https://github.com/OSU-NLP-Group/Online-Mind2Web">Online-Mind2Web</a> (COLM 2025): 300 tasks on 136 live websites, written by researchers at Ohio State. The sample keeps the dataset's mix of easy, medium and hard tasks and uses a fixed seed, so anyone can recompute it. No tool vendor, including us, picked these tasks.</p>
</article>
<article class="card">
<h3>An independent judge</h3>
Expand All @@ -131,7 +131,7 @@ <h3>What we measure</h3>
<h3>Limitations</h3>
<ul class="tight">
<li>Live websites change, block bots and show CAPTCHAs, so runs are not exactly repeatable. A blocked run counts as a failure for every tool it hits.</li>
<li>50 tasks and one run per task give intervals about ±13 points wide. Close results are ties.</li>
<li>100 tasks and one run per task give intervals about ±10 points wide. Close results are ties.</li>
<li>WebJudge disagrees with human reviewers on about one run in seven.</li>
<li>A tool that does several steps in one call leaves fewer screenshots for the judge to check.</li>
<li>This benchmark is maintained by the WebdriverIO project, which builds two of the tools. That is why the code, prompts, raw data and judgments are all public. If a setup looks wrong, <a href="https://github.com/webdriverio/benchmark/issues">tell us</a>.</li>
Expand Down
6 changes: 3 additions & 3 deletions src/mind2web.ts
Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,7 @@
* we commit only the ids of the sampled tasks and download the rest at run
* time.
*
* node src/mind2web.ts sample [--size 50] [--seed 1] pick the sample and write tasks/online-mind2web.json
* node src/mind2web.ts sample [--size 100] [--seed 1] pick the sample and write tasks/online-mind2web.json
*/
import fs from 'node:fs/promises'
import path from 'node:path'
Expand Down Expand Up @@ -108,10 +108,10 @@ export function sample (tasks: Mind2WebTask[], size: number, seed: number): Samp
if (import.meta.url === `file://${process.argv[1]}`) {
const { positionals, values } = parseArgs({
allowPositionals: true,
options: { size: { type: 'string', default: '50' }, seed: { type: 'string', default: '1' } }
options: { size: { type: 'string', default: '100' }, seed: { type: 'string', default: '1' } }
})
if (positionals[0] !== 'sample') {
console.error('usage: node src/mind2web.ts sample [--size 50] [--seed 1]')
console.error('usage: node src/mind2web.ts sample [--size 100] [--seed 1]')
process.exit(1)
}
const tasks = await loadDataset()
Expand Down
200 changes: 200 additions & 0 deletions tasks/online-mind2web.json
Original file line number Diff line number Diff line change
Expand Up @@ -55,6 +55,62 @@
"task_id": "f00e7accfb4a5e09680bdb326e6274ad",
"level": "easy"
},
{
"task_id": "271b36efd4346721b5542488ff997042",
"level": "easy"
},
{
"task_id": "95cad96f2e43f3c0d8efad1331c77c8c",
"level": "easy"
},
{
"task_id": "7b182a5087347d494b48a29dbc0f1d3e",
"level": "easy"
},
{
"task_id": "f2097f92a10d42a842c14179f422311e",
"level": "easy"
},
{
"task_id": "ade4c09ad3fdb1607209750924cd232f",
"level": "easy"
},
{
"task_id": "bb314cb80f0f8489135cbf59074d11e2",
"level": "easy"
},
{
"task_id": "a96fca87a17d792644e736d1d10d3cbe",
"level": "easy"
},
{
"task_id": "bf3b311cc8dce16d3de844f4b5875dfd",
"level": "easy"
},
{
"task_id": "a11ecdff735b51372d536c866011af6f",
"level": "easy"
},
{
"task_id": "bb518416a786fdb9b9bbf0c78515595e",
"level": "easy"
},
{
"task_id": "461ab9b0c7b20ac5f912704480979c65",
"level": "easy"
},
{
"task_id": "783ce6a3499fa7cf25bc12f8f0ecbbbb",
"level": "easy"
},
{
"task_id": "52efbab520734ef9bf7c09ba0f62cdc8",
"level": "easy"
},
{
"task_id": "a172a5d9ffaf5ef02bd550ec4fe24e6d",
"level": "easy"
},
{
"task_id": "9d46ccb915eff39ee1ae1e7328f5f20d",
"level": "medium"
Expand Down Expand Up @@ -151,6 +207,98 @@
"task_id": "864244b6969e0f8733b0eb1ca06cd51f",
"level": "medium"
},
{
"task_id": "07ec4a12cba8090e2dc524d558ac7675",
"level": "medium"
},
{
"task_id": "63d6866fc000fcb1f153e07604bd1395",
"level": "medium"
},
{
"task_id": "6ca20f1da01edeb49a7a42c816d8c6fe",
"level": "medium"
},
{
"task_id": "9f1cba613830ca1c6a58f9498c06e679_110325",
"level": "medium"
},
{
"task_id": "aa4b5cb7114fcc138ade82b4b9716d24",
"level": "medium"
},
{
"task_id": "8689af4d33ce00bf2cdd8987d3bbfd86",
"level": "medium"
},
{
"task_id": "43a1ca251f11c6b0bdd0379766cc49e6",
"level": "medium"
},
{
"task_id": "da8f3823a827c7d3a492f383808e7912",
"level": "medium"
},
{
"task_id": "47bfe8a7e0e4e7efc837287b407fbe90",
"level": "medium"
},
{
"task_id": "330cd04c773ac498f51afa4665461ec8",
"level": "medium"
},
{
"task_id": "b922508886ded315c9835457a6eb43ea",
"level": "medium"
},
{
"task_id": "d9d8b7d84a3f8d057e368254fe8d65e2",
"level": "medium"
},
{
"task_id": "5d542a7ec1fa142ba73cc87d970caf39",
"level": "medium"
},
{
"task_id": "75146b7b67388b9244e0f21a1527c022",
"level": "medium"
},
{
"task_id": "2fc51dd3febd447f0fdcdabca8d944ce_110325",
"level": "medium"
},
{
"task_id": "47e314cc452c540524ffb7cf520285a3",
"level": "medium"
},
{
"task_id": "9d09bc948462db032bac98968b11b008",
"level": "medium"
},
{
"task_id": "8103786e0e5976ebf961bd062d5f39cd",
"level": "medium"
},
{
"task_id": "92a3d4236f167af4afdc08876a902ba6",
"level": "medium"
},
{
"task_id": "6b2cfae0ef25c73d1224b6ab74cb8b63",
"level": "medium"
},
{
"task_id": "a8b9edd598561d2de901864d5f40fe67",
"level": "medium"
},
{
"task_id": "e9f4dfc67e0e6aa37f05f7cc5aa7428c",
"level": "medium"
},
{
"task_id": "48c73f3f53e2611c4a1052457c1033db",
"level": "medium"
},
{
"task_id": "27fa3ac20745d3d35e89fae157f63069",
"level": "hard"
Expand Down Expand Up @@ -202,6 +350,58 @@
{
"task_id": "1b867afecf072cb877ebfa4069263746",
"level": "hard"
},
{
"task_id": "f2be37a9a60fbc25b6b11cf622d17352_110325",
"level": "hard"
},
{
"task_id": "a0a18ca6a3529f3e97c771aadd42d3a0_110325",
"level": "hard"
},
{
"task_id": "1c3b747ae12ccee895745f82e3f2ef8a",
"level": "hard"
},
{
"task_id": "3ef64f34eae59c9fac7ee9a4f18b4a0c",
"level": "hard"
},
{
"task_id": "d1970c16271496cbbe166ecbecc0a1d8",
"level": "hard"
},
{
"task_id": "f05e87c5b92d9869e08806103c1c15a1",
"level": "hard"
},
{
"task_id": "c3a333968fc3c43d7f2688f425a0d633",
"level": "hard"
},
{
"task_id": "c801d1c951f59297f526bab84fa86c6e",
"level": "hard"
},
{
"task_id": "547f5729c59d5d12a457a3ebb74c31c6_010226",
"level": "hard"
},
{
"task_id": "dcd26e662a616d373ddd339747c6ce5b",
"level": "hard"
},
{
"task_id": "b64f938af842f6a1b4489d0e49a785a7_121125",
"level": "hard"
},
{
"task_id": "3443e9c3151fef19a3c3a45eb2c13640_070826",
"level": "hard"
},
{
"task_id": "905cb53061c33aa2d77e485fe1fca516",
"level": "hard"
}
]
}
Loading