--- title: DarwinX emoji: 🧬 colorFrom: blue colorTo: indigo sdk: static app_file: index.html pinned: false license: cc-by-4.0 short_description: Evolving Agent Harnesses Through Natural Selection --- # DarwinX: Evolving Agent Harnesses Through Natural Selection Project page for [arXiv:2608.07545](https://arxiv.org/abs/2608.07545). Paper page: [huggingface.co/papers/2608.07545](https://huggingface.co/papers/2608.07545). Official DarwinX implementation (in Beagle): [github.com/SalesforceAIResearch/Beagle](https://github.com/SalesforceAIResearch/Beagle). DarwinX treats agent self-evolution as selection over a *population* of harnesses with the base model frozen. Across four benchmarks, one loop adds ~17 points on average: | Benchmark | Result | Gain | What it isolates | | --- | --- | --- | --- | | Terminal-Bench 2.1 (avg@5) | 84.7% | +7.7 on matched base | in-domain evolution | | TerminalWorld (held-out) | 68.3% | +7.3 | disjoint held-out task split | | WebArena-Infinity (audit-clean pass@1) | 93.0% | +49.5 pp | synthetic → real intent shift | | SWE-bench Verified (zero-shot transfer) | 84.2% | +3.4 | cross-benchmark transfer | ## Page sections `Demo` · `Abstract` · `Method` · `Results at a glance` · `Benchmark detail` · `Ablation: what evolution changes` · `Limitations` · `BibTeX` The demo is a 17-second animated walkthrough of the loop (`assets/demo.mp4`, 1080p H.264, 2.9 MB, faststart so it streams rather than buffering). It sits under the teaser, autoplays muted on loop, and carries narration for anyone who unmutes. It was re-encoded from a 2560×1440 master, which is kept out of the Space: at 9.99 MiB the master is within kilobytes of Hugging Face's 10 MiB limit for files not tracked with LFS. To rebuild it from a new master: ``` ffmpeg -i master.mp4 -vf "scale=1920:-2:flags=lanczos" \ -c:v libx264 -profile:v high -level 4.0 -pix_fmt yuv420p -crf 23 -preset slow \ -c:a aac -b:a 128k -ac 2 -movflags +faststart assets/demo.mp4 ffmpeg -i master.mp4 -vf "thumbnail=48,scale=1920:-2" -frames:v 1 -q:v 3 assets/demo_poster.jpg ``` The ablation section reproduces the paper's framing: it is an *exploratory attribution, not a per-skill causal ablation*, since the skills were co-selected rather than independently randomized. The limitations section is carried over from the paper's Discussion. ## Interactive figures Eight figures are interactive, and each is driven by real run artifacts or the paper's own plotting scripts rather than illustrative numbers. `tools/build_data.py` regenerates `assets/data.js` from the original files, so nothing is hand-transcribed. **1. TerminalWorld merge explorer.** Toggle any subset of the four evolved specialists and watch which of the 41 held-out tasks the selection covers. Built from the `per_task_results` arrays of the five evaluation runs, joined on `task_id` (the run files list the tasks in different orders, so position-based joining would silently misalign them). It surfaces something the paper states only in aggregate. The four specialists solve 24/25/26/27 tasks and their **union is 29**, but the harness recombination actually produced solves **28**: | | tasks | | --- | --- | | best single specialist (D) | 27/41 | | union of all four | 29/41 | | realized merge | 28/41 | | solved by merge, by no specialist | `tw_448247` | | solved by some specialist, not by merge | `tw_449421`, `tw_498533` | | solved by all four | 21/41 | | solved by none | 12/41 | So recombination is not a free set union: it adds a capability no parent had and loses two. These are single attempts and one task is worth 2.4 points, so individual flips sit inside the noise band — the page says so next to the widget. **2. TB2.1 cluster explorer.** Per-cluster base vs. evolved avg@5, sortable by gain, base rate, cluster size, or name, with exact rates on hover. Numbers from `notes/TB21_RESULTS.md` (paired protocol, 88 tasks). Deltas are carried over as reported rather than recomputed, because the source rounds them from unrounded rates. **3. WebArena-Infinity evolution curve.** All 37 evaluated variants with a best-so-far envelope, from `notes/tw_dynamics.json` → `wai_adaptive_scores`. Toggle either series. **4. Headline four-benchmark panels.** Replaces the right half of the teaser, and sits directly under it. Clicking the schematic expands it in place to the paper's full figure, so while expanded the paper's static bars and this interactive version of the same chart are both on screen. From `scripts/gen_summary_figure.py`. The reason to make this one interactive is that the paper's figure has to *state* its truncation in the caption: on a shared axis WebArena-Infinity's +49.5 flattens the two terminal benchmarks into slivers, so each panel is scaled to its own range. Here the reader can switch between per-panel and shared 0–100 axes and see both framings. The best-prior-agent bar can be hidden, since those systems use different models and effort settings and are context rather than a controlled comparison. Per-panel labels carry the exceptions: on SWE-bench Verified the grey bar is the fix-skill reference, not an unevolved Monet, and there is no prior-agent bar. **5. TerminalWorld specialist bars.** From `notes/tw_dynamics.json` → `tw_heldout`. Same axis question in miniature: the six arms span 58.5–68.3, so a 0–100 track renders them near-identical. Defaults to a 55–70 axis with the full axis one click away and the truncation named in both notes. Spec. D and the Claude Code reference are both 27/41, so D's bar landing exactly on the dashed reference line is a built-in check that the axis transform is right. **6. TB2.1 compute.** From `scripts/gen_tb21_compute.py` (medians over clean attempts). Switch between turns and tokens; the bars share one scale across both task groups so the newly-solved vs. already-solved contrast is not rescaled away. **7. WAI invalid-trajectory composition.** From `scripts/gen_wai_invalid_composition.py`. Break the 293 → 17 collapse down by application or by mechanism. Both rows sit on a shared absolute scale by default, which makes the "after" row a near-invisible sliver — that *is* the finding; normalizing each row to its own total then shows what the remainder consists of. `build_data.py` asserts both decompositions still total 293 and 17. **8. WAI audit dumbbells.** From `scripts/gen_wai_audit_by_app.py`. One line per application from raw pass@1 to audited pass@1, so line length reads directly as how much a harness was leaning on trajectories the audit rejects. Sortable by audit loss. `build_data.py` cross-checks the audited DarwinX column against the per-application table rendered elsewhere on the page. Three figures stay static because they are conceptual diagrams with no underlying data: the teaser's selection schematic, the method overview, and the per-generation operators. The archive lineage tree also stays static for a different reason — its node/edge data lives in a `state.db` on the cluster, not on this machine, so there is nothing truthful to make interactive yet. Tables are click-to-sort. Numeric columns open descending, text columns A–Z, and `Overall` rows stay pinned to the bottom. ## Local preview ```bash python3 -m http.server 8000 # open http://localhost:8000 ``` ## Regenerate and test ```bash python3 tools/build_data.py # rebuild assets/data.js from the run artifacts # then, with the server running, open: # http://localhost:8000/tools/interaction_test.html # it drives every widget with real click events and writes PASS/FAIL into the page title ``` Headless run: ```bash python3 -m http.server 8000 & "/Applications/Google Chrome.app/Contents/MacOS/Google Chrome" --headless --disable-gpu \ --virtual-time-budget=6000 --dump-dom http://localhost:8000/tools/interaction_test.html \ | grep -o '