| --- |
| language: |
| - en |
| tags: |
| - browser-agent |
| - cpu |
| - experimental |
| - custom-model |
| library_name: baim |
| --- |
| |
| # Devils Agent / BAIM — experimental research checkpoints |
|
|
| **Research prototype, not a production-ready general browser agent.** Four custom |
| checkpoints are stored under `models/`. They require the accompanying Python code; |
| this repository is not a standard Transformers `AutoModel` or hosted-inference |
| package. No production champion has been selected. |
|
|
| The checkpoints were trained on generated single-step click/type/select fixtures. |
| The mean encoder matched only 1 of 46 scorable action/target pairs in a small |
| Mind2Web diagnostic. Target Linux VPS validation is unfinished. Do not interpret |
| the high synthetic scores as real-world browser reliability. |
|
|
| No code/model license grant has been selected for this release. Third-party |
| pretrained weights and raw Mind2Web data are not bundled. See `docs/RESEARCH.md` |
| and `docs/PRETRAINED.md` for source attribution and external-model provenance. |
|
|
| CPU-first browser action research in progress. **Small policies have been trained, |
| but there is no validated production checkpoint.** The project includes a browser |
| runtime, synthetic training pipeline, learned action/pointer policies and measured |
| evaluations. Real-data transfer is poor; this is not yet a general browser agent. |
|
|
| ## Run |
|
|
| Python 3.12 or newer: |
|
|
| ```sh |
| python -m venv .venv |
| # Linux: source .venv/bin/activate |
| # Windows PowerShell: .venv\Scripts\Activate.ps1 |
| python -m pip install -e ".[browser]" |
| python -m playwright install chromium |
| python -m unittest discover -s tests -v |
| python -m baim.bench_runtime --output reports/runtime-baseline.json |
| ``` |
|
|
| The current development host uses `py -3.12` in place of `python` without a virtual |
| environment for the runtime-only tests. Training uses a virtual environment at |
| `../../work/baim-venv`. Tests launch isolated headless Chromium contexts and local fixtures. |
|
|
| ## Train and compare |
|
|
| ```sh |
| python -m pip install torch --index-url https://download.pytorch.org/whl/cpu |
| python -m pip install -e ".[training,browser]" numpy |
| python -m baim.synthetic |
| python -m baim.train --epochs 16 --output models/v000-mean |
| python -m baim.train --epochs 16 --no-lexical --output models/v001-no-lexical |
| python -m baim.train --epochs 16 --encoder gru --output models/v002-gru |
| python -m baim.train --epochs 16 --encoder transformer --output models/v003-transformer |
| python -m baim.experiment_suite |
| python -m baim.evaluate_policy --limit 120 |
| python -m baim.evaluate_policy --data datasets/synthetic-v1/novel_wording.jsonl --output reports/policy-browser-novel-v000.json |
| python -m baim.evaluate_policy --baseline --data datasets/synthetic-v1/novel_wording.jsonl --output reports/baseline-browser-novel.json |
| ``` |
|
|
| Mean/GRU/Transformer policies use Hugging Face `PyTorchModelHubMixin` checkpoint |
| serialization. Nothing is uploaded automatically. FP32 checkpoints use safetensors. |
| The quantization harness quantizes Linear layers and tests restricted state-dict |
| reloading; embeddings and recurrent/attention encoders remain FP32. |
|
|
| See [DATASET.md](docs/DATASET.md) for split design and known shortcuts, |
| [ACCEPTANCE.md](docs/ACCEPTANCE.md) for the full remaining scope, and |
| `requirements-observed.txt` for package versions measured on this Windows host. |
|
|
| ## Current boundaries |
|
|
| Model output uses a compact action opcode plus a JSON array (for example |
| `C["e17"]` and `T["e4","hello"]`). It cannot supply selectors or executable JS. |
| The host supplies the task authority, permission callback and completion verifier. |
| Task tickets bind session, task epoch, action sequence and observation revision. |
| References resolve to retained DOM nodes, rechecked before interaction. The adapter |
| includes frames and open shadow roots. Browser actionability checks still apply. |
|
|
| Recorded telemetry contains metadata and keyed hashes; it excludes goal text, |
| typed values, raw observations and extracted content. This is **not yet sufficient |
| training data**. The key is a separate local file; protect both files and set an |
| appropriate retention policy. Sensitive-target detection is incomplete and must |
| not be mistaken for comprehensive PII detection. |
|
|
| This is not a browser security sandbox. The permission callback must enforce the |
| deployment's trusted action policy. Network isolation, redirect restrictions, |
| download policy and rich redacted trajectories remain to be implemented. A page |
| can change between validation and interaction; hostile timing attacks are not |
| solved by node handles. Accessible names are an approximation, not full ARIA |
| accessible-name computation. Closed shadow roots and canvas require fallback. |
|
|
| The learned baseline covers CLICK, TYPE and SELECT, with literal copying from |
| one quoted user-goal value. It has no general planner, history model or reliable |
| unsupported-task detector. Validation temperature scaling fails under distribution |
| shift; confidence is not a security boundary. Keep it on isolated research fixtures. |
|
|
| FINISH requires a host verifier bound to the current task. The local tests supply |
| fixture-specific verifiers; arbitrary user-goal completion is still unresolved. |
|
|
| See [docs/STATUS.md](docs/STATUS.md) and the full [requirements](docs/requirements.txt). |
|
|