---
license: mit
library_name: transformers
pipeline_tag: token-classification
language:
- multilingual
tags:
- autofill
- password-manager
- form-classification
- field-classification
- browser-extension
- bert
- onnx
---
# Zap ⚡ — a tiny web-form classifier for autofill
**Zap** looks at a web form and tells a password manager what every input wants and what the form is for.
It runs locally, has **20.7M parameters**, and ships **bf16 weights (41 MB)**, plus ONNX files (fp32 83 MB, int8 21 MB).
* **Field types (26):** `username`, `email`, `current-password`, `new-password`, `one-time-code`, `tel`, names, address parts, payment-card parts, `bday`, `search`, `captcha`, `other`
* **Form purposes (11):** `login`, `signup`, `password-change`, `password-recovery`, `otp`, `payment`, `address`, `newsletter`, `contact`, `search`, `other`
So your autofill can answer questions like *"is this a login form?"*, *"which box gets the saved password, and which one gets a newly generated password?"*, *"is this the email field, or is it a honeypot?"*
**Held-out websites:** 6,838 real forms and 19,811 fields from Common Crawl pages on hosts never seen in training:
| | **Zap** | keyword heuristic¹ |
|---|---|---|
| field accuracy (26 classes) | **97.2%** | 82.7% |
| field macro-F1 | **0.945** | 0.723 |
| form-purpose accuracy (11 classes) | **95.7%** | 71.6% |
| password-field detection F1 | **0.996** | 0.968 |
| current vs new password (on password fields) | **99.2%** | 85.5% |
| login-identifier (`username`/`email`) F1 | **0.994** | 0.945 |
| `autocomplete` removed from the input: recovers the developer's own token² | **97.8%** | 78.7% |
¹ A baseline of the kind password managers ship: it trusts a valid `autocomplete`, then the input `type`, then multilingual keyword regexes over name/id/placeholder/label/nearby text.
² On 1,622 test fields where the site's `autocomplete` token agrees with the reference label, the token is deleted from the model input. The score is how often Zap still predicts it. This check doesn't depend on the LLM teacher.
**Live websites:** 147 real login, sign-up and password-reset pages rendered in Chromium. Every visible field was labeled blind from a marked screenshot, and Zap's predictions were compared with the same keyword heuristic ([details](#live-site-benchmark)):

## Quick start (Python)
```python
# pip install transformers torch lxml selectolax huggingface_hub
import sys
from huggingface_hub import snapshot_download
sys.path.insert(0, snapshot_download("ProCreations/Zap"))
from zap_infer import Zap
zap = Zap.from_pretrained("ProCreations/Zap") # CPU, fp32 compute from the bf16 weights
for form in zap.classify_html(html, url="https://northwind.example/login"):
print(form["form"], round(form["form_score"], 3))
for f in form["fields"]:
print(" ", f["type"], f["name"], "->", f["label"], round(f["score"], 3))
```
Output for a page that has a sign-in form and a separate "New here? Create account" block:
```
login 0.966
input text login -> username 0.96
input password pw -> current-password 0.962
signup 0.961
input text email -> email 0.957
input password pass 1 -> new-password 0.956
input password pass 2 -> new-password 0.955
```
Neither block sets a useful `autocomplete` attribute: the login box has `autocomplete="off"`, and the signup inputs are typeless ``s with placeholders. Zap tells them apart from context.
## In a browser extension (JavaScript)
Zap doesn't read raw HTML. It reads a compact text rendering of each form: page title, URL words, form attributes, heading, buttons and links, and for every field its type, name, id, autocomplete, placeholder, label, nearby text and so on. `zap-extract.js` builds that rendering from the **live DOM**, so it sees exactly what the model was trained on. `zap-infer.js` handles windowing and decoding with transformers.js.
```js
import * as transformers from "@huggingface/transformers";
// load zap-extract.js and zap-infer.js first (content scripts), or require() them in Node
const zap = await ZapModel.load(transformers, "ProCreations/Zap"); // onnx/model.onnx; pass { dtype: "q8" } for the int8 file
const results = await zap.classify(Zap.zapExtract(document, location.href));
for (const r of results) {
console.log(r.form, r.formScore);
for (const f of r.fields) console.log(f.element, f.label, f.score); // f.element is the /