Configuration Parsing Warning:In adapter_config.json: "peft.task_type" must be a string

Reflex Instinct 0.6B

Reflex Instinct is the fast decision model behind Reflex, a local browser agent by Atlas AI. Given a task and a web page's accessibility tree, it picks the next action and the element to act on in a single forward pass, and returns calibrated probabilities, never generated text.

  • About 0.2 s per decision on a consumer GPU (191 ms p50 on an RTX 5070, real pages of about 4.5k tokens)
  • Calibrated: expected calibration error (ECE) of 0.025 on 2,000 held-out steps, so "0.9" means about 90%
  • Runs locally: 0.6B parameters, small enough for any recent consumer GPU
  • Decides, never writes: there is no language-model head, and every answer is a probability over the options you supply

In Reflex, Instinct handles every step it is confident about. When its confidence falls below a threshold, the step goes to Reflex Reason, a 2B vision-language model that looks at a screenshot and reasons the step through.

Code: github.com/leemadov/reflex: the Reflex browser agent that runs this model (Chrome extension and desktop app), plus the training, evaluation and JevBench scripts.

How it works

  • Backbone: Qwen/Qwen3-0.6B with a LoRA adapter (rank 64, alpha 128, all attention and MLP projections). The base model's LM head is not used.
  • Pointer head: two 256-d projections (q, k) plus a calibration temperature T, about 0.5M parameters (head.safetensors).
  • Answer options are positions in the input. A text option is the last token of its line in the question block. An element option is the last token of that element's line inside the page, so elements are never copied. logit = <W_q h_question, W_k h_option> / sqrt(256), followed by a softmax over that question's options.
  • Three question types:
    • choice: one of the given options, or "elements" for every [id] element on the page
    • noul: a yes/no probability
    • score: ordered levels, returning the expected level
  • Batching: any number of questions are answered in the same single pass.

Operations seen in training: CLICK, TYPE_TEXT, HOVER, PRESS_KEY, SCROLL_DOWN, SCROLL_UP, GOTO_URL, GO_BACK, GO_FORWARD, SWITCH_TAB, DONE, BLOCKED. Training randomized option names, descriptions, order and subsets, so you can pass your own operation list.

Results

Measured on 2,000 held-out test steps from NNetNav, disjoint from the steps used to pick the checkpoint:

Metric Reflex Instinct Reference
Operation, top-1 / top-5 62.1% / 95.2% 56.0% for always predicting CLICK; 61% for TF-IDF + logistic regression
Target element, top-1 / top-5 41.6% / 70.2% about 0.3% for a random guess (about 300 elements per page)
Step (operation and element both right) 31.0% WebArena sites 32.0%, live web 30.0%
Calibration error (ECE, 15 bins) 0.025
Decisions with p ≥ 0.9 operation: 90.4% right; element: 95.1% right about 6–7% of steps
Latency 162 ms p50 / 452 ms p90 in-process, RTX 5070

How to read this: the labels come from an LLM explorer, so a step often has several reasonable next actions and the logged one is only one of them. Top-1 accuracy against these labels understates how useful the model is; the calibrated probabilities are the point. Act automatically when confidence is high, and send low-confidence steps to a bigger model or to the user.

Other checks, on held-out pages:

  • 1,200 answers with 2 to 600+ options, including option names never seen in training, gave 0 answers outside the supplied options.
  • Asking a question alone or in a different order moves its probabilities by a median of 0.001.
  • The same 3 decisions took 0.19 s each here, against 14 s and 380 generated tokens for the base Qwen3-0.6B reasoning them out in text.

Usage

Run it

serve.py in this repo serves the model on Jev's /v1/systemone wire format (TypeSafe-compatible), on CPU or GPU:

pip install torch transformers peft safetensors fastapi uvicorn huggingface_hub
python serve.py --port 8765

It downloads the base model and this adapter, then answers POST http://127.0.0.1:8765/v1/systemone.

Request and response

A request has a state (the task, the action history and the accessibility tree) and named questions:

{
  "state": {
    "task": "Find reviews for a product in the Electronics category.",
    "url": "http://shop.local/",
    "history": [],
    "page": "RootWebArea 'One Stop Market'\n\t[227] link 'My Account'\n\t[815] menuitem 'Electronics'\n\t[272] combobox 'Search'"
  },
  "questions": {
    "operation": {"type": "choice", "instructions": "What is the next operation?",
                  "criteria": {"CLICK": "click an element", "TYPE_TEXT": "type text into a field",
                               "SCROLL": "scroll the page", "DONE": "the task is complete"}},
    "target": {"type": "choice", "instructions": "Which element should the next operation act on?",
               "criteria": "elements"},
    "done": {"type": "noul", "instructions": "Is the task already accomplished?"}
  }
}

Response:

{"operation": {"type": "choice", "choice": "CLICK", "confidence": 0.369,
               "probabilities": {"CLICK": 0.5181, "TYPE_TEXT": 0.4322, "SCROLL": 0.0413, "DONE": 0.0085}},
 "target": {"type": "choice", "choice": "272", "confidence": 0.455,
            "probabilities": {"227": 0.0049, "815": 0.2563, "272": 0.7387}},
 "done": {"type": "noul", "noul": 0.0062}}

On this three-element toy page the model hesitates between searching (272) and opening the category menu (815), and that split is visible in the probabilities instead of being hidden. confidence is 1 - normalized entropy.

Input format

The page should look like a WebArena/BrowserGym accessibility tree: one element per line as [id] role 'name' properties, tab-indented by depth. That is what the model was trained on. The context limit is 16k tokens; longer pages are cut off at the end, by whole lines.

Files

File What
adapter_config.json, adapter_model.safetensors LoRA adapter for Qwen/Qwen3-0.6B (load with PEFT, then merge)
head.safetensors Pointer head (q.weight, q.bias, k.weight, k.bias) and calibration temperature T
serve.py The inference code and server: builds the input, runs the one pass, reads out the answers

The model is not a standard transformers architecture: the backbone is the bare Qwen3Model, and the scores come from the pointer head, so it needs the inference code in serve.py.

JevBench (self-measured, public items only)

On JevBench's 231 public items, through JevBench's own typesafe adapter: easy 91.7%, standard 66.7%, hard 28.8% (below that tier's 33.6% chance). It was trained only on web navigation, so general reasoning questions are outside what it learned. These are our own measurements, not JevBench's.

Training

  • Data: stanfordnlp/nnetnav-wa (44k steps on WebArena's self-hosted sites) and stanfordnlp/nnetnav-live (46k steps on live websites), both Apache-2.0, from NNetNav. Only the train splits were used.
  • Preprocessing: per-link URLs, icon-font glyphs and repeated StaticText were stripped, taking a live-web page from a median of 11k to 4.5k tokens.
  • Run: 2,000 optimizer steps of 8 sequences each (16k steps, about 18% of one epoch), 4 h 20 min on one RTX 5070, 4 GB peak memory. Calibrated afterwards with temperature scaling. Accuracy was still rising when training stopped.

Limitations

  • It is specialized for web pages as accessibility trees. It does not see pixels; screenshots are Reflex Reason's job.
  • The training labels come from an LLM explorer, so the model learned that explorer's habits along with the task.
  • It picks among similar-looking elements (for example 19 identical "Add to cart" buttons) poorly on its own. Reflex resolves these by matching the task against each button's surrounding card.
  • It is English only.
  • The context limit is 16k tokens.
  • Act on its decisions automatically only above a confidence threshold, and never let an agent type passwords or payment details on its own.

License

Apache 2.0, the same as the base model (Qwen3-0.6B) and the training data (NNetNav).

Built by Atlas AI.

Downloads last month
30
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Atlas-AI-research/reflex-instinct-0.6b

Finetuned
Qwen/Qwen3-0.6B
Adapter
(639)
this model

Datasets used to train Atlas-AI-research/reflex-instinct-0.6b

Paper for Atlas-AI-research/reflex-instinct-0.6b