# Helium Browser Environment

This example shows the main pieces of an OpenEnv environment by wrapping a
Chromium browser. A client sends one typed action at a time. The environment
performs it with Helium and Selenium, then returns a typed observation with a
fresh screenshot and the current URL.

The browser is useful here because the boundary is easy to see: OpenEnv owns
the `reset`/`step` protocol, while the environment owns all browser-specific
behavior. The example does not include an agent, so the OpenEnv parts stay in
focus.

![illustrated](https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/helium_demo.png)

## How browser agents work

A natively multimodal browser agent operates in a loop:

1. **Observe:** receive the current page, as a screenshot.
2. **Decide:** choose one small action that moves the task forward.
3. **Act:** click, type, press a key, scroll, or navigate.
4. **Observe again:** inspect the page after the action instead of assuming it
   worked.
5. **Stop:** finish when the task is complete or the episode reaches a limit.

This example uses screenshots as the main observation. The agent sees the same
800×600 browser image that the environment uses for coordinates, then returns
one structured action. A click such as `x=320, y=180` means a click position in that
image; it is not a CSS selector or a reference to the page's HTML.

One action per step matters because web pages change. A click may open a menu,
show a cookie banner, navigate to another page, or fail. Returning a fresh
screenshot and URL after every action lets the caller react to what actually
happened. It also produces a clear trajectory: observation, action, next
observation.

That loop explains the design choices used later in this example:

- **Typed actions** give OpenEnv a small, validated vocabulary for controlling
  the browser.
- **A fixed viewport** keeps screenshot pixels and click coordinates aligned.
- **Real pointer events** let clicks reach visual controls such as overlays and
  canvases without exposing DOM selectors to the agent.
- **An isolated browser** keeps temporary profiles and untrusted pages away
  from the user's personal browser session.
- **Step limits and `done`** give every browser episode a definite boundary.

The agent or policy that chooses actions is intentionally outside this
environment. OpenEnv provides the repeatable interaction contract; the caller
decides how to turn an observation into the next `BrowserAction`.

## Tasks

Websites always keep on changing so using a regular dataset doesn't work. You can try the samples [in this dataset](https://huggingface.co/datasets/merve/browser-tasks) as they are created late 2026. They are divided into three:
- **Information:** getting an information from a website by browsing it.
- **Navigation:** navigating a specific part of the website. 
- **Interaction:** clicking, filtering etc.

## The OpenEnv pieces

An OpenEnv environment has four small parts in this example:

| OpenEnv concept | Implementation | Responsibility |
| --- | --- | --- |
| Data contract | `BrowserAction`, `BrowserObservation`, `BrowserState` | Defines the validated objects that cross the client/server boundary. |
| Environment | `BrowserEnvironment` | Implements the episode lifecycle: `reset`, `step`, `state`, and `close`. |
| Server | `create_app(...)` in `server/app.py` | Turns the environment class and data models into an HTTP/WebSocket service. |
| Client | `BrowserClient` | Serializes actions and reconstructs typed observations and step results. |

These parts separate domain logic from transport. `BrowserEnvironment` knows
how to drive a browser, but it does not define HTTP routes. `BrowserClient`
knows how to exchange OpenEnv messages, but it does not import Helium or
Selenium.

### 1. Define the data contract

`models.py` subclasses OpenEnv's three base models with the actions an agent can take: "click", "type", "key", "scroll", "back", "wait", "finish". x and y are clicking coordinates, dy is change in vertical direction (scroll dy pixels), text is text to be typed on e.g. search bars.

```python
class BrowserAction(Action):
    op: Literal["click", "type", "key", "scroll", "back", "wait", "finish"]
    x: int | None = Field(default=None, ge=0, lt=800)
    y: int | None = Field(default=None, ge=0, lt=600)
    text: str = Field(default="", max_length=2000)
    dy: int = Field(default=0, ge=-1200, le=1200)

class BrowserObservation(Observation):
    screenshot: str
    url: str
    error: str = ""

class BrowserState(State):
    pass
```

The action describes what may enter the environment. The observation describes
what comes back. The state is separate: OpenEnv uses it to report episode
metadata such as the episode ID and step count without adding those fields to
every browser observation.

Because these are Pydantic models, validation happens at the boundary. For
example, a click outside the 800×600 viewport is rejected before browser logic
runs, and the model validator makes sure a `click` has both coordinates.

### 2. Implement the environment lifecycle

`BrowserEnvironment` subclasses OpenEnv's `Environment`. Its methods have the
same roles they would have in a game, simulator, or other environment:

- `reset(...)` starts a new episode and returns its first observation. Here it
  creates a temporary browser profile, opens `start_url`, and takes a
  screenshot.
- `step(action)` applies exactly one validated action and returns the next
  observation.
- `state` exposes the current `BrowserState`, including OpenEnv's episode ID
  and step count.
- `close()` releases the resources owned by the episode.

This is the Gym-like core of OpenEnv. Browser setup and Helium calls are
ordinary implementation details behind that interface.

### 3. Turn it into a server

`server/app.py` passes the environment class and its wire types to `create_app`:

```python
app = create_app(
    BrowserEnvironment,
    BrowserAction,
    BrowserObservation,
    env_name="helium_browser_env",
    max_concurrent_envs=1,
)
```

`create_app` supplies the FastAPI application and the standard OpenEnv
endpoints, including health, reset, step, state, and WebSocket access. The
environment code only implements the lifecycle; it does not duplicate those
routes.

The environment class is passed instead of a pre-built instance so OpenEnv can
own its lifecycle. Concurrency is limited to one because Helium stores the
active Selenium driver globally and Selenium drivers are not thread-safe.

### 4. Add the typed client

`BrowserClient` subclasses OpenEnv's generic `EnvClient` with the action,
observation, and state types:

```python
class BrowserClient(EnvClient[BrowserAction, BrowserObservation, BrowserState]):
    def _step_payload(self, action):
        return action.model_dump()

    def _parse_result(self, payload):
        ...

    def _parse_state(self, payload):
        return BrowserState.model_validate(payload)
```

The base client handles the connection and the common `reset`, `step`, and
state operations. This subclass only teaches it how this environment's models
map to JSON. Callers therefore work with `BrowserAction` and
`BrowserObservation`, not untyped dictionaries.

## How it fits together

```text
BrowserClient -> OpenEnv server -> BrowserEnvironment -> Helium/Selenium -> Chromium
BrowserClient <- screenshot, URL, error, done <- BrowserEnvironment
```

For a step, `BrowserClient` serializes a `BrowserAction`. OpenEnv validates it,
calls `BrowserEnvironment.step()`, and serializes the returned
`BrowserObservation`. The client then rebuilds a typed `StepResult` containing
the observation, reward, and `done` flag.

## Files

| File | Purpose |
| --- | --- |
| `models.py` | Defines the OpenEnv action, observation, and state contract. |
| `client.py` | Adapts the OpenEnv client to the environment's typed wire format. |
| `hf_sandbox.py` | Starts a packaged environment image with OpenEnv's HF Sandbox provider. |
| `server/helium_browser_environment.py` | Implements OpenEnv's lifecycle while owning Chromium. |
| `server/desktop.py` | Starts a private X display and sends real pointer clicks with `xdotool`. |
| `server/app.py` | Gives the environment and its models to OpenEnv's `create_app`. |
| `server/Dockerfile` | Builds the runnable image with Chromium and its system tools. |
| `pyproject.toml` | Declares Python dependencies, package layout, and the `server` command. |
| `uv.lock` | Pins the resolved Python dependency versions for reproducible builds. |
| `openenv.yaml` | Describes the environment to the OpenEnv CLI. |

## Why the package has both a Dockerfile and an HF Sandbox runner

The Dockerfile and HF Sandbox solve different parts of deployment.

`server/Dockerfile` is the image recipe. It installs Chromium, ChromeDriver,
Xvfb, and `xdotool`, then installs this Python package. The `server` command
declared in `pyproject.toml` starts `server/app.py`. This makes the image
self-contained: a runtime does not have to install browser software or copy
source files when an episode begins.

`hf_sandbox.py` is a client-side launcher. It asks `HFSandboxProvider` to start
that already-built image in an isolated HF Sandbox, waits for the OpenEnv
health endpoint, and connects `BrowserClient`. The Sandbox provides isolation
and a secure proxy; the image provides the software that runs inside it.

This separation is why `app.py` and `client.py` are both present:

- `server/app.py` runs inside the image and exposes the OpenEnv protocol;
- `client.py` runs on the caller side and converts typed Python objects to and
  from that protocol.

The client is not browser automation code. Helium, Selenium, Chromium, and the
virtual display stay behind the server boundary.

## Package and publish the environment

From the OpenEnv repository root, enter the environment directory:

```sh
cd envs/helium_browser_env
```

The usual OpenEnv packaging commands are:

```sh
openenv build
openenv push --repo-id <namespace>/helium-browser-env
```

`openenv build` uses `server/Dockerfile` with the environment directory as its
build context. `openenv push` publishes the same package as a Hugging Face
Space. Once the Space image has built, HF Sandboxes can refer to it as
`hf.co/spaces/<namespace>/helium-browser-env`.

## Run it in an HF Sandbox

`hf_sandbox.py` follows OpenEnv's provider pattern: start a packaged environment
image in an HF Sandbox, wait for its OpenEnv server, then connect the same typed
client used with any other OpenEnv deployment.

```python
import os

from openenv.core.containers.runtime.hf_sandbox_provider import HFSandboxProvider

from helium_browser_env.client import BrowserClient
from helium_browser_env.models import BrowserAction

with HFSandboxProvider(
    image=os.environ["HELIUM_BROWSER_IMAGE"],
    flavor="cpu-basic",
    env_vars={"BROWSER_SANDBOX": "hf"},
) as provider:
    base_url = provider.start_container()
    provider.wait_for_ready(base_url, timeout_s=300.0)

    with BrowserClient(base_url=base_url).sync() as browser:
        result = browser.reset(start_url="https://example.com")
        result = browser.step(BrowserAction(op="scroll", dy=500))
        print(result.observation.url)
```

Set `HELIUM_BROWSER_IMAGE` to the image published in the previous section. A
Hugging Face Space image uses the form `hf.co/spaces/<namespace>/<space>`.

```sh
export HELIUM_BROWSER_IMAGE=hf.co/spaces/<namespace>/<space>
uv run helium-browser-hf
```

The runner also forwards the existing browser settings for the step limit and
action waits. `BROWSER_SANDBOX=hf` satisfies the environment's safety guard.
The provider handles the authenticated proxy and destroys the Sandbox when its
context exits. This is the same lifecycle shown in OpenEnv's
[HF Sandbox example](https://github.com/huggingface/OpenEnv/blob/main/examples/hf_sandbox_coding_env.py).

## The OpenEnv loop

Once the server is running, reset the browser to a URL and send one action at a
time. The local URL below is useful when explaining the OpenEnv protocol on its
own; the HF Sandbox launcher supplies a proxied URL instead.

```python
from helium_browser_env.client import BrowserClient
from helium_browser_env.models import BrowserAction

with BrowserClient(base_url="http://localhost:8000").sync() as browser:
    result = browser.reset(start_url="https://example.com")
    observation = result.observation

    result = browser.step(BrowserAction(op="scroll", dy=500))
    observation = result.observation

    print(observation.url)
    print(observation.error)
    print(result.done, result.reward)
    print(observation.screenshot)  # Base64-encoded PNG
```

`reset` must receive an HTTP or HTTPS URL without embedded credentials. It
starts a fresh browser profile, opens the page, and returns the first
observation. `step` performs one action and returns an OpenEnv `StepResult`.
The result groups the next observation with the reward and episode termination
flag, which gives callers the same loop shape across different environments.

## Actions

| Operation | Fields | What it does |
| --- | --- | --- |
| `click` | `x`, `y` | Clicks one point in the 800×600 browser viewport. |
| `type` | `text` | Types into the currently focused element. |
| `key` | `text` | Presses a supported key such as `ENTER`, `TAB`, or `CTRL+A`. |
| `scroll` | `dy` | Scrolls down for positive values and up for negative values. |
| `back` | none | Goes back in browser history. |
| `wait` | none | Waits for the page without performing another action. |
| `finish` | optional `text` | Ends the episode. |

Click coordinates are browser pixels: `x` is from 0 to 799 and `y` is from 0
to 599. Scroll values are limited to the range -1200 to 1200.

Supported keys are `ENTER`, `TAB`, `ESCAPE`, `BACKSPACE`, `CTRL+A`, `UP`,
`DOWN`, `LEFT`, and `RIGHT`.

## Why the viewport is fixed

Chromium runs on a private Xvfb display with an 800×600 viewport. The
environment waits until the browser reports that exact geometry before it
opens the requested page. This keeps screenshot coordinates aligned with
click coordinates.

Clicks go through `xdotool` instead of a DOM selector. They are real pointer
events, so they interact with overlays, canvas elements, and other visual
controls in the same coordinate system as the screenshot. Helium handles text,
keys, scrolling, navigation, and access to the Selenium driver.

Helium stores the active driver globally, and Selenium drivers are not
thread-safe. The server therefore allows one environment session at a time and
the environment asks OpenEnv to keep its work on one thread.

## Observations and episode endings

Every observation contains:

- `screenshot`: the current browser window as a base64-encoded PNG;
- `url`: the current page URL;
- `error`: an empty string or a short error code;
- `done`: whether the episode has ended.

The OpenEnv step result carries a reward of `0` because this example does not
score browser actions.

The episode ends when it receives `finish`, reaches the configured step limit,
encounters a renderer timeout, or detects a supported bot-verification page.
An action error is reported by exception class name so the caller can inspect
the next screenshot and decide what to do.

## Settings

| Environment variable | Default | Accepted values |
| --- | --- | --- |
| `BROWSER_MAX_STEPS` | `20` | An integer from 1 to 200. |
| `BROWSER_ACTION_WAIT_SECONDS` | `1.5` | Seconds from 0 to 10 after ordinary actions. |
| `BROWSER_EXPLICIT_WAIT_SECONDS` | `3` | Seconds from 0 to 10 for a `wait` action. |

## Boundaries

The environment is designed to run inside its browser sandbox, not against a
personal desktop browser. Chromium uses a temporary profile for each episode.
Downloads and password storage are disabled, and a reset rejects URLs that
contain a username or password.

Web pages are untrusted input. Do not enter credentials or sensitive data, and
expect some sites to show automation challenges. This example provides the
browser interaction loop; it does not provide tasks, scoring, or a reward
function.

