OpenEnv documentation

Helium Browser Environment

You are viewing main version, which requires installation from source. If you'd like regular pip install, checkout the latest stable version (v0.7.0).
Hugging Face's logo
Join the Hugging Face community

and get access to the augmented documentation experience

to get started

Helium Browser Environment

This example shows the main pieces of an OpenEnv environment by wrapping a Chromium browser. A client sends one typed action at a time. The environment performs it with Helium and Selenium, then returns a typed observation with a fresh screenshot and the current URL.

The browser is useful here because the boundary is easy to see: OpenEnv owns the reset/step protocol, while the environment owns all browser-specific behavior. The example does not include an agent, so the OpenEnv parts stay in focus.

illustrated

How browser agents work

A natively multimodal browser agent operates in a loop:

  1. Observe: receive the current page, as a screenshot.
  2. Decide: choose one small action that moves the task forward.
  3. Act: click, type, press a key, scroll, or navigate.
  4. Observe again: inspect the page after the action instead of assuming it worked.
  5. Stop: finish when the task is complete or the episode reaches a limit.

This example uses screenshots as the main observation. The agent sees the same 800×600 browser image that the environment uses for coordinates, then returns one structured action. A click such as x=320, y=180 means a click position in that image; it is not a CSS selector or a reference to the page’s HTML.

One action per step matters because web pages change. A click may open a menu, show a cookie banner, navigate to another page, or fail. Returning a fresh screenshot and URL after every action lets the caller react to what actually happened. It also produces a clear trajectory: observation, action, next observation.

That loop explains the design choices used later in this example:

  • Typed actions give OpenEnv a small, validated vocabulary for controlling the browser.
  • A fixed viewport keeps screenshot pixels and click coordinates aligned.
  • Real pointer events let clicks reach visual controls such as overlays and canvases without exposing DOM selectors to the agent.
  • An isolated browser keeps temporary profiles and untrusted pages away from the user’s personal browser session.
  • Step limits and done give every browser episode a definite boundary.

The agent or policy that chooses actions is intentionally outside this environment. OpenEnv provides the repeatable interaction contract; the caller decides how to turn an observation into the next BrowserAction.

Tasks

Websites always keep on changing so using a regular dataset doesn’t work. You can try the samples in this dataset as they are created late 2026. They are divided into three:

  • Information: getting an information from a website by browsing it.
  • Navigation: navigating a specific part of the website.
  • Interaction: clicking, filtering etc.

The OpenEnv pieces

An OpenEnv environment has four small parts in this example:

OpenEnv conceptImplementationResponsibility
Data contractBrowserAction, BrowserObservation, BrowserStateDefines the validated objects that cross the client/server boundary.
EnvironmentBrowserEnvironmentImplements the episode lifecycle: reset, step, state, and close.
Servercreate_app(...) in server/app.pyTurns the environment class and data models into an HTTP/WebSocket service.
ClientBrowserClientSerializes actions and reconstructs typed observations and step results.

These parts separate domain logic from transport. BrowserEnvironment knows how to drive a browser, but it does not define HTTP routes. BrowserClient knows how to exchange OpenEnv messages, but it does not import Helium or Selenium.

1. Define the data contract

models.py subclasses OpenEnv’s three base models with the actions an agent can take: “click”, “type”, “key”, “scroll”, “back”, “wait”, “finish”. x and y are clicking coordinates, dy is change in vertical direction (scroll dy pixels), text is text to be typed on e.g. search bars.

class BrowserAction(Action):
    op: Literal["click", "type", "key", "scroll", "back", "wait", "finish"]
    x: int | None = Field(default=None, ge=0, lt=800)
    y: int | None = Field(default=None, ge=0, lt=600)
    text: str = Field(default="", max_length=2000)
    dy: int = Field(default=0, ge=-1200, le=1200)


class BrowserObservation(Observation):
    screenshot: str
    url: str
    error: str = ""


class BrowserState(State):
    pass

The action describes what may enter the environment. The observation describes what comes back. The state is separate: OpenEnv uses it to report episode metadata such as the episode ID and step count without adding those fields to every browser observation.

Because these are Pydantic models, validation happens at the boundary. For example, a click outside the 800×600 viewport is rejected before browser logic runs, and the model validator makes sure a click has both coordinates.

2. Implement the environment lifecycle

BrowserEnvironment subclasses OpenEnv’s Environment. Its methods have the same roles they would have in a game, simulator, or other environment:

  • reset(...) starts a new episode and returns its first observation. Here it creates a temporary browser profile, opens start_url, and takes a screenshot.
  • step(action) applies exactly one validated action and returns the next observation.
  • state exposes the current BrowserState, including OpenEnv’s episode ID and step count.
  • close() releases the resources owned by the episode.

This is the Gym-like core of OpenEnv. Browser setup and Helium calls are ordinary implementation details behind that interface.

3. Turn it into a server

server/app.py passes the environment class and its wire types to create_app:

app = create_app(
    BrowserEnvironment,
    BrowserAction,
    BrowserObservation,
    env_name="helium_browser_env",
    max_concurrent_envs=1,
)

create_app supplies the FastAPI application and the standard OpenEnv endpoints, including health, reset, step, state, and WebSocket access. The environment code only implements the lifecycle; it does not duplicate those routes.

The environment class is passed instead of a pre-built instance so OpenEnv can own its lifecycle. Concurrency is limited to one because Helium stores the active Selenium driver globally and Selenium drivers are not thread-safe.

4. Add the typed client

BrowserClient subclasses OpenEnv’s generic EnvClient with the action, observation, and state types:

class BrowserClient(EnvClient[BrowserAction, BrowserObservation, BrowserState]):
    def _step_payload(self, action):
        return action.model_dump()

    def _parse_result(self, payload):
        ...

    def _parse_state(self, payload):
        return BrowserState.model_validate(payload)

The base client handles the connection and the common reset, step, and state operations. This subclass only teaches it how this environment’s models map to JSON. Callers therefore work with BrowserAction and BrowserObservation, not untyped dictionaries.

How it fits together

BrowserClient -> OpenEnv server -> BrowserEnvironment -> Helium/Selenium -> Chromium
BrowserClient <- screenshot, URL, error, done <- BrowserEnvironment

For a step, BrowserClient serializes a BrowserAction. OpenEnv validates it, calls BrowserEnvironment.step(), and serializes the returned BrowserObservation. The client then rebuilds a typed StepResult containing the observation, reward, and done flag.

Files

FilePurpose
models.pyDefines the OpenEnv action, observation, and state contract.
client.pyAdapts the OpenEnv client to the environment’s typed wire format.
hf_sandbox.pyStarts a packaged environment image with OpenEnv’s HF Sandbox provider.
server/helium_browser_environment.pyImplements OpenEnv’s lifecycle while owning Chromium.
server/desktop.pyStarts a private X display and sends real pointer clicks with xdotool.
server/app.pyGives the environment and its models to OpenEnv’s create_app.
server/DockerfileBuilds the runnable image with Chromium and its system tools.
pyproject.tomlDeclares Python dependencies, package layout, and the server command.
uv.lockPins the resolved Python dependency versions for reproducible builds.
openenv.yamlDescribes the environment to the OpenEnv CLI.

Why the package has both a Dockerfile and an HF Sandbox runner

The Dockerfile and HF Sandbox solve different parts of deployment.

server/Dockerfile is the image recipe. It installs Chromium, ChromeDriver, Xvfb, and xdotool, then installs this Python package. The server command declared in pyproject.toml starts server/app.py. This makes the image self-contained: a runtime does not have to install browser software or copy source files when an episode begins.

hf_sandbox.py is a client-side launcher. It asks HFSandboxProvider to start that already-built image in an isolated HF Sandbox, waits for the OpenEnv health endpoint, and connects BrowserClient. The Sandbox provides isolation and a secure proxy; the image provides the software that runs inside it.

This separation is why app.py and client.py are both present:

  • server/app.py runs inside the image and exposes the OpenEnv protocol;
  • client.py runs on the caller side and converts typed Python objects to and from that protocol.

The client is not browser automation code. Helium, Selenium, Chromium, and the virtual display stay behind the server boundary.

Package and publish the environment

From the OpenEnv repository root, enter the environment directory:

cd envs/helium_browser_env

The usual OpenEnv packaging commands are:

openenv build
openenv push --repo-id <namespace>/helium-browser-env

openenv build uses server/Dockerfile with the environment directory as its build context. openenv push publishes the same package as a Hugging Face Space. Once the Space image has built, HF Sandboxes can refer to it as hf.co/spaces/<namespace>/helium-browser-env.

Run it in an HF Sandbox

hf_sandbox.py follows OpenEnv’s provider pattern: start a packaged environment image in an HF Sandbox, wait for its OpenEnv server, then connect the same typed client used with any other OpenEnv deployment.

import os

from openenv.core.containers.runtime.hf_sandbox_provider import HFSandboxProvider

from helium_browser_env.client import BrowserClient
from helium_browser_env.models import BrowserAction

with HFSandboxProvider(
    image=os.environ["HELIUM_BROWSER_IMAGE"],
    flavor="cpu-basic",
    env_vars={"BROWSER_SANDBOX": "hf"},
) as provider:
    base_url = provider.start_container()
    provider.wait_for_ready(base_url, timeout_s=300.0)

    with BrowserClient(base_url=base_url).sync() as browser:
        result = browser.reset(start_url="https://example.com")
        result = browser.step(BrowserAction(op="scroll", dy=500))
        print(result.observation.url)

Set HELIUM_BROWSER_IMAGE to the image published in the previous section. A Hugging Face Space image uses the form hf.co/spaces/<namespace>/<space>.

export HELIUM_BROWSER_IMAGE=hf.co/spaces/<namespace>/<space>
uv run helium-browser-hf

The runner also forwards the existing browser settings for the step limit and action waits. BROWSER_SANDBOX=hf satisfies the environment’s safety guard. The provider handles the authenticated proxy and destroys the Sandbox when its context exits. This is the same lifecycle shown in OpenEnv’s HF Sandbox example.

The OpenEnv loop

Once the server is running, reset the browser to a URL and send one action at a time. The local URL below is useful when explaining the OpenEnv protocol on its own; the HF Sandbox launcher supplies a proxied URL instead.

from helium_browser_env.client import BrowserClient
from helium_browser_env.models import BrowserAction

with BrowserClient(base_url="http://localhost:8000").sync() as browser:
    result = browser.reset(start_url="https://example.com")
    observation = result.observation

    result = browser.step(BrowserAction(op="scroll", dy=500))
    observation = result.observation

    print(observation.url)
    print(observation.error)
    print(result.done, result.reward)
    print(observation.screenshot)  # Base64-encoded PNG

reset must receive an HTTP or HTTPS URL without embedded credentials. It starts a fresh browser profile, opens the page, and returns the first observation. step performs one action and returns an OpenEnv StepResult. The result groups the next observation with the reward and episode termination flag, which gives callers the same loop shape across different environments.

Actions

OperationFieldsWhat it does
clickx, yClicks one point in the 800×600 browser viewport.
typetextTypes into the currently focused element.
keytextPresses a supported key such as ENTER, TAB, or CTRL+A.
scrolldyScrolls down for positive values and up for negative values.
backnoneGoes back in browser history.
waitnoneWaits for the page without performing another action.
finishoptional textEnds the episode.

Click coordinates are browser pixels: x is from 0 to 799 and y is from 0 to 599. Scroll values are limited to the range -1200 to 1200.

Supported keys are ENTER, TAB, ESCAPE, BACKSPACE, CTRL+A, UP, DOWN, LEFT, and RIGHT.

Why the viewport is fixed

Chromium runs on a private Xvfb display with an 800×600 viewport. The environment waits until the browser reports that exact geometry before it opens the requested page. This keeps screenshot coordinates aligned with click coordinates.

Clicks go through xdotool instead of a DOM selector. They are real pointer events, so they interact with overlays, canvas elements, and other visual controls in the same coordinate system as the screenshot. Helium handles text, keys, scrolling, navigation, and access to the Selenium driver.

Helium stores the active driver globally, and Selenium drivers are not thread-safe. The server therefore allows one environment session at a time and the environment asks OpenEnv to keep its work on one thread.

Observations and episode endings

Every observation contains:

  • screenshot: the current browser window as a base64-encoded PNG;
  • url: the current page URL;
  • error: an empty string or a short error code;
  • done: whether the episode has ended.

The OpenEnv step result carries a reward of 0 because this example does not score browser actions.

The episode ends when it receives finish, reaches the configured step limit, encounters a renderer timeout, or detects a supported bot-verification page. An action error is reported by exception class name so the caller can inspect the next screenshot and decide what to do.

Settings

Environment variableDefaultAccepted values
BROWSER_MAX_STEPS20An integer from 1 to 200.
BROWSER_ACTION_WAIT_SECONDS1.5Seconds from 0 to 10 after ordinary actions.
BROWSER_EXPLICIT_WAIT_SECONDS3Seconds from 0 to 10 for a wait action.

Boundaries

The environment is designed to run inside its browser sandbox, not against a personal desktop browser. Chromium uses a temporary profile for each episode. Downloads and password storage are disabled, and a reset rejects URLs that contain a username or password.

Web pages are untrusted input. Do not enter credentials or sensitive data, and expect some sites to show automation challenges. This example provides the browser interaction loop; it does not provide tasks, scoring, or a reward function.

Update on GitHub