OpenEnv documentation
Helium Browser Environment
Helium Browser Environment
This example shows the main pieces of an OpenEnv environment by wrapping a Chromium browser. A client sends one typed action at a time. The environment performs it with Helium and Selenium, then returns a typed observation with a fresh screenshot and the current URL.
The browser is useful here because the boundary is easy to see: OpenEnv owns
the reset/step protocol, while the environment owns all browser-specific
behavior. The example does not include an agent, so the OpenEnv parts stay in
focus.

How browser agents work
A natively multimodal browser agent operates in a loop:
- Observe: receive the current page, as a screenshot.
- Decide: choose one small action that moves the task forward.
- Act: click, type, press a key, scroll, or navigate.
- Observe again: inspect the page after the action instead of assuming it worked.
- Stop: finish when the task is complete or the episode reaches a limit.
This example uses screenshots as the main observation. The agent sees the same
800×600 browser image that the environment uses for coordinates, then returns
one structured action. A click such as x=320, y=180 means a click position in that
image; it is not a CSS selector or a reference to the page’s HTML.
One action per step matters because web pages change. A click may open a menu, show a cookie banner, navigate to another page, or fail. Returning a fresh screenshot and URL after every action lets the caller react to what actually happened. It also produces a clear trajectory: observation, action, next observation.
That loop explains the design choices used later in this example:
- Typed actions give OpenEnv a small, validated vocabulary for controlling the browser.
- A fixed viewport keeps screenshot pixels and click coordinates aligned.
- Real pointer events let clicks reach visual controls such as overlays and canvases without exposing DOM selectors to the agent.
- An isolated browser keeps temporary profiles and untrusted pages away from the user’s personal browser session.
- Step limits and
donegive every browser episode a definite boundary.
The agent or policy that chooses actions is intentionally outside this
environment. OpenEnv provides the repeatable interaction contract; the caller
decides how to turn an observation into the next BrowserAction.
Tasks
Websites always keep on changing so using a regular dataset doesn’t work. You can try the samples in this dataset as they are created late 2026. They are divided into three:
- Information: getting an information from a website by browsing it.
- Navigation: navigating a specific part of the website.
- Interaction: clicking, filtering etc.
The OpenEnv pieces
An OpenEnv environment has four small parts in this example:
| OpenEnv concept | Implementation | Responsibility |
|---|---|---|
| Data contract | BrowserAction, BrowserObservation, BrowserState | Defines the validated objects that cross the client/server boundary. |
| Environment | BrowserEnvironment | Implements the episode lifecycle: reset, step, state, and close. |
| Server | create_app(...) in server/app.py | Turns the environment class and data models into an HTTP/WebSocket service. |
| Client | BrowserClient | Serializes actions and reconstructs typed observations and step results. |
These parts separate domain logic from transport. BrowserEnvironment knows
how to drive a browser, but it does not define HTTP routes. BrowserClient knows how to exchange OpenEnv messages, but it does not import Helium or
Selenium.
1. Define the data contract
models.py subclasses OpenEnv’s three base models with the actions an agent can take: “click”, “type”, “key”, “scroll”, “back”, “wait”, “finish”. x and y are clicking coordinates, dy is change in vertical direction (scroll dy pixels), text is text to be typed on e.g. search bars.
class BrowserAction(Action):
op: Literal["click", "type", "key", "scroll", "back", "wait", "finish"]
x: int | None = Field(default=None, ge=0, lt=800)
y: int | None = Field(default=None, ge=0, lt=600)
text: str = Field(default="", max_length=2000)
dy: int = Field(default=0, ge=-1200, le=1200)
class BrowserObservation(Observation):
screenshot: str
url: str
error: str = ""
class BrowserState(State):
passThe action describes what may enter the environment. The observation describes what comes back. The state is separate: OpenEnv uses it to report episode metadata such as the episode ID and step count without adding those fields to every browser observation.
Because these are Pydantic models, validation happens at the boundary. For
example, a click outside the 800×600 viewport is rejected before browser logic
runs, and the model validator makes sure a click has both coordinates.
2. Implement the environment lifecycle
BrowserEnvironment subclasses OpenEnv’s Environment. Its methods have the
same roles they would have in a game, simulator, or other environment:
reset(...)starts a new episode and returns its first observation. Here it creates a temporary browser profile, opensstart_url, and takes a screenshot.step(action)applies exactly one validated action and returns the next observation.stateexposes the currentBrowserState, including OpenEnv’s episode ID and step count.close()releases the resources owned by the episode.
This is the Gym-like core of OpenEnv. Browser setup and Helium calls are ordinary implementation details behind that interface.
3. Turn it into a server
server/app.py passes the environment class and its wire types to create_app:
app = create_app(
BrowserEnvironment,
BrowserAction,
BrowserObservation,
env_name="helium_browser_env",
max_concurrent_envs=1,
)create_app supplies the FastAPI application and the standard OpenEnv
endpoints, including health, reset, step, state, and WebSocket access. The
environment code only implements the lifecycle; it does not duplicate those
routes.
The environment class is passed instead of a pre-built instance so OpenEnv can own its lifecycle. Concurrency is limited to one because Helium stores the active Selenium driver globally and Selenium drivers are not thread-safe.
4. Add the typed client
BrowserClient subclasses OpenEnv’s generic EnvClient with the action,
observation, and state types:
class BrowserClient(EnvClient[BrowserAction, BrowserObservation, BrowserState]):
def _step_payload(self, action):
return action.model_dump()
def _parse_result(self, payload):
...
def _parse_state(self, payload):
return BrowserState.model_validate(payload)The base client handles the connection and the common reset, step, and
state operations. This subclass only teaches it how this environment’s models
map to JSON. Callers therefore work with BrowserAction and BrowserObservation, not untyped dictionaries.
How it fits together
BrowserClient -> OpenEnv server -> BrowserEnvironment -> Helium/Selenium -> Chromium BrowserClient <- screenshot, URL, error, done <- BrowserEnvironment
For a step, BrowserClient serializes a BrowserAction. OpenEnv validates it,
calls BrowserEnvironment.step(), and serializes the returned BrowserObservation. The client then rebuilds a typed StepResult containing
the observation, reward, and done flag.
Files
| File | Purpose |
|---|---|
models.py | Defines the OpenEnv action, observation, and state contract. |
client.py | Adapts the OpenEnv client to the environment’s typed wire format. |
hf_sandbox.py | Starts a packaged environment image with OpenEnv’s HF Sandbox provider. |
server/helium_browser_environment.py | Implements OpenEnv’s lifecycle while owning Chromium. |
server/desktop.py | Starts a private X display and sends real pointer clicks with xdotool. |
server/app.py | Gives the environment and its models to OpenEnv’s create_app. |
server/Dockerfile | Builds the runnable image with Chromium and its system tools. |
pyproject.toml | Declares Python dependencies, package layout, and the server command. |
uv.lock | Pins the resolved Python dependency versions for reproducible builds. |
openenv.yaml | Describes the environment to the OpenEnv CLI. |
Why the package has both a Dockerfile and an HF Sandbox runner
The Dockerfile and HF Sandbox solve different parts of deployment.
server/Dockerfile is the image recipe. It installs Chromium, ChromeDriver,
Xvfb, and xdotool, then installs this Python package. The server command
declared in pyproject.toml starts server/app.py. This makes the image
self-contained: a runtime does not have to install browser software or copy
source files when an episode begins.
hf_sandbox.py is a client-side launcher. It asks HFSandboxProvider to start
that already-built image in an isolated HF Sandbox, waits for the OpenEnv
health endpoint, and connects BrowserClient. The Sandbox provides isolation
and a secure proxy; the image provides the software that runs inside it.
This separation is why app.py and client.py are both present:
server/app.pyruns inside the image and exposes the OpenEnv protocol;client.pyruns on the caller side and converts typed Python objects to and from that protocol.
The client is not browser automation code. Helium, Selenium, Chromium, and the virtual display stay behind the server boundary.
Package and publish the environment
From the OpenEnv repository root, enter the environment directory:
cd envs/helium_browser_envThe usual OpenEnv packaging commands are:
openenv build openenv push --repo-id <namespace>/helium-browser-env
openenv build uses server/Dockerfile with the environment directory as its
build context. openenv push publishes the same package as a Hugging Face
Space. Once the Space image has built, HF Sandboxes can refer to it as hf.co/spaces/<namespace>/helium-browser-env.
Run it in an HF Sandbox
hf_sandbox.py follows OpenEnv’s provider pattern: start a packaged environment
image in an HF Sandbox, wait for its OpenEnv server, then connect the same typed
client used with any other OpenEnv deployment.
import os
from openenv.core.containers.runtime.hf_sandbox_provider import HFSandboxProvider
from helium_browser_env.client import BrowserClient
from helium_browser_env.models import BrowserAction
with HFSandboxProvider(
image=os.environ["HELIUM_BROWSER_IMAGE"],
flavor="cpu-basic",
env_vars={"BROWSER_SANDBOX": "hf"},
) as provider:
base_url = provider.start_container()
provider.wait_for_ready(base_url, timeout_s=300.0)
with BrowserClient(base_url=base_url).sync() as browser:
result = browser.reset(start_url="https://example.com")
result = browser.step(BrowserAction(op="scroll", dy=500))
print(result.observation.url)Set HELIUM_BROWSER_IMAGE to the image published in the previous section. A
Hugging Face Space image uses the form hf.co/spaces/<namespace>/<space>.
export HELIUM_BROWSER_IMAGE=hf.co/spaces/<namespace>/<space>
uv run helium-browser-hfThe runner also forwards the existing browser settings for the step limit and
action waits. BROWSER_SANDBOX=hf satisfies the environment’s safety guard.
The provider handles the authenticated proxy and destroys the Sandbox when its
context exits. This is the same lifecycle shown in OpenEnv’s HF Sandbox example.
The OpenEnv loop
Once the server is running, reset the browser to a URL and send one action at a time. The local URL below is useful when explaining the OpenEnv protocol on its own; the HF Sandbox launcher supplies a proxied URL instead.
from helium_browser_env.client import BrowserClient
from helium_browser_env.models import BrowserAction
with BrowserClient(base_url="http://localhost:8000").sync() as browser:
result = browser.reset(start_url="https://example.com")
observation = result.observation
result = browser.step(BrowserAction(op="scroll", dy=500))
observation = result.observation
print(observation.url)
print(observation.error)
print(result.done, result.reward)
print(observation.screenshot) # Base64-encoded PNGreset must receive an HTTP or HTTPS URL without embedded credentials. It
starts a fresh browser profile, opens the page, and returns the first
observation. step performs one action and returns an OpenEnv StepResult.
The result groups the next observation with the reward and episode termination
flag, which gives callers the same loop shape across different environments.
Actions
| Operation | Fields | What it does |
|---|---|---|
click | x, y | Clicks one point in the 800×600 browser viewport. |
type | text | Types into the currently focused element. |
key | text | Presses a supported key such as ENTER, TAB, or CTRL+A. |
scroll | dy | Scrolls down for positive values and up for negative values. |
back | none | Goes back in browser history. |
wait | none | Waits for the page without performing another action. |
finish | optional text | Ends the episode. |
Click coordinates are browser pixels: x is from 0 to 799 and y is from 0
to 599. Scroll values are limited to the range -1200 to 1200.
Supported keys are ENTER, TAB, ESCAPE, BACKSPACE, CTRL+A, UP, DOWN, LEFT, and RIGHT.
Why the viewport is fixed
Chromium runs on a private Xvfb display with an 800×600 viewport. The environment waits until the browser reports that exact geometry before it opens the requested page. This keeps screenshot coordinates aligned with click coordinates.
Clicks go through xdotool instead of a DOM selector. They are real pointer
events, so they interact with overlays, canvas elements, and other visual
controls in the same coordinate system as the screenshot. Helium handles text,
keys, scrolling, navigation, and access to the Selenium driver.
Helium stores the active driver globally, and Selenium drivers are not thread-safe. The server therefore allows one environment session at a time and the environment asks OpenEnv to keep its work on one thread.
Observations and episode endings
Every observation contains:
screenshot: the current browser window as a base64-encoded PNG;url: the current page URL;error: an empty string or a short error code;done: whether the episode has ended.
The OpenEnv step result carries a reward of 0 because this example does not
score browser actions.
The episode ends when it receives finish, reaches the configured step limit,
encounters a renderer timeout, or detects a supported bot-verification page.
An action error is reported by exception class name so the caller can inspect
the next screenshot and decide what to do.
Settings
| Environment variable | Default | Accepted values |
|---|---|---|
BROWSER_MAX_STEPS | 20 | An integer from 1 to 200. |
BROWSER_ACTION_WAIT_SECONDS | 1.5 | Seconds from 0 to 10 after ordinary actions. |
BROWSER_EXPLICIT_WAIT_SECONDS | 3 | Seconds from 0 to 10 for a wait action. |
Boundaries
The environment is designed to run inside its browser sandbox, not against a personal desktop browser. Chromium uses a temporary profile for each episode. Downloads and password storage are disabled, and a reset rejects URLs that contain a username or password.
Web pages are untrusted input. Do not enter credentials or sensitive data, and expect some sites to show automation challenges. This example provides the browser interaction loop; it does not provide tasks, scoring, or a reward function.
Update on GitHub