OpenEnv documentation
Harnesses in OpenEnv
Harnesses in OpenEnv
A harness is the loop around a model: it sends the conversation to the model, runs the tools the model calls, and decides when the episode ends. There are three ways to put a harness and an OpenEnv environment together. They differ in who owns the loop, what the environment provides and who is on the other side, and that decides what each one is for. “Black-box” is the term RFC 006 uses, and “white-box” the one openenv.core.harness uses.
| White-box | Black-box | Agent inside the environment | |
|---|---|---|---|
| Who owns the loop | The trainer (for example TRL’s environment_factory) | A real agent (OpenCode, Claude Code, Codex, …) | A real agent (Claude Code in the tutorial) |
| What the environment provides | Tools and a reward for the model being trained | Tasks: an instruction, a sandbox, and a verifier | Its own tools and state, injected into the agent over MCP |
| Who is on the other side | Nobody, the model works through the task | Nobody, the agent runs to completion | A user, simulated or real, over several turns |
| Trains a model | Yes | Yes, from the captured token ids and logprobs in a TrainingTrace (against vLLM or SGLang) | No, nothing is captured |
| Serves real users | No | No | Yes, over WS /harness |
| Example | Web tasks in BrowserGym | “Fix this repository”, “analyze these files” | Customer service on τ²-bench |
| Start here | BrowserGym | Harbor | Claude Code on τ²-bench |
Which One to Use
- You want to train a model on your environment and the loop can be simple. Use the white-box path. The trainer runs the tool loop, so token ids and logprobs come for free, and the environment only supplies tools and a reward. The BrowserGym tutorial does it with TRL’s
GRPOTrainer, and Training with OpenEnv lists the other frameworks that train on OpenEnv environments. - The job is a task that an agent does on its own, and a verifier checks the result. Use Harbor, to train the model behind the agent or to evaluate it. The agent keeps its own planner, tools and context management in a sandbox, and a proxy records every model call it makes. The resulting
TrainingTraceis plain data that doesn’t depend on a trainer, and TRL’sAsyncGRPOTrainerhas a worked example. - The job is a conversation, and the agent has to use your environment’s tools. Use
HarnessEnvironmentfrom RFC 005, to evaluate the agent against a simulated user or to serve it to real ones. The environment injects its tools, answers between turns and scores the conversation with a rubric.
Harbor or HarnessEnvironment ?
Both run a real agent, and the same agent can go either way: Claude Code is one of Harbor’s agents and the agent in the RFC 005 tutorial. What decides is the shape of the job. If an instruction goes in and a verifier checks what comes out, it is a Harbor task. If someone answers the agent between turns, or the agent works on your environment’s own state (a database, an API), use HarnessEnvironment. τ²-bench is the second kind: a simulated customer replies after every agent turn, which a Harbor task has no place for. To train, only Harbor records the tokens.
The ultimate guide to multi-harness RL goes deeper into the white-box and black-box paths and why training across several agents matters.
Update on GitHub
openenv.core.harnessalso has a session runtime (ResourceSession,MCPHarnessAdapter,build_harness_rollout_func), where OpenEnv runs a small tool loop and a model generates each turn.openenv collectuses it to collect rollouts for SFT. ItsHarnessAdapteris a different class from RFC 005’sAgenticHarnessAdapter.