OpenEnv documentation
Learn
Get Started
Concepts
Learn
Basics
Training
Harnesses
OverviewWhite-Box: Train a Web Agent on BrowserGymBlack-Box: Train Real Agents with HarborEvaluate Claude Code in an EnvironmentOpenCode (deprecated)Pi (deprecated)
Evals
Environments
API Reference
Project
You are viewing main version, which requires installation from source. If you'd like
regular pip install, checkout the latest stable version (v0.8.0).
Learn
Choose a learning path from the sidebar:
- Basics: start with Hello World, build Your First Environment and deploy it.
- Training: the ways to train and the supported frameworks, then train a reasoning model, play Wordle or 2048 with GRPO, or collect rollouts for SFT. Each framework’s own examples (ART, Miles, Oumi, SkyRL, torchforge, TRL, Unsloth, …) are listed in Integrations.
- Harnesses: pick a path: train with the trainer’s own loop on BrowserGym (white-box), train real agents through Harbor (black-box), or evaluate Claude Code inside an environment. The OpenCode and Pi tutorials are deprecated.
- Evals: follow Evaluating with Environments.
For MCP environments and rubrics, see Concepts in the sidebar.
New to OpenEnv? Start Here
The Getting Started Series walks you from zero to deploying your own environment in five short parts. No GPU required.
| Part | What it covers | Notebook |
|---|---|---|
| 1 — Introduction & Quick Start | What OpenEnv is, why it exists, and your first environment in under 10 minutes | |
| 2 — Using Environments | Connect to environments, create policies, run evaluations | |
| 3 — Building Environments | Create a custom environment from scratch | |
| 4 — Deploying an Environment | Package with Docker and deploy to Hugging Face | — |
| 5 — Contributing Environments | Publish, fork, and share environments on the Hub | — |
Topic Tutorials
Already familiar with the basics? These tutorials cover specific workflows in depth.
| Tutorial | What it covers | GPU | Notebook |
|---|---|---|---|
| Hello World | Install OpenEnv, run an environment on a local server, and build one from scratch: models, environment, server app and client. | No | |
| Train a Reasoning Model | The full pipeline: connect to reasoning_gym, wire it into TRL via environment_factory, fine-tune with GRPO, and push the checkpoint to the Hub. | Yes | |
| MCP Environments | Consume and build MCP-backed environments: list and call tools through step(), register Python functions as tools with FastMCP. | No | |
| Rubrics | Compose reward functions from reusable pieces using Gate, WeightedSum, LLMJudge, and TrajectoryRubric. | No | |
| Play Wordle with GRPO | Train an agent to play Wordle using GRPO via TRL’s environment_factory. | Yes | |
| Play 2048 with GRPO | Train a language model to play 2048 on the OpenSpiel environment with TRL’s GRPOTrainer and environment_factory. Unsloth’s own 2048 notebook trains gpt-oss-20b to write a 2048 strategy. | Yes | — |
| Evaluating with Environments | Wrap an OpenEnv environment in an Inspect AI Task, run it via InspectAIHarness, and get a structured EvalResult. | No | |
| White-Box: Train a Web Agent on BrowserGym | Train a vision-language model on BrowserGym web tasks with TRL’s GRPOTrainer and environment_factory, so TRL runs the multi-turn tool loop. | Yes | — |
| Black-Box: Train Real Agents with Harbor | Run real agents (OpenCode, Claude Code, Codex, …) on Harbor tasks with their token ids and logprobs captured, and train the model behind them with TRL’s AsyncGRPOTrainer. | Yes | — |
| Evaluate Claude Code in an Environment | Run Claude Code inside an OpenEnv environment with HarnessEnvironment (RFC 005), inject the environment’s tools over MCP, and evaluate it on τ²-bench against a simulated customer. Also serves it in production mode. | No | — |
| Training a Real Coding Agent (deprecated) | Deprecated, removed in OpenEnv 0.9.0: use Harbor with harness="opencode". Train the actual OpenCode agent (black-box, loop-owning) with TRL’s AsyncGRPOTrainer: a transparent proxy captures each turn’s token ids and logprobs while the agent runs its own tool loop. | Yes | — |
| Collect Rollouts for SFT | Run a teacher model to collect reward-labeled rollouts, filter them, and fine-tune a student with TRL’s SFTTrainer as a warm-start for GRPO. | Yes | |
| Training a Real Coding Agent (Pi, deprecated) | Deprecated, removed in OpenEnv 0.9.0: use Harbor with harness="pi". Train the actual Pi agent (black-box, loop-owning) with TRL’s AsyncGRPOTrainer: a transparent proxy captures each turn’s token ids and logprobs while the agent runs its own tool loop. | Yes | — |