OpenEnv documentation

Black-Box: Train Real Agents with Harbor

You are viewing main version, which requires installation from source. If you'd like regular pip install, checkout the latest stable version (v0.8.0).
Hugging Face's logo
Join the Hugging Face community

and get access to the augmented documentation experience

to get started

Black-Box: Train Real Agents with Harbor

Some agents can’t be reimplemented as a trainer’s tool loop. A coding agent such as OpenCode, Claude Code or Codex has its own planner, tools, context management and stop condition, and the point is to train the model that drives that agent. This is the black-box harness path: the agent owns its loop, and OpenEnv records every model call it makes. Episodes are tasks: an instruction goes in, the agent works on its own, and a verifier checks the result. If the job is a conversation instead, or the agent has to use your environment’s own tools, see Evaluate Claude Code in an Environment. Harnesses in OpenEnv compares the paths.

OpenEnv does this through Harbor, which supplies the tasks, the sandboxes, the agents and the verifiers. One harbor_env server runs any of 16 validated agents on any Harbor dataset, in any of Harbor’s sandbox backends, and its capture proxy returns the token ids and logprobs of each model call together with the task’s reward.

Try One Rollout

With an OpenAI-compatible endpoint in $LLM and Python 3.12 or later:

pip install "openenv[harbor]"

openenv harbor rollout \
  --llm-url $LLM \
  --dataset AdithyaSK/data_agent_rl_environment_eval \
  --task-index 0 --harness opencode --sandbox e2b \
  --out rollout.json

Against a hosted provider you get an evaluation rollout: the reward and the full trace. Against vLLM or SGLang you also get the token ids and logprobs that training needs. openenv harbor info checks first which agents, sandboxes and capture level this machine can use.

Train on the Captures

Each rollout comes back as a TrainingTrace: the token ids, logprobs and loss masks of every model call, plus the task’s reward. It doesn’t depend on a trainer, so any framework can train on it. TRL has a worked example: AsyncGRPOTrainer with a HarnessRolloutWorker runs each task to completion, reads back the trace, and syncs the new weights into the same vLLM server.

Learn More

Update on GitHub