OpenEnv documentation
τ²-bench Environment
τ²-bench Environment
τ²-bench as an OpenEnv environment. Each episode is a customer-service conversation: an LLM-simulated user comes with a request, and the agent has to solve it with the domain’s tools while following the domain’s policy. The user replies to what the agent says, changes their mind, or pushes for things the policy does not allow. When the conversation ends, τ²-bench’s own evaluator scores it, mostly from the final state of the database.
That makes it multi-turn (a real back-and-forth with a user), multi-step (several tool calls per turn), and verifiable. It works for evaluation and as an RL environment.
| Domain | Tasks (train / test) | Tools | Reward |
|---|---|---|---|
airline | 30 / 20 | book, cancel, change flights, baggage, certificates… | database state + information given to the user |
retail | 74 / 40 | orders, returns, exchanges, addresses, payments… | database state + natural-language assertions (judged by an LLM) |
telecom | 74 / 40 | account and line management; the user also has tools on their phone | assertions on the final state of the account and the phone |
telecom-workflow, banking_knowledge and mock are also available (banking_knowledge needs τ²-bench’s knowledge extra).
Web UI
The Space opens on a τ²-bench tab. The domain, the split and the simulated customer’s model can be changed at the top, and there are four views:
- Tasks: an explorer of the split’s tasks, with search and filters (changes the database or read-only). A task page shows what the agent does not see (why the customer calls, what they know, how they behave) and what the conversation is scored on.
- Run a model: pick an agent model on Hugging Face Inference Providers (the list shows each model’s providers and price) and watch the conversation, with each tool call as a collapsible row and τ²-bench’s reward breakdown at the end. A run can be stopped.
- Play as the agent: talk to the simulated customer yourself and call the domain’s tools from a form, then see your reward.
- Runs: the runs of your session. Open one to read it again or download it as JSON, or compare two to four side by side.
On a Space, visitors sign in with Hugging Face, and their conversations run on their own Inference Providers credits (the Space asks for the inference-api scope). Running a conversation requires signing in, and browsing the tasks doesn’t.
The runs live in the browser session and are not kept across restarts. The generic OpenEnv playground stays available as a second tab.
Tools
The agent acts through MCP tools:
- The domain’s tools, e.g.
get_user_details,search_direct_flight,book_reservation. They read and write the task’s database. respond_to_user(message): says something to the user and returns their reply.done(): ends the conversation from the agent’s side.
The episode ends when the user is done (or the agent calls done), and the last observation carries done=True, the reward, and τ²-bench’s breakdown in metadata["reward_info"].
The simulated user
The user (and the judge for retail’s natural-language assertions) is an LLM. It runs on Hugging Face Inference Providers by default, with deepseek-ai/DeepSeek-V4.1-Flash:
TAU2_USER_PROVIDER | Credential | Default TAU2_USER_MODEL |
|---|---|---|
hf (default) | HF_TOKEN | deepseek-ai/DeepSeek-V4.1-Flash |
openai | OPENAI_API_KEY | gpt-4.1 |
anthropic | ANTHROPIC_API_KEY | claude-sonnet-4-5 |
Any model of the provider works through TAU2_USER_MODEL (for hf, any Inference Providers chat model). Non-reasoning models answer fastest, which matters when the user is called thousands of times during training.
τ²-bench’s published results use
gpt-4.1as the user. Scores with a different user model are not comparable to the leaderboard.
Quick start
from openenv.core.env_server.mcp_types import CallToolAction
from tau2_env import Tau2Env
async with Tau2Env(base_url="http://localhost:8000") as env:
result = await env.reset(task_id="2")
print(result.observation.metadata["user_message"]) # the user's opening message
policy = result.observation.metadata["policy"] # what the agent must follow
tools = await env.list_tools()
reply = await env.call_tool(
"respond_to_user", message="Happy to help. What's your user id?"
)
user = await env.call_tool("get_user_details", user_id="noah_muller_9847")
# `call_tool` returns the tool's output only. `step` also returns whether the
# episode is over and, once it is, the reward.
result = await env.step(CallToolAction(tool_name="done", arguments={}))
print(result.done, result.reward, result.observation.metadata["reward_info"])reset() takes task_id to run a given task, or picks one of the split at random (seed makes it reproducible). Give the agent the policy as its system prompt. With the default hf provider, reset(hf_token=...) makes the session’s user and judge run on your token, which is how to use a server that has none, like the public Space (base_url="https://sergiopaniego-tau2-env.hf.space").
Running the server
τ²-bench’s tasks, databases and policies live in its repository rather than in the package, so outside Docker fetch them once (the commit matches the tau2 pin in pyproject.toml) and point TAU2_DATA_DIR at them:
git clone --filter=blob:none --no-checkout https://github.com/sierra-research/tau2-bench
git -C tau2-bench sparse-checkout set --no-cone data/tau2/domains data/tau2/user_simulator
git -C tau2-bench checkout 5bfa7e37b36656b37dc6d022156be6563c1007f3
cd envs/tau2_env
TAU2_DATA_DIR=../../tau2-bench/data HF_TOKEN=hf_... ENABLE_WEB_INTERFACE=true uv run serverThe Docker image does this at build time.
| Variable | Default | |
|---|---|---|
TAU2_DOMAIN | airline | τ²-bench domain |
TAU2_SPLIT | test | train, test or base (all tasks) |
TAU2_USER_PROVIDER | hf | hf, openai or anthropic |
TAU2_USER_MODEL | the provider’s default | model for the user and the judge |
TAU2_DATA_DIR | set in the Docker image | τ²-bench’s data folder |
MAX_CONCURRENT_ENVS | 8 | API sessions at once |
With Docker:
docker build -t tau2-env -f envs/tau2_env/server/Dockerfile envs/tau2_env docker run -p 8000:8000 -e HF_TOKEN=hf_... tau2-env
On a Space, the UI’s runs use each visitor’s own token once they sign in, so the Space needs no secret. They always run on Inference Providers, whatever TAU2_USER_PROVIDER is. API clients pass their own with reset(hf_token=...), so an HF_TOKEN secret is only a fallback for clients that don’t, and the UI never uses it. Without credentials the server and the task explorer still run, and reset() explains what is missing.
The image starts from a Python 3.12 base rather than openenv-base, because τ²-bench requires Python 3.12+.
Notes
- τ²-bench is installed from GitHub, pinned to a commit. The
tau2package on PyPI is an unrelated project. - To evaluate Claude Code on τ²-bench, see Evaluate Claude Code in an Environment.
- Train on the
trainsplit and evaluate ontest, so training does not leak into the benchmark. - litellm has no prices for Inference Providers models, so τ²-bench reports the cost of those runs as $0.