Datasets documentation
Load agent evaluation tasks
Load agent evaluation tasks
A Harbor dataset is a collection of tasks used to evaluate AI agents: each task gives an agent an instruction, a sandboxed environment to work in, and a verifier that scores its work.
Unlike tabular datasets, a Harbor dataset isnβt a file with one row per example: every task is a directory.
The harbor builder reads such a repository and yields one row per task, so a benchmark can be explored,
filtered and shared like any other dataset:
>>> from datasets import load_dataset
>>> dataset = load_dataset("harborframework/terminal-bench")
>>> dataset
DatasetDict({
test: Dataset({
features: ['name', 'instruction', 'description', 'keywords', 'schema_version', 'config', 'metadata', 'files', 'location'],
num_rows: 66
})
})Loading a Harbor dataset only reads the tasks: it never runs an agent, builds an image or executes a verifier, and it doesnβt require Harbor to be installed. Use the Harbor library to actually run the tasks.
Harbor task configurations are TOML files. Python 3.11 and later read them with the standard library (
tomllib); on Python 3.10, you need to install a TOML parser as an extra dependency. Check out the installation guide to learn how to install it.
Repository structure
A task directory is identified by two files, task.toml and instruction.md, and usually also contains the
environment definition, the verifier and an oracle solution:
my-harbor-dataset/
βββ README.md
βββ tasks/
βββ task-a/
β βββ task.toml # configuration: metadata, resources, timeouts...
β βββ instruction.md # what the agent is asked to do
β βββ README.md # optional human-readable description
β βββ environment/ # Dockerfile / docker-compose.yaml + files of the sandbox
β βββ tests/ # verifier
β βββ solution/ # oracle solution
βββ task-b/
βββ ...A directory is recognized as a Harbor dataset as soon as one of its sub-directories, however deep, is a task
directory (a task.toml file and an instruction.md file side by side): the tasks are found automatically,
without having to list any file.
Load tasks
Since Harbor datasets are made of directories rather than of data files, they are loaded by path β no data_files argument:
>>> from datasets import load_dataset
>>> # a dataset repository on the Hub
>>> dataset = load_dataset("harborframework/terminal-bench")
>>> # a local directory of tasks
>>> dataset = load_dataset("/path/to/my-harbor-dataset")
>>> # the tasks of a repository that are in a sub-directory
>>> dataset = load_dataset("harborframework/terminal-bench", data_dir="tasks")Passing data_files explicitly is supported too, as long as the pattern includes the task.toml and instruction.md files β one
matched file pair file gives one row:
>>> dataset = load_dataset(
... "/path/to/my-harbor-dataset",
... data_files={"test": ["my-tasks-dir/*/task.toml", "my-tasks-dir/*/instruction.md"]},
... )Columns
Every row describes one task:
| Column | Type | Description |
|---|---|---|
name | string | [task].name of task.toml (usually org/name), or the task directory name. |
description | string | [task].description of task.toml. |
instruction | string | Contents of instruction.md: what the agent has to do. |
keywords | list of string | [task].keywords of task.toml. |
schema_version | string | schema_version declared by task.toml. |
config | dict | The whole task.toml, as a JSON object: [task], [metadata], [verifier], [agent], [environment], [steps]β¦ |
metadata | dict | The free-form [metadata] section of task.toml. |
files | list of string | All the files of the task, with paths relative to the task directory. |
location | string | Location of the task directory in the dataset. |
>>> dataset = load_dataset("harborframework/terminal-bench", split="test")
>>> dataset[0]["name"]
'terminal-bench/atrx-vep-crispr'
>>> dataset[0]["instruction"][:90]
'You are provided with a region of genomic DNA representing the human ATRX locus plus sequence'
>>> dataset[0]["config"]["task"]["authors"]
[{'name': 'ScaleAI', 'email': 'tbench@scale.com'}]
>>> dataset[0]["metadata"]["category"]
'Science'
>>> dataset[0]["config"]["environment"]
{'build_timeout_sec': 1800.0, 'cpus': 2, 'memory_mb': 4096, 'storage_mb': 10240, 'gpus': 0}
>>> dataset[0]["files"][:4]
['LICENSE.md', 'README.md', 'environment/Dockerfile', 'environment/data/CDS-information.txt']
>>> dataset[0]["location"]
'tasks/atrx-vep-crispr'The whole task.toml is available in config, so nothing is lost when a task uses fields that arenβt
promoted to their own column:
>>> dataset = dataset.map(lambda task: {"timeout_min": task["config"]["verifier"]["timeout_sec"] / 60})
>>> sorted(set(dataset["timeout_min"]))[:6]
[1.0, 2.0, 3.0, 4.0, 5.0, 7.0]Filter a benchmark
Because the task metadata is loaded as a dictionary, a benchmark can be explored with the usual filter() and map() methods:
>>> security_tasks = dataset.filter(lambda task: task["metadata"].get("category") == "Security")
>>> [task["name"] for task in security_tasks]
['terminal-bench/formal-crypto', 'terminal-bench/html-js-filter', 'terminal-bench/interleaved-vigenere', 'terminal-bench/shadow-relay', 'terminal-bench/uefi-bootkit']The files column lists everything the task ships, which is handy to look at how a task is verified:
>>> pytest_verifiers = dataset.filter(lambda task: any(f.startswith("tests/") and f.endswith(".py") for f in task["files"]))
>>> pytest_verifiers.num_rows
63Read the other files of a task
Only the instruction and the configuration of a task are loaded as columns: its environment, verifier and
solution files are listed in the files column, as paths relative to the task directory. Combine task_id and files to open any of them with fsspec:
>>> from huggingface_hub import hffs
>>> dataset_path = "hf://datasets/harborframework/terminal-bench"
>>> with hffs.open(f"{dataset_path}/tasks/cumulative-layout-shift/environment/Dockerfile", "r") as f:
... print(f.read()[:20])
'# harbor-canary GUIDConfiguration
HarborConfig accepts a few arguments on top of the usual load_dataset ones:
strip_canary(defaultTrue): remove the leading βbenchmark data canaryβ comment lines from the instructions, the same way Harbor does when it sends them to an agent. Set it toFalseto keep the original file contents.include_files(defaultTrue): list all the files of every task in thefilescolumn. Set it toFalseto skip listing the task directories, which is faster when the dataset is on a remote filesystem or has thousands of files per task.
Tasks can also be streamed, which avoids downloading the whole benchmark upfront:
>>> from datasets import load_dataset
>>> dataset = load_dataset("harborframework/terminal-bench", streaming=True)
>>> for task in dataset["test"]:
... print(task["name"])See also
- The reference documentation of Harbor and HarborConfig.
- The Harbor documentation to run the tasks of a benchmark with agents and verifiers.