Datasets documentation

Load agent evaluation tasks

You are viewing main version, which requires installation from source. If you'd like regular pip install, checkout the latest stable version (v4.8.4).
Hugging Face's logo
Join the Hugging Face community

and get access to the augmented documentation experience

to get started

Load agent evaluation tasks

A Harbor dataset is a collection of tasks used to evaluate AI agents: each task gives an agent an instruction, a sandboxed environment to work in, and a verifier that scores its work.

Unlike tabular datasets, a Harbor dataset isn’t a file with one row per example: every task is a directory. The harbor builder reads such a repository and yields one row per task, so a benchmark can be explored, filtered and shared like any other dataset:

>>> from datasets import load_dataset
>>> dataset = load_dataset("harborframework/terminal-bench")
>>> dataset
DatasetDict({
    test: Dataset({
        features: ['name', 'instruction', 'description', 'keywords', 'schema_version', 'config', 'metadata', 'files', 'location'],
        num_rows: 66
    })
})

Loading a Harbor dataset only reads the tasks: it never runs an agent, builds an image or executes a verifier, and it doesn’t require Harbor to be installed. Use the Harbor library to actually run the tasks.

Harbor task configurations are TOML files. Python 3.11 and later read them with the standard library (tomllib); on Python 3.10, you need to install a TOML parser as an extra dependency. Check out the installation guide to learn how to install it.

Repository structure

A task directory is identified by two files, task.toml and instruction.md, and usually also contains the environment definition, the verifier and an oracle solution:

my-harbor-dataset/
β”œβ”€β”€ README.md
└── tasks/
    β”œβ”€β”€ task-a/
    β”‚   β”œβ”€β”€ task.toml          # configuration: metadata, resources, timeouts...
    β”‚   β”œβ”€β”€ instruction.md     # what the agent is asked to do
    β”‚   β”œβ”€β”€ README.md          # optional human-readable description
    β”‚   β”œβ”€β”€ environment/       # Dockerfile / docker-compose.yaml + files of the sandbox
    β”‚   β”œβ”€β”€ tests/             # verifier
    β”‚   └── solution/          # oracle solution
    └── task-b/
        └── ...

A directory is recognized as a Harbor dataset as soon as one of its sub-directories, however deep, is a task directory (a task.toml file and an instruction.md file side by side): the tasks are found automatically, without having to list any file.

Load tasks

Since Harbor datasets are made of directories rather than of data files, they are loaded by path β€” no data_files argument:

>>> from datasets import load_dataset

>>> # a dataset repository on the Hub
>>> dataset = load_dataset("harborframework/terminal-bench")
>>> # a local directory of tasks
>>> dataset = load_dataset("/path/to/my-harbor-dataset")
>>> # the tasks of a repository that are in a sub-directory
>>> dataset = load_dataset("harborframework/terminal-bench", data_dir="tasks")

Passing data_files explicitly is supported too, as long as the pattern includes the task.toml and instruction.md files β€” one matched file pair file gives one row:

>>> dataset = load_dataset(
...     "/path/to/my-harbor-dataset",
...     data_files={"test": ["my-tasks-dir/*/task.toml", "my-tasks-dir/*/instruction.md"]},
... )

Columns

Every row describes one task:

ColumnTypeDescription
namestring[task].name of task.toml (usually org/name), or the task directory name.
descriptionstring[task].description of task.toml.
instructionstringContents of instruction.md: what the agent has to do.
keywordslist of string[task].keywords of task.toml.
schema_versionstringschema_version declared by task.toml.
configdictThe whole task.toml, as a JSON object: [task], [metadata], [verifier], [agent], [environment], [steps]…
metadatadictThe free-form [metadata] section of task.toml.
fileslist of stringAll the files of the task, with paths relative to the task directory.
locationstringLocation of the task directory in the dataset.
>>> dataset = load_dataset("harborframework/terminal-bench", split="test")
>>> dataset[0]["name"]
'terminal-bench/atrx-vep-crispr'
>>> dataset[0]["instruction"][:90]
'You are provided with a region of genomic DNA representing the human ATRX locus plus sequence'
>>> dataset[0]["config"]["task"]["authors"]
[{'name': 'ScaleAI', 'email': 'tbench@scale.com'}]
>>> dataset[0]["metadata"]["category"]
'Science'
>>> dataset[0]["config"]["environment"]
{'build_timeout_sec': 1800.0, 'cpus': 2, 'memory_mb': 4096, 'storage_mb': 10240, 'gpus': 0}
>>> dataset[0]["files"][:4]
['LICENSE.md', 'README.md', 'environment/Dockerfile', 'environment/data/CDS-information.txt']
>>> dataset[0]["location"]
'tasks/atrx-vep-crispr'

The whole task.toml is available in config, so nothing is lost when a task uses fields that aren’t promoted to their own column:

>>> dataset = dataset.map(lambda task: {"timeout_min": task["config"]["verifier"]["timeout_sec"] / 60})
>>> sorted(set(dataset["timeout_min"]))[:6]
[1.0, 2.0, 3.0, 4.0, 5.0, 7.0]

Filter a benchmark

Because the task metadata is loaded as a dictionary, a benchmark can be explored with the usual filter() and map() methods:

>>> security_tasks = dataset.filter(lambda task: task["metadata"].get("category") == "Security")
>>> [task["name"] for task in security_tasks]
['terminal-bench/formal-crypto', 'terminal-bench/html-js-filter', 'terminal-bench/interleaved-vigenere', 'terminal-bench/shadow-relay', 'terminal-bench/uefi-bootkit']

The files column lists everything the task ships, which is handy to look at how a task is verified:

>>> pytest_verifiers = dataset.filter(lambda task: any(f.startswith("tests/") and f.endswith(".py") for f in task["files"]))
>>> pytest_verifiers.num_rows
63

Read the other files of a task

Only the instruction and the configuration of a task are loaded as columns: its environment, verifier and solution files are listed in the files column, as paths relative to the task directory. Combine task_id and files to open any of them with fsspec:

>>> from huggingface_hub import hffs
>>> dataset_path = "hf://datasets/harborframework/terminal-bench"
>>> with hffs.open(f"{dataset_path}/tasks/cumulative-layout-shift/environment/Dockerfile", "r") as f:
...     print(f.read()[:20])
'# harbor-canary GUID

Configuration

HarborConfig accepts a few arguments on top of the usual load_dataset ones:

  • strip_canary (default True): remove the leading β€œbenchmark data canary” comment lines from the instructions, the same way Harbor does when it sends them to an agent. Set it to False to keep the original file contents.
  • include_files (default True): list all the files of every task in the files column. Set it to False to skip listing the task directories, which is faster when the dataset is on a remote filesystem or has thousands of files per task.

Tasks can also be streamed, which avoids downloading the whole benchmark upfront:

>>> from datasets import load_dataset
>>> dataset = load_dataset("harborframework/terminal-bench", streaming=True)
>>> for task in dataset["test"]:
...     print(task["name"])

See also

Update on GitHub