beam-pi-programbench / source /study /PROGRAMBENCH.md
burtenshaw's picture
burtenshaw HF Staff
feat: publish beam pi study source
5741b22 verified
|
Raw History Blame Contribute Delete
7.38 kB
# ProgramBench controller adapter
`programbench_adapter.py` is controller/evaluator code. Never expose this checkout,
the manifest, calibration output, test cache, Docker socket, or provider credentials
to solver tools. Pi's bash backend must invoke `docker exec --user agent` in an
assigned container checkout. Normal Docker CLI environment, including `DOCKER_HOST`,
is inherited; no sudo or privileged container is required by this adapter.
On Slurm hosts where rootless Docker reports `CgroupDriver=none` and CPU/memory
limit support is false, use `--resource-mode slurm --cpus 16`: resource flags are explicitly omitted,
while Slurm supplies the shared 16-CPU/30-GiB allocation and the evaluator uses
16 pytest workers. The same allocation must apply to every study condition, with
one episode/evaluator container active at a time. Every result records this mode
and the Slurm resource environment; it does not claim container-specific limits.
Use a clean checkout of `facebookresearch/ProgramBench` at
`27f02157c785f8da3647aa6dbbe6b9137f99f10e`, with its Python dependencies installed.
Python 3.11 is supported. Prefer the upstream lockfile (`uv sync --frozen`). The
adapter verifies the Git revision and all five task/test manifest hashes. Dataset
downloads are pinned to `de0ddfb637590c7ecb54fa0b5301f6dc7dfbcee5` and anonymous.
They occur only in the controller process. Pre-pull the five immutable image refs
from the JSON manifest on the compute node; every adapter container uses
`--pull never`, so missing images fail explicitly.
Build the offline dependency cache with `evaluator-wheelhouse.py --root STUDY_ROOT`.
It downloads pinned PyPI releases, verifies release hashes, and records every wheel
hash (including the controller-built pytest-dependency wheel). Set
`PROGRAMBENCH_EVALUATOR_WHEELHOUSE=STUDY_ROOT/runtime/evaluator-wheelhouse`, or pass
`--evaluator-wheelhouse PATH` to every calibration/grade command. The reusable
rootless helper exports this setting when the cache exists. The adapter verifies
all wheel/lock hashes and requires exactly the calibration dependency fingerprint
when grading. It copies this cache only into evaluator containers, installs with
`--no-index --require-hashes`, and constrains later upstream setup installs to those
versions. Solver images/tools never receive the cache or evaluator metadata.
Global arguments precede the command. Each command emits one JSON object to stdout;
errors also emit JSON and return a nonzero exit status.
```bash
python study/programbench_adapter.py --programbench-root /path/to/ProgramBench \
--cpus 8 --memory 16g calibrate --output-dir /path/to/study/calibration --repetitions 2
python study/programbench_adapter.py --cpus 8 --memory 16g prepare \
--instance-id burntsushi__ripgrep.3b7fd44 --episode-id ripgrep-single \
--output-dir /path/to/study/episodes/ripgrep-single
python study/programbench_adapter.py snapshot --instance-id burntsushi__ripgrep.3b7fd44 \
--container CONTAINER_ID --output-dir /path/to/study/snapshots/ripgrep-single/0001
python study/programbench_adapter.py --programbench-root /path/to/ProgramBench \
--cpus 8 --memory 16g grade --instance-id burntsushi__ripgrep.3b7fd44 \
--submission /path/to/study/snapshots/ripgrep-single/0001/burntsushi__ripgrep.3b7fd44/submission.tar.gz \
--calibration /path/to/study/calibration/calibration.json \
--output-dir /path/to/study/grades/ripgrep-single/0001
```
Prepare starts a fresh, network-disabled container as `agent`, with no host mounts
or injected environment credentials. Root setup moves the original `/workspace`
into read-only, root-owned `/reference`. `/reference/executable` and bundled docs
are the solver's specification. The empty seed repository is `/workspace/solution`.
Pi creates `/workspace/.pi-study/shared.git` and per-agent checkouts under
`/workspace/.pi-study/agents/`. Only the shared repository's committed `main` is
snapshotted, using its resolved immutable commit. A solver must write `compile.sh`
which works at `/workspace` in the fresh evaluator and creates `./executable`.
Prepare emits the task prompt and full layout contract in `prepared.json`.
Calibration runs all five references, even when one fails, before inference is
admitted by the supervisor. The default is two repetitions and a 0.9 minimum
fraction on every repetition. The adapter changes only reference setup: it stashes
the cleanroom's original executable and then uses upstream branch evaluation.
Candidate builds use the upstream compiler path, including removal of prebuilt and
reference-hash artifacts. Both paths use the same immutable base image, CPU/memory
limits, official branch runner/retries, and **network disabled**. This is stricter
than upstream's test-container network default and is an explicit pilot policy.
The offline cache supplies upstream `pytest-rerunfailures==16.4`, timeout/xdist/
dependency plugins, libtmux, and pinned dependencies. The original offline attempt
failed because test setup could not install pytest-timeout; its artifacts are
preserved as `calibration-attempt-1`. The `probe` command checks only the first
active reference branch and is explicitly diagnostic, never full calibration.
Scores use upstream `test_results_map()` and `score_from_tests()`, applying active
branches and ignored-test exclusions. The frozen mask is the official mask; failed
reference tests are not removed to improve scores. Calibration fails on missing
outcomes, harness errors, unexpected tests, warnings, or a fraction below 0.9.
Failures retain all task IDs; there is no automatic replacement. Grading requires
a passed calibration report for the identical manifest and verifies mask identity.
Each grade retains the raw eval JSON, aggregate fraction, errors, completeness,
submission hash, and image digest. A candidate build failure remains a recorded
failure/zero score; it must not be dropped from analysis because `complete` is false.
Use `valid` and `analysis_score` for analysis: known candidate compile/output
failures are valid zeroes, while incomplete infrastructure observations have
`valid=false` and `analysis_score=null`. `score` always retains the raw official
fraction, and `scoring_status` distinguishes these cases.
No grading results are returned to agents.
The official evaluator requires Docker-compatible `run`, `exec`, `cp`, `commit`,
`stop`, `rm`, and `rmi`; a working rootless Docker daemon can provide these. Container
creation/commit support must be checked on the actual compute node. Local offline
tests do not establish runtime compatibility or successful reference calibration.
Offline tests (with the pinned package dependencies available):
```bash
PROGRAMBENCH_TEST_ROOT=/path/to/ProgramBench python -m pytest -q tests/test_programbench_adapter.py
```
Primary references: [official usage](https://github.com/facebookresearch/ProgramBench/blob/27f02157c785f8da3647aa6dbbe6b9137f99f10e/docs/README.md),
[evaluator](https://github.com/facebookresearch/ProgramBench/blob/27f02157c785f8da3647aa6dbbe6b9137f99f10e/src/programbench/eval/eval.py),
[scoring](https://github.com/facebookresearch/ProgramBench/blob/27f02157c785f8da3647aa6dbbe6b9137f99f10e/src/programbench/submission.py),
[baseline task rules](https://github.com/SWE-agent/mini-swe-agent/blob/04d809ceab9df28f9adaed044884180159172930/src/minisweagent/config/benchmarks/programbench.yaml).