beam-pi-programbench / source /study /PROGRAMBENCH.md
burtenshaw's picture
burtenshaw HF Staff
feat: publish beam pi study source
5741b22 verified
|
Raw History Blame Contribute Delete
7.38 kB

A newer version of the Gradio SDK is available: 6.30.0

Upgrade

ProgramBench controller adapter

programbench_adapter.py is controller/evaluator code. Never expose this checkout, the manifest, calibration output, test cache, Docker socket, or provider credentials to solver tools. Pi's bash backend must invoke docker exec --user agent in an assigned container checkout. Normal Docker CLI environment, including DOCKER_HOST, is inherited; no sudo or privileged container is required by this adapter.

On Slurm hosts where rootless Docker reports CgroupDriver=none and CPU/memory limit support is false, use --resource-mode slurm --cpus 16: resource flags are explicitly omitted, while Slurm supplies the shared 16-CPU/30-GiB allocation and the evaluator uses 16 pytest workers. The same allocation must apply to every study condition, with one episode/evaluator container active at a time. Every result records this mode and the Slurm resource environment; it does not claim container-specific limits.

Use a clean checkout of facebookresearch/ProgramBench at 27f02157c785f8da3647aa6dbbe6b9137f99f10e, with its Python dependencies installed. Python 3.11 is supported. Prefer the upstream lockfile (uv sync --frozen). The adapter verifies the Git revision and all five task/test manifest hashes. Dataset downloads are pinned to de0ddfb637590c7ecb54fa0b5301f6dc7dfbcee5 and anonymous. They occur only in the controller process. Pre-pull the five immutable image refs from the JSON manifest on the compute node; every adapter container uses --pull never, so missing images fail explicitly.

Build the offline dependency cache with evaluator-wheelhouse.py --root STUDY_ROOT. It downloads pinned PyPI releases, verifies release hashes, and records every wheel hash (including the controller-built pytest-dependency wheel). Set PROGRAMBENCH_EVALUATOR_WHEELHOUSE=STUDY_ROOT/runtime/evaluator-wheelhouse, or pass --evaluator-wheelhouse PATH to every calibration/grade command. The reusable rootless helper exports this setting when the cache exists. The adapter verifies all wheel/lock hashes and requires exactly the calibration dependency fingerprint when grading. It copies this cache only into evaluator containers, installs with --no-index --require-hashes, and constrains later upstream setup installs to those versions. Solver images/tools never receive the cache or evaluator metadata.

Global arguments precede the command. Each command emits one JSON object to stdout; errors also emit JSON and return a nonzero exit status.

python study/programbench_adapter.py --programbench-root /path/to/ProgramBench \
  --cpus 8 --memory 16g calibrate --output-dir /path/to/study/calibration --repetitions 2

python study/programbench_adapter.py --cpus 8 --memory 16g prepare \
  --instance-id burntsushi__ripgrep.3b7fd44 --episode-id ripgrep-single \
  --output-dir /path/to/study/episodes/ripgrep-single

python study/programbench_adapter.py snapshot --instance-id burntsushi__ripgrep.3b7fd44 \
  --container CONTAINER_ID --output-dir /path/to/study/snapshots/ripgrep-single/0001

python study/programbench_adapter.py --programbench-root /path/to/ProgramBench \
  --cpus 8 --memory 16g grade --instance-id burntsushi__ripgrep.3b7fd44 \
  --submission /path/to/study/snapshots/ripgrep-single/0001/burntsushi__ripgrep.3b7fd44/submission.tar.gz \
  --calibration /path/to/study/calibration/calibration.json \
  --output-dir /path/to/study/grades/ripgrep-single/0001

Prepare starts a fresh, network-disabled container as agent, with no host mounts or injected environment credentials. Root setup moves the original /workspace into read-only, root-owned /reference. /reference/executable and bundled docs are the solver's specification. The empty seed repository is /workspace/solution. Pi creates /workspace/.pi-study/shared.git and per-agent checkouts under /workspace/.pi-study/agents/. Only the shared repository's committed main is snapshotted, using its resolved immutable commit. A solver must write compile.sh which works at /workspace in the fresh evaluator and creates ./executable. Prepare emits the task prompt and full layout contract in prepared.json.

Calibration runs all five references, even when one fails, before inference is admitted by the supervisor. The default is two repetitions and a 0.9 minimum fraction on every repetition. The adapter changes only reference setup: it stashes the cleanroom's original executable and then uses upstream branch evaluation. Candidate builds use the upstream compiler path, including removal of prebuilt and reference-hash artifacts. Both paths use the same immutable base image, CPU/memory limits, official branch runner/retries, and network disabled. This is stricter than upstream's test-container network default and is an explicit pilot policy. The offline cache supplies upstream pytest-rerunfailures==16.4, timeout/xdist/ dependency plugins, libtmux, and pinned dependencies. The original offline attempt failed because test setup could not install pytest-timeout; its artifacts are preserved as calibration-attempt-1. The probe command checks only the first active reference branch and is explicitly diagnostic, never full calibration.

Scores use upstream test_results_map() and score_from_tests(), applying active branches and ignored-test exclusions. The frozen mask is the official mask; failed reference tests are not removed to improve scores. Calibration fails on missing outcomes, harness errors, unexpected tests, warnings, or a fraction below 0.9. Failures retain all task IDs; there is no automatic replacement. Grading requires a passed calibration report for the identical manifest and verifies mask identity. Each grade retains the raw eval JSON, aggregate fraction, errors, completeness, submission hash, and image digest. A candidate build failure remains a recorded failure/zero score; it must not be dropped from analysis because complete is false. Use valid and analysis_score for analysis: known candidate compile/output failures are valid zeroes, while incomplete infrastructure observations have valid=false and analysis_score=null. score always retains the raw official fraction, and scoring_status distinguishes these cases. No grading results are returned to agents.

The official evaluator requires Docker-compatible run, exec, cp, commit, stop, rm, and rmi; a working rootless Docker daemon can provide these. Container creation/commit support must be checked on the actual compute node. Local offline tests do not establish runtime compatibility or successful reference calibration.

Offline tests (with the pinned package dependencies available):

PROGRAMBENCH_TEST_ROOT=/path/to/ProgramBench python -m pytest -q tests/test_programbench_adapter.py

Primary references: official usage, evaluator, scoring, baseline task rules.