Spaces:
Running
Download source/study/PROGRAMBENCH.md from burtenshaw/beam-pi-programbench: direct link, hf CLI and curl.
- Browser
- Download file 7.38 kB
-
https://huggingface.co/spaces/burtenshaw/beam-pi-programbench/resolve/main/source/study/PROGRAMBENCH.md
- Command line
-
hf download hf://spaces/burtenshaw/beam-pi-programbench/source/study/PROGRAMBENCH.md
-
curl -L -o PROGRAMBENCH.md https://huggingface.co/spaces/burtenshaw/beam-pi-programbench/resolve/main/source/study/PROGRAMBENCH.md
A newer version of the Gradio SDK is available: 6.30.0
ProgramBench controller adapter
programbench_adapter.py is controller/evaluator code. Never expose this checkout,
the manifest, calibration output, test cache, Docker socket, or provider credentials
to solver tools. Pi's bash backend must invoke docker exec --user agent in an
assigned container checkout. Normal Docker CLI environment, including DOCKER_HOST,
is inherited; no sudo or privileged container is required by this adapter.
On Slurm hosts where rootless Docker reports CgroupDriver=none and CPU/memory
limit support is false, use --resource-mode slurm --cpus 16: resource flags are explicitly omitted,
while Slurm supplies the shared 16-CPU/30-GiB allocation and the evaluator uses
16 pytest workers. The same allocation must apply to every study condition, with
one episode/evaluator container active at a time. Every result records this mode
and the Slurm resource environment; it does not claim container-specific limits.
Use a clean checkout of facebookresearch/ProgramBench at
27f02157c785f8da3647aa6dbbe6b9137f99f10e, with its Python dependencies installed.
Python 3.11 is supported. Prefer the upstream lockfile (uv sync --frozen). The
adapter verifies the Git revision and all five task/test manifest hashes. Dataset
downloads are pinned to de0ddfb637590c7ecb54fa0b5301f6dc7dfbcee5 and anonymous.
They occur only in the controller process. Pre-pull the five immutable image refs
from the JSON manifest on the compute node; every adapter container uses
--pull never, so missing images fail explicitly.
Build the offline dependency cache with evaluator-wheelhouse.py --root STUDY_ROOT.
It downloads pinned PyPI releases, verifies release hashes, and records every wheel
hash (including the controller-built pytest-dependency wheel). Set
PROGRAMBENCH_EVALUATOR_WHEELHOUSE=STUDY_ROOT/runtime/evaluator-wheelhouse, or pass
--evaluator-wheelhouse PATH to every calibration/grade command. The reusable
rootless helper exports this setting when the cache exists. The adapter verifies
all wheel/lock hashes and requires exactly the calibration dependency fingerprint
when grading. It copies this cache only into evaluator containers, installs with
--no-index --require-hashes, and constrains later upstream setup installs to those
versions. Solver images/tools never receive the cache or evaluator metadata.
Global arguments precede the command. Each command emits one JSON object to stdout; errors also emit JSON and return a nonzero exit status.
python study/programbench_adapter.py --programbench-root /path/to/ProgramBench \
--cpus 8 --memory 16g calibrate --output-dir /path/to/study/calibration --repetitions 2
python study/programbench_adapter.py --cpus 8 --memory 16g prepare \
--instance-id burntsushi__ripgrep.3b7fd44 --episode-id ripgrep-single \
--output-dir /path/to/study/episodes/ripgrep-single
python study/programbench_adapter.py snapshot --instance-id burntsushi__ripgrep.3b7fd44 \
--container CONTAINER_ID --output-dir /path/to/study/snapshots/ripgrep-single/0001
python study/programbench_adapter.py --programbench-root /path/to/ProgramBench \
--cpus 8 --memory 16g grade --instance-id burntsushi__ripgrep.3b7fd44 \
--submission /path/to/study/snapshots/ripgrep-single/0001/burntsushi__ripgrep.3b7fd44/submission.tar.gz \
--calibration /path/to/study/calibration/calibration.json \
--output-dir /path/to/study/grades/ripgrep-single/0001
Prepare starts a fresh, network-disabled container as agent, with no host mounts
or injected environment credentials. Root setup moves the original /workspace
into read-only, root-owned /reference. /reference/executable and bundled docs
are the solver's specification. The empty seed repository is /workspace/solution.
Pi creates /workspace/.pi-study/shared.git and per-agent checkouts under
/workspace/.pi-study/agents/. Only the shared repository's committed main is
snapshotted, using its resolved immutable commit. A solver must write compile.sh
which works at /workspace in the fresh evaluator and creates ./executable.
Prepare emits the task prompt and full layout contract in prepared.json.
Calibration runs all five references, even when one fails, before inference is
admitted by the supervisor. The default is two repetitions and a 0.9 minimum
fraction on every repetition. The adapter changes only reference setup: it stashes
the cleanroom's original executable and then uses upstream branch evaluation.
Candidate builds use the upstream compiler path, including removal of prebuilt and
reference-hash artifacts. Both paths use the same immutable base image, CPU/memory
limits, official branch runner/retries, and network disabled. This is stricter
than upstream's test-container network default and is an explicit pilot policy.
The offline cache supplies upstream pytest-rerunfailures==16.4, timeout/xdist/
dependency plugins, libtmux, and pinned dependencies. The original offline attempt
failed because test setup could not install pytest-timeout; its artifacts are
preserved as calibration-attempt-1. The probe command checks only the first
active reference branch and is explicitly diagnostic, never full calibration.
Scores use upstream test_results_map() and score_from_tests(), applying active
branches and ignored-test exclusions. The frozen mask is the official mask; failed
reference tests are not removed to improve scores. Calibration fails on missing
outcomes, harness errors, unexpected tests, warnings, or a fraction below 0.9.
Failures retain all task IDs; there is no automatic replacement. Grading requires
a passed calibration report for the identical manifest and verifies mask identity.
Each grade retains the raw eval JSON, aggregate fraction, errors, completeness,
submission hash, and image digest. A candidate build failure remains a recorded
failure/zero score; it must not be dropped from analysis because complete is false.
Use valid and analysis_score for analysis: known candidate compile/output
failures are valid zeroes, while incomplete infrastructure observations have
valid=false and analysis_score=null. score always retains the raw official
fraction, and scoring_status distinguishes these cases.
No grading results are returned to agents.
The official evaluator requires Docker-compatible run, exec, cp, commit,
stop, rm, and rmi; a working rootless Docker daemon can provide these. Container
creation/commit support must be checked on the actual compute node. Local offline
tests do not establish runtime compatibility or successful reference calibration.
Offline tests (with the pinned package dependencies available):
PROGRAMBENCH_TEST_ROOT=/path/to/ProgramBench python -m pytest -q tests/test_programbench_adapter.py
Primary references: official usage, evaluator, scoring, baseline task rules.