Spaces:
Running
Running
|
Download source/study/PROGRAMBENCH.md from burtenshaw/beam-pi-programbench: direct link, hf CLI and curl.
- Browser
- Download file 7.38 kB
-
https://huggingface.co/spaces/burtenshaw/beam-pi-programbench/resolve/main/source/study/PROGRAMBENCH.md
- Command line
-
hf download hf://spaces/burtenshaw/beam-pi-programbench/source/study/PROGRAMBENCH.md
-
curl -L -o PROGRAMBENCH.md https://huggingface.co/spaces/burtenshaw/beam-pi-programbench/resolve/main/source/study/PROGRAMBENCH.md
7.38 kB
| # ProgramBench controller adapter | |
| `programbench_adapter.py` is controller/evaluator code. Never expose this checkout, | |
| the manifest, calibration output, test cache, Docker socket, or provider credentials | |
| to solver tools. Pi's bash backend must invoke `docker exec --user agent` in an | |
| assigned container checkout. Normal Docker CLI environment, including `DOCKER_HOST`, | |
| is inherited; no sudo or privileged container is required by this adapter. | |
| On Slurm hosts where rootless Docker reports `CgroupDriver=none` and CPU/memory | |
| limit support is false, use `--resource-mode slurm --cpus 16`: resource flags are explicitly omitted, | |
| while Slurm supplies the shared 16-CPU/30-GiB allocation and the evaluator uses | |
| 16 pytest workers. The same allocation must apply to every study condition, with | |
| one episode/evaluator container active at a time. Every result records this mode | |
| and the Slurm resource environment; it does not claim container-specific limits. | |
| Use a clean checkout of `facebookresearch/ProgramBench` at | |
| `27f02157c785f8da3647aa6dbbe6b9137f99f10e`, with its Python dependencies installed. | |
| Python 3.11 is supported. Prefer the upstream lockfile (`uv sync --frozen`). The | |
| adapter verifies the Git revision and all five task/test manifest hashes. Dataset | |
| downloads are pinned to `de0ddfb637590c7ecb54fa0b5301f6dc7dfbcee5` and anonymous. | |
| They occur only in the controller process. Pre-pull the five immutable image refs | |
| from the JSON manifest on the compute node; every adapter container uses | |
| `--pull never`, so missing images fail explicitly. | |
| Build the offline dependency cache with `evaluator-wheelhouse.py --root STUDY_ROOT`. | |
| It downloads pinned PyPI releases, verifies release hashes, and records every wheel | |
| hash (including the controller-built pytest-dependency wheel). Set | |
| `PROGRAMBENCH_EVALUATOR_WHEELHOUSE=STUDY_ROOT/runtime/evaluator-wheelhouse`, or pass | |
| `--evaluator-wheelhouse PATH` to every calibration/grade command. The reusable | |
| rootless helper exports this setting when the cache exists. The adapter verifies | |
| all wheel/lock hashes and requires exactly the calibration dependency fingerprint | |
| when grading. It copies this cache only into evaluator containers, installs with | |
| `--no-index --require-hashes`, and constrains later upstream setup installs to those | |
| versions. Solver images/tools never receive the cache or evaluator metadata. | |
| Global arguments precede the command. Each command emits one JSON object to stdout; | |
| errors also emit JSON and return a nonzero exit status. | |
| ```bash | |
| python study/programbench_adapter.py --programbench-root /path/to/ProgramBench \ | |
| --cpus 8 --memory 16g calibrate --output-dir /path/to/study/calibration --repetitions 2 | |
| python study/programbench_adapter.py --cpus 8 --memory 16g prepare \ | |
| --instance-id burntsushi__ripgrep.3b7fd44 --episode-id ripgrep-single \ | |
| --output-dir /path/to/study/episodes/ripgrep-single | |
| python study/programbench_adapter.py snapshot --instance-id burntsushi__ripgrep.3b7fd44 \ | |
| --container CONTAINER_ID --output-dir /path/to/study/snapshots/ripgrep-single/0001 | |
| python study/programbench_adapter.py --programbench-root /path/to/ProgramBench \ | |
| --cpus 8 --memory 16g grade --instance-id burntsushi__ripgrep.3b7fd44 \ | |
| --submission /path/to/study/snapshots/ripgrep-single/0001/burntsushi__ripgrep.3b7fd44/submission.tar.gz \ | |
| --calibration /path/to/study/calibration/calibration.json \ | |
| --output-dir /path/to/study/grades/ripgrep-single/0001 | |
| ``` | |
| Prepare starts a fresh, network-disabled container as `agent`, with no host mounts | |
| or injected environment credentials. Root setup moves the original `/workspace` | |
| into read-only, root-owned `/reference`. `/reference/executable` and bundled docs | |
| are the solver's specification. The empty seed repository is `/workspace/solution`. | |
| Pi creates `/workspace/.pi-study/shared.git` and per-agent checkouts under | |
| `/workspace/.pi-study/agents/`. Only the shared repository's committed `main` is | |
| snapshotted, using its resolved immutable commit. A solver must write `compile.sh` | |
| which works at `/workspace` in the fresh evaluator and creates `./executable`. | |
| Prepare emits the task prompt and full layout contract in `prepared.json`. | |
| Calibration runs all five references, even when one fails, before inference is | |
| admitted by the supervisor. The default is two repetitions and a 0.9 minimum | |
| fraction on every repetition. The adapter changes only reference setup: it stashes | |
| the cleanroom's original executable and then uses upstream branch evaluation. | |
| Candidate builds use the upstream compiler path, including removal of prebuilt and | |
| reference-hash artifacts. Both paths use the same immutable base image, CPU/memory | |
| limits, official branch runner/retries, and **network disabled**. This is stricter | |
| than upstream's test-container network default and is an explicit pilot policy. | |
| The offline cache supplies upstream `pytest-rerunfailures==16.4`, timeout/xdist/ | |
| dependency plugins, libtmux, and pinned dependencies. The original offline attempt | |
| failed because test setup could not install pytest-timeout; its artifacts are | |
| preserved as `calibration-attempt-1`. The `probe` command checks only the first | |
| active reference branch and is explicitly diagnostic, never full calibration. | |
| Scores use upstream `test_results_map()` and `score_from_tests()`, applying active | |
| branches and ignored-test exclusions. The frozen mask is the official mask; failed | |
| reference tests are not removed to improve scores. Calibration fails on missing | |
| outcomes, harness errors, unexpected tests, warnings, or a fraction below 0.9. | |
| Failures retain all task IDs; there is no automatic replacement. Grading requires | |
| a passed calibration report for the identical manifest and verifies mask identity. | |
| Each grade retains the raw eval JSON, aggregate fraction, errors, completeness, | |
| submission hash, and image digest. A candidate build failure remains a recorded | |
| failure/zero score; it must not be dropped from analysis because `complete` is false. | |
| Use `valid` and `analysis_score` for analysis: known candidate compile/output | |
| failures are valid zeroes, while incomplete infrastructure observations have | |
| `valid=false` and `analysis_score=null`. `score` always retains the raw official | |
| fraction, and `scoring_status` distinguishes these cases. | |
| No grading results are returned to agents. | |
| The official evaluator requires Docker-compatible `run`, `exec`, `cp`, `commit`, | |
| `stop`, `rm`, and `rmi`; a working rootless Docker daemon can provide these. Container | |
| creation/commit support must be checked on the actual compute node. Local offline | |
| tests do not establish runtime compatibility or successful reference calibration. | |
| Offline tests (with the pinned package dependencies available): | |
| ```bash | |
| PROGRAMBENCH_TEST_ROOT=/path/to/ProgramBench python -m pytest -q tests/test_programbench_adapter.py | |
| ``` | |
| Primary references: [official usage](https://github.com/facebookresearch/ProgramBench/blob/27f02157c785f8da3647aa6dbbe6b9137f99f10e/docs/README.md), | |
| [evaluator](https://github.com/facebookresearch/ProgramBench/blob/27f02157c785f8da3647aa6dbbe6b9137f99f10e/src/programbench/eval/eval.py), | |
| [scoring](https://github.com/facebookresearch/ProgramBench/blob/27f02157c785f8da3647aa6dbbe6b9137f99f10e/src/programbench/submission.py), | |
| [baseline task rules](https://github.com/SWE-agent/mini-swe-agent/blob/04d809ceab9df28f9adaed044884180159172930/src/minisweagent/config/benchmarks/programbench.yaml). | |