# Verification — 2026-10-05 ## Passed - Python 3.12 in a fresh isolated virtual environment using the pinned project dependencies - 75 pytest cases, including the original ten acceptance requirements, clean controls, explicit known failures, scoped authority, timestamps, exact arguments, single-use approvals, execution-time revalidation, malformed inputs, deterministic audits, oracle isolation, UI grading validity, repeat/concurrent callbacks, and exports - Ruff static checks - Dependency consistency (`pip check`) - Eleven registered scenarios through all four comparison modes - Deliberately unsafe policy experiment (`experiments/bad_policy.yaml`) - Actual loopback HTTP/API smoke: page/config response, four-mode comparison, malformed-input error, successful repeated clean run, and downloadable artifacts - Combined JSONL export plus all original per-mode event chains verify successfully - Independent code and test review; identified approval, grading, audit robustness, and export issues were corrected and retested ## Observed authored-suite results All four modes complete all four initially sufficient clean controls. Full BORDER completes seven of eleven total fixtures; it resolves the three supplied recoverable adverse scenarios and correctly abstains on unresolved conflict, expired evidence, and absent approval. It executes eight goal actions, of which one is oracle-incorrect: `compromised_official`, the intentional authoritative-source failure. Its three attempted recoveries are oracle-correct. The simple action gate blocks missing approval. Provenance Only uses the same scripted proposal as Baseline and only annotates metadata, so their behavior is identical here. The deliberately weak email-authority policy produces another visible incorrect action while complying with that policy. See `results/suite.json` for full numerator/denominator counts, per-mode traces and local harness timing. These are regression fixtures, not independent benchmark evidence. No token savings, general security effectiveness, or production-readiness claim is made. ## Not verified - Visual browser layout or screenshots. Installed Chromium could not create the required Unix socket in this execution environment. The supported cloud browser independently blocked the loopback preview. No visual pass is claimed. - Hugging Face build, public hosting, GitHub push, CI, or deployment. The repository has the Gradio Spaces configuration but has not been published. - Real service APIs, real transactions, real user approval authentication, cryptographic source authenticity, or resistance to hostile code inside the Python process. A software license remains to be chosen before publication.