HumanPermit Neural Gate โ research preview
A small experimental policy head for a frozen Qwen2.5-0.5B-Instruct backbone. It evaluates a proposed action against operator rules and predicts PROCEED, WAIT or DENY. This repository contains custom head weights, not a standalone language model, a PEFT adapter or a production authorization system.
Source and attribution: sdellava/humanpermit-ai.
Reference code revision: c079751b0c5b22e185eb820fca865eec7f7b5d74.
Architecture and intended use
Five separately trained probes read internal representations of candidate actions, evidence and individual rules. They predict applicability, prohibition, approval requirement, missing evidence and approval coverage. Fixed thresholds (0.5) and programmed precedence combine predictions. The transformer backbone is unchanged. Policies contain one to six nonempty rule lines.
Intended for local research into policy-conditioned action gating. Actions are supplied by a user or script; this adapter does not plan tasks, generate tool calls, control machinery or assess arbitrary generated text. A host implements approval and pause/resume. WAIT does not suspend a transformer layer until a human responds.
Measured results and failures
| Evaluation | Result |
|---|---|
| Repaired staged diagnostic | 72/72 on already-observed cases; not independent validation |
| Cross-domain challenge | 25/35 correct initial decisions; 15/35 correct complete flows |
| Forbidden release in that challenge | 1, a payment explicitly forbidden by a paraphrased rule |
| Fresh approval flows | All 10 remained WAIT after valid consent |
| Follow-up stale autonomous actions | Unsafe PROCEED in finance, infrastructure and industry |
Results are synthetic, narrow and not proof of general-purpose safety. Rule wording and new domains cause failures. Scores are not calibrated probabilities of safety. The GitHub demo adds a deterministic host freshness constraint; that containment is not part of these learned weights and does not repair them.
Full methods and counterexamples and training diagnostics are preserved in the source repository. Training used synthetic policies and separate fit-family probes, followed by a fit-only coverage repair after observing an integration failure. No new cross-domain training was performed.
Use locally
Use Python 3.11+ and install an appropriate PyTorch build for your device. Install the reference implementation:
pip install "humanpermit-prototype[research] @ git+https://github.com/sdellava/humanpermit-ai.git@c079751b0c5b22e185eb820fca865eec7f7b5d74"
python run_example.py
The included script downloads the pinned base model and this adapter on the first
run. Later inference loads local files. CPU is the portable default; use
--placement cuda-stream for GPU inference with RAM offload, or --placement cuda
when sufficient GPU memory is available. Offload may be slower. After the first
successful download, --offline disallows Hub lookups and uses cached files.
Loading requires humanpermit.staged_layer.StagedModel. Do not load these weights
using AutoModel.from_pretrained() or a PEFT adapter loader. Automatic hosted
inference is not configured. No arbitrary Hub Python code is executed by the loader.
The example prints both raw model and effective host decisions. Human approval is simulated with the local demo identity, not production authentication. A new model evaluation follows approval; approval is not a forced EXECUTE.
Files and provenance
policy_head.safetensors: repaired staged head only; no Qwen backbone weights.adapter_config.json: base revision, dimensions, thresholds and training metadata.run_example.py: local loading and simulated approval example.LICENSE: MIT with source-repository attribution.SHA256SUMS: checksums of the supplied adapter files.
License
Project-authored artifacts use MIT. Retain the copyright and permission notices, including the repository URL in the copyright notice, in copies or substantial portions. Qwen and third-party dependencies retain their separate licenses. No safety warranty, certification or suitability for real authorization is claimed.