laya-cli-gate / README.md
uzuw's picture
Upload README.md with huggingface_hub
f846bb5 verified
|
Raw History Blame Contribute Delete
2.61 kB
---
license: apache-2.0
library_name: laya
tags:
- text-classification
- routing
- guardrails
- terminal
- cli
- safetensors
---
# laya-cli-gate
Head-only finetune of [`convaiinnovations/laya`](https://huggingface.co/convaiinnovations/laya)
(Apache-2.0, © Convai Innovations; encoder `answerdotai/ModernBERT-large`) as a
**pre-execution safety gate + tool router for terminal commands**.
```python
from laya import Agent
agent = Agent("uzuw/laya-cli-gate")
d = agent.system_one("git push origin main", {
"tool": {"type": "choice", "instructions": "Which tool owns this command?",
"criteria": {"git": "git", "docker": "docker", "node": "node",
"k8s": "k8s", "shell": "shell"}},
"safe": {"type": "noul", "instructions": "Is this command safe to run without human review?"}})["answers"]
# tool -> git (1.00), safe -> 0.13 -> ask the human
```
## Results (train = 65 commands, see GitHub for data + recipe)
| | base `convaiinnovations/laya` | this checkpoint |
|---|---|---|
| tool acc | 0.68 | **0.985** |
| safe acc @0.5 | 0.42 | **1.000** |
| safe acc @0.7 | ~0.47 | **1.000** |
| ECE | 0.339 | **0.029** |
Recipe: head-only (encoder frozen), Adam lr 2e-4, cosine, grad-clip, best-on-combined
checkpointing. Full log: [uzuw/laya-cli-gate](https://github.com/uzuw/laya-cli-gate).
## Honest limits
- Train-set eval is a ceiling. Holdout on 31 unseen commands: tool ~1.0, safe ~0.7.
- Known FP: `npm unpublish --force` reads safe; novel read-only commands can read unsafe (fail-closed).
- Good **gate** (human reviews `safe:false`), not an auto-approver.
- The checkpoint ships one out-of-range upstream temperature (`choice:11+`); `laya`
clamps it with a warning — treat that bucket as uncalibrated.
## v2 — round 4 (subfolder `v2/`, 184 commands balanced 92 safe / 92 unsafe)
Root files above are unchanged (v1). Load v2 with:
```python
from laya import Agent
agent = Agent("uzuw/laya-cli-gate", subfolder="v2")
```
| | v1 (root, 65 cmds) | v2 (`v2/`, 184 cmds) |
|---|---|---|
| train tool acc | 0.985 | **0.995** |
| train safe acc @0.5 | 1.000 | **1.000** (@0.4–0.8) |
| train ECE | 0.029 | **0.015** |
| unseen safe (model only) | 0.75 | **10/12** |
| unseen safe (with deny net in `gate.py`) | — | **12/12, 0 FP** |
v2 adds template-swept unsafe rows (new pkgs/paths/flags per dangerous verb) and ships
alongside a deterministic deny net (`gate.py` in the code repo) that forces `safe: false`
on known-dangerous verbs — `--dry-run` exempts. Same honest limits apply: use as a
**gate** (human reviews `safe:false`), not an auto-approver.