Wald-4B / docs /lora-cli.md
Harry19081's picture
Release Wald-Q4B22D0-f7: full DI0.2.1 54.59 and clean calibration
503dfd6 verified
|
Raw History Blame Contribute Delete
7.5 kB
> Historical v1.0 documentation. Use HF revision `v1.0-legacy` and each adapter's recorded parent. These results and adapters are not validated on the current 22D0-f7 default.
# Fine-tune Wald-4B on your task (CLI, releasing soon)
> **Preview.** The CLI is not public yet; we plan to release it soon. The commands below are preview syntax and may
> change before the release.
**One dataset in, one task model out, with an honest comparison to Jev.** The `vertical` CLI takes a labelled dataset (a
Hugging Face repo, a URL, or your own file) and trains a LoRA on Wald-4B for that one task. It then reads the task's
fixed test items with your LoRA, with Jev and with zero-shot baselines, using byte-identical requests, and writes a report
with paired confidence intervals and calibration. The verticals in the [README](../README.md#verticals-a-quick-lora-per-task)
were all made this way. One LoRA costs $0.12–$1.81 of GPU time (under 2 GPU-hours).
![Architecture of the vertical CLI](../figures/lora-cli-architecture.svg)
## The pipeline
Free stages run on your machine. Paid stages run on a GPU backend you choose: any SSH GPU box, RunPod or Modal. Every
stage is resumable. Reads are cached per item, a finished LoRA is not trained again, and each run keeps a manifest with
the sha256 of every output.
| stage | command | what it does |
|---|---|---|
| task spec | `init` · `inspect` · `validate` | Scaffold `task.yaml` from a dataset: the source pinned to a commit, fields, labels, question, metric and split sizes. `inspect` renders example records as the model will see them. |
| prep | `prep` | Download the data and hash every file. Render each item once, both as a decision record and as the exact request Jev receives. Make a seeded, label-stratified, group-aware split into train · calib · dev · test, and dedupe across splits. |
| overlap check | `prep` | Compare held-out items with the train pool, with Decision Index items and with Wald-4B's training data. The counts go in the data manifest. |
| audit | `audit` · `approve` | Free, deterministic rule checks before and after every stage, plus a review of a sample by a cheap model. A FAIL blocks the next paid stage. |
| augment (optional) | `augment` | Extra training rows: soft labels for unlabelled texts, synthetic items, or paraphrases. The teacher is your own coding agent or any OpenAI-compatible API. Every row is re-checked on ingest (schema, labels, dedupe, overlap with the held-out splits). |
| baseline | `baseline` | Read calib, dev and test with Jev (cached per request, so re-runs are free) and with zero-shot Wald-4B and the raw base. Optionally add frontier LLMs, plus random and class-prior floors. |
| train | `train` · `watch` | One LoRA per arm. An arm is a training size, for example 300 labels, 1,000 labels, or all of them. The recipe is rank 32 on every projection, lr 1e-4, 2 epochs, and KL replay toward the base's own answers, so the base keeps its other skills. `watch` tails the log and stops a bad run. |
| eval | `eval` | One vLLM server holds Wald-4B plus every adapter. Reads use the serving reader: letter readout, knockout above 26 options, and optional thinking effort. |
| calibrate | `calibrate` | Fit one temperature per system on the calib split, never on test. It changes confidence, never the answer. Jev is scored as served. |
| report | `report` · `compare` | `report.md` + `report.json`: the test metric with 2,000 paired bootstrap resamples against Jev and the base, ECE, high-confidence errors, and dev numbers for picking a configuration. `compare` gives paired differences between two runs, for example Wald-4B vs the raw base. |
| deploy | `serve` · `try` · `latency` | Serve the adapter and its temperature on Wald-4B behind the same `/v1/systemone` decision API. The adapter stays a separate file, and Wald-4B itself is never changed. `try` sends one item through Jev and your model side by side, and `latency` measures serial latency on an otherwise idle card. |
## Guards
A task model is only useful if its numbers are real. The CLI enforces these rules itself:
- **Same bytes everywhere.** Each item is rendered once, into one request. Jev receives that request, our reader reads
it, and your LoRA trains on the same bytes.
- **Leak gate.** Before training, any test item whose word 5-gram Jaccard similarity with a train or calib item is 0.8 or
higher stops the run. You can drop those items from this run (`--drop-near-dups`), or a person can approve keeping them.
- **Audit gate.** A FAIL (exit code 4) blocks every paid stage. Only a person can override it: `vertical approve` asks
y/N on a real terminal, or you click *Approve override* in the UI. The approval is signed, and it expires when the
data or the audit result changes. An agent cannot sign it.
- **Held-out discipline.** Temperatures are fitted on calib. Configurations are chosen on dev. Test is reported once per
configuration, and a best-of-N read on test is labelled as best-of-N.
- **Cost gate.** Paid stages print a cost estimate and run only with `--yes`, or when the estimate fits under `--max-usd`.
`--dry-run` prints the plan and changes nothing.
- **Training watch.** A NaN or infinite loss, a zero learning rate, the wrong base, or a cost above twice the estimate
stops the job.
- **Secrets.** Keys are passed at run time and never reach the GPU box's disk or the logs.
## Built for coding agents
Every command takes `--json` (one JSON object on stdout, with a `next` hint) and has fixed exit codes: 0 ok, 2 user or
config error, 3 remote failure, 4 audit FAIL. Spending money and overriding an audit still need a person's yes.
## Cost and time
- **One LoRA:** $0.12–$1.81 of GPU time on one RTX PRO 6000 at about $1 per hour (under 2 GPU-hours). Training runs at
about 5,000 tokens/s.
- **A full run** has cost $0.25–$2.15 on our public tasks. That covers the baseline reads, one to three training sizes,
eval and the report. Jev reads are billed by Jev and cached, so re-runs do not pay for them again.
## Example (preview syntax)
A task is one `task.yaml`:
```yaml
name: banking77
title: Bank customer intent routing
source: {hf: legacy-datasets/banking77, revision: <commit>, license: CC-BY-4.0, splits: {train: train, test: test}}
fields: {text: text, label: label}
labels: {from_features: true}
state_template: "Customer message: {text}"
question: Which intent does this bank customer's message express?
metric: macro_f1
sizes: {calib: 300, test: 500}
arms: [n300, n1000, all]
```
```sh
vertical init banking77 --hf legacy-datasets/banking77 # scaffold task.yaml
vertical inspect banking77 # rendered records, labels, lengths
vertical validate banking77
vertical prep banking77 # download, split, dedupe, overlap checks
vertical audit banking77 --stage data
vertical run banking77 --dry-run # cost of every paid stage; nothing runs
vertical run banking77 --max-usd 3 # baseline → train → eval → calibrate → report
vertical report banking77 # report.md + report.json
vertical serve banking77 --arm all --yes # the adapter behind /v1/systemone
vertical try banking77 --arm all --text "My card still hasn't arrived"
```
Other commands: `list`, `status`, `stop`, `backends --check`, `compare`, `latency`, `note`, `import-read` (adds a read
made outside the CLI to a run's report).