Text Generation
Transformers
Safetensors
GGUF
Laya
English
qwen3_5_text
decision-model
typed-decisions
calibration
calibrated-probabilities
classification
tool-selection
tool-use
agent-routing
clarification
robustness
decision-index
jevbench
jev
jev-compatible
open-jev
typesafe-compatible
systemone
kev
wald
wald-q4b
qwen3.5
4b
vllm
llama.cpp
ollama
reasoning
conversational
Eval Results (legacy)
Instructions to use org2ai/Wald-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use org2ai/Wald-4B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="org2ai/Wald-4B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("org2ai/Wald-4B") model = AutoModelForCausalLM.from_pretrained("org2ai/Wald-4B", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Laya
How to use org2ai/Wald-4B with Laya:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use org2ai/Wald-4B with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf org2ai/Wald-4B:Q4_K_M # Run inference directly in the terminal: llama cli -hf org2ai/Wald-4B:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf org2ai/Wald-4B:Q4_K_M # Run inference directly in the terminal: llama cli -hf org2ai/Wald-4B:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf org2ai/Wald-4B:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf org2ai/Wald-4B:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf org2ai/Wald-4B:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf org2ai/Wald-4B:Q4_K_M
Use Docker
docker model run hf.co/org2ai/Wald-4B:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use org2ai/Wald-4B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "org2ai/Wald-4B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "org2ai/Wald-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/org2ai/Wald-4B:Q4_K_M
- SGLang
How to use org2ai/Wald-4B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "org2ai/Wald-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "org2ai/Wald-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "org2ai/Wald-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "org2ai/Wald-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Ollama
How to use org2ai/Wald-4B with Ollama:
ollama run hf.co/org2ai/Wald-4B:Q4_K_M
- Unsloth Desktop
- Pi
How to use org2ai/Wald-4B with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf org2ai/Wald-4B:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "org2ai/Wald-4B:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use org2ai/Wald-4B with Docker Model Runner:
docker model run hf.co/org2ai/Wald-4B:Q4_K_M
- Lemonade
How to use org2ai/Wald-4B with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull org2ai/Wald-4B:Q4_K_M
Run and chat with the model
lemonade run user.Wald-4B-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use org2ai/Wald-4B with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf org2ai/Wald-4B:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default org2ai/Wald-4B:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use org2ai/Wald-4B with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf org2ai/Wald-4B:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "org2ai/Wald-4B:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
File size: 7,501 Bytes
503dfd6 78ebf43 503dfd6 78ebf43 503dfd6 78ebf43 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 | > Historical v1.0 documentation. Use HF revision `v1.0-legacy` and each adapter's recorded parent. These results and adapters are not validated on the current 22D0-f7 default.
# Fine-tune Wald-4B on your task (CLI, releasing soon)
> **Preview.** The CLI is not public yet; we plan to release it soon. The commands below are preview syntax and may
> change before the release.
**One dataset in, one task model out, with an honest comparison to Jev.** The `vertical` CLI takes a labelled dataset (a
Hugging Face repo, a URL, or your own file) and trains a LoRA on Wald-4B for that one task. It then reads the task's
fixed test items with your LoRA, with Jev and with zero-shot baselines, using byte-identical requests, and writes a report
with paired confidence intervals and calibration. The verticals in the [README](../README.md#verticals-a-quick-lora-per-task)
were all made this way. One LoRA costs $0.12–$1.81 of GPU time (under 2 GPU-hours).

## The pipeline
Free stages run on your machine. Paid stages run on a GPU backend you choose: any SSH GPU box, RunPod or Modal. Every
stage is resumable. Reads are cached per item, a finished LoRA is not trained again, and each run keeps a manifest with
the sha256 of every output.
| stage | command | what it does |
|---|---|---|
| task spec | `init` · `inspect` · `validate` | Scaffold `task.yaml` from a dataset: the source pinned to a commit, fields, labels, question, metric and split sizes. `inspect` renders example records as the model will see them. |
| prep | `prep` | Download the data and hash every file. Render each item once, both as a decision record and as the exact request Jev receives. Make a seeded, label-stratified, group-aware split into train · calib · dev · test, and dedupe across splits. |
| overlap check | `prep` | Compare held-out items with the train pool, with Decision Index items and with Wald-4B's training data. The counts go in the data manifest. |
| audit | `audit` · `approve` | Free, deterministic rule checks before and after every stage, plus a review of a sample by a cheap model. A FAIL blocks the next paid stage. |
| augment (optional) | `augment` | Extra training rows: soft labels for unlabelled texts, synthetic items, or paraphrases. The teacher is your own coding agent or any OpenAI-compatible API. Every row is re-checked on ingest (schema, labels, dedupe, overlap with the held-out splits). |
| baseline | `baseline` | Read calib, dev and test with Jev (cached per request, so re-runs are free) and with zero-shot Wald-4B and the raw base. Optionally add frontier LLMs, plus random and class-prior floors. |
| train | `train` · `watch` | One LoRA per arm. An arm is a training size, for example 300 labels, 1,000 labels, or all of them. The recipe is rank 32 on every projection, lr 1e-4, 2 epochs, and KL replay toward the base's own answers, so the base keeps its other skills. `watch` tails the log and stops a bad run. |
| eval | `eval` | One vLLM server holds Wald-4B plus every adapter. Reads use the serving reader: letter readout, knockout above 26 options, and optional thinking effort. |
| calibrate | `calibrate` | Fit one temperature per system on the calib split, never on test. It changes confidence, never the answer. Jev is scored as served. |
| report | `report` · `compare` | `report.md` + `report.json`: the test metric with 2,000 paired bootstrap resamples against Jev and the base, ECE, high-confidence errors, and dev numbers for picking a configuration. `compare` gives paired differences between two runs, for example Wald-4B vs the raw base. |
| deploy | `serve` · `try` · `latency` | Serve the adapter and its temperature on Wald-4B behind the same `/v1/systemone` decision API. The adapter stays a separate file, and Wald-4B itself is never changed. `try` sends one item through Jev and your model side by side, and `latency` measures serial latency on an otherwise idle card. |
## Guards
A task model is only useful if its numbers are real. The CLI enforces these rules itself:
- **Same bytes everywhere.** Each item is rendered once, into one request. Jev receives that request, our reader reads
it, and your LoRA trains on the same bytes.
- **Leak gate.** Before training, any test item whose word 5-gram Jaccard similarity with a train or calib item is 0.8 or
higher stops the run. You can drop those items from this run (`--drop-near-dups`), or a person can approve keeping them.
- **Audit gate.** A FAIL (exit code 4) blocks every paid stage. Only a person can override it: `vertical approve` asks
y/N on a real terminal, or you click *Approve override* in the UI. The approval is signed, and it expires when the
data or the audit result changes. An agent cannot sign it.
- **Held-out discipline.** Temperatures are fitted on calib. Configurations are chosen on dev. Test is reported once per
configuration, and a best-of-N read on test is labelled as best-of-N.
- **Cost gate.** Paid stages print a cost estimate and run only with `--yes`, or when the estimate fits under `--max-usd`.
`--dry-run` prints the plan and changes nothing.
- **Training watch.** A NaN or infinite loss, a zero learning rate, the wrong base, or a cost above twice the estimate
stops the job.
- **Secrets.** Keys are passed at run time and never reach the GPU box's disk or the logs.
## Built for coding agents
Every command takes `--json` (one JSON object on stdout, with a `next` hint) and has fixed exit codes: 0 ok, 2 user or
config error, 3 remote failure, 4 audit FAIL. Spending money and overriding an audit still need a person's yes.
## Cost and time
- **One LoRA:** $0.12–$1.81 of GPU time on one RTX PRO 6000 at about $1 per hour (under 2 GPU-hours). Training runs at
about 5,000 tokens/s.
- **A full run** has cost $0.25–$2.15 on our public tasks. That covers the baseline reads, one to three training sizes,
eval and the report. Jev reads are billed by Jev and cached, so re-runs do not pay for them again.
## Example (preview syntax)
A task is one `task.yaml`:
```yaml
name: banking77
title: Bank customer intent routing
source: {hf: legacy-datasets/banking77, revision: <commit>, license: CC-BY-4.0, splits: {train: train, test: test}}
fields: {text: text, label: label}
labels: {from_features: true}
state_template: "Customer message: {text}"
question: Which intent does this bank customer's message express?
metric: macro_f1
sizes: {calib: 300, test: 500}
arms: [n300, n1000, all]
```
```sh
vertical init banking77 --hf legacy-datasets/banking77 # scaffold task.yaml
vertical inspect banking77 # rendered records, labels, lengths
vertical validate banking77
vertical prep banking77 # download, split, dedupe, overlap checks
vertical audit banking77 --stage data
vertical run banking77 --dry-run # cost of every paid stage; nothing runs
vertical run banking77 --max-usd 3 # baseline → train → eval → calibrate → report
vertical report banking77 # report.md + report.json
vertical serve banking77 --arm all --yes # the adapter behind /v1/systemone
vertical try banking77 --arm all --text "My card still hasn't arrived"
```
Other commands: `list`, `status`, `stop`, `backends --check`, `compare`, `latency`, `note`, `import-read` (adds a read
made outside the CLI to a run's report).
|