Text Generation
Transformers
Safetensors
GGUF
Laya
English
qwen3_5_text
decision-model
typed-decisions
calibration
calibrated-probabilities
classification
tool-selection
tool-use
agent-routing
clarification
robustness
decision-index
jevbench
jev
jev-compatible
open-jev
typesafe-compatible
systemone
kev
wald
wald-q4b
qwen3.5
4b
vllm
llama.cpp
ollama
reasoning
conversational
Eval Results (legacy)
Instructions to use org2ai/Wald-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use org2ai/Wald-4B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="org2ai/Wald-4B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("org2ai/Wald-4B") model = AutoModelForCausalLM.from_pretrained("org2ai/Wald-4B", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Laya
How to use org2ai/Wald-4B with Laya:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use org2ai/Wald-4B with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf org2ai/Wald-4B:Q4_K_M # Run inference directly in the terminal: llama cli -hf org2ai/Wald-4B:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf org2ai/Wald-4B:Q4_K_M # Run inference directly in the terminal: llama cli -hf org2ai/Wald-4B:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf org2ai/Wald-4B:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf org2ai/Wald-4B:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf org2ai/Wald-4B:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf org2ai/Wald-4B:Q4_K_M
Use Docker
docker model run hf.co/org2ai/Wald-4B:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use org2ai/Wald-4B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "org2ai/Wald-4B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "org2ai/Wald-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/org2ai/Wald-4B:Q4_K_M
- SGLang
How to use org2ai/Wald-4B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "org2ai/Wald-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "org2ai/Wald-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "org2ai/Wald-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "org2ai/Wald-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Ollama
How to use org2ai/Wald-4B with Ollama:
ollama run hf.co/org2ai/Wald-4B:Q4_K_M
- Unsloth Desktop
- Pi
How to use org2ai/Wald-4B with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf org2ai/Wald-4B:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "org2ai/Wald-4B:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use org2ai/Wald-4B with Docker Model Runner:
docker model run hf.co/org2ai/Wald-4B:Q4_K_M
- Lemonade
How to use org2ai/Wald-4B with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull org2ai/Wald-4B:Q4_K_M
Run and chat with the model
lemonade run user.Wald-4B-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use org2ai/Wald-4B with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf org2ai/Wald-4B:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default org2ai/Wald-4B:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use org2ai/Wald-4B with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf org2ai/Wald-4B:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "org2ai/Wald-4B:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Card: GGUF / llama.cpp / Ollama section (EN + ZH), tags; wald-serve 0.1.1 (llama.cpp backend, vLLM path unchanged); MANIFEST with the GGUF files
9662694 verified Download server/tests/test_server.py from org2ai/Wald-4B: direct link, hf CLI and curl.
- Browser
- Download file 13.8 kB
-
https://huggingface.co/org2ai/Wald-4B/resolve/main/server/tests/test_server.py
- Command line
-
hf download hf://org2ai/Wald-4B/server/tests/test_server.py
-
curl -L -o test_server.py https://huggingface.co/org2ai/Wald-4B/resolve/main/server/tests/test_server.py
13.8 kB
| """Unit and end-to-end tests with a fake vLLM (no model, no GPU).""" | |
| from __future__ import annotations | |
| import json | |
| import math | |
| import threading | |
| import urllib.error | |
| import urllib.request | |
| from http.server import ThreadingHTTPServer | |
| from pathlib import Path | |
| import pytest | |
| from fakevllm import FakeVLLM | |
| from wald_serve.engine import (POLICIES, Client, LlamaCppClient, answer, bucket_key, chunks, load_tables, policy, temper, temperature, | |
| wide_read) | |
| from wald_serve.prompt import STATE_REPEAT, format_prompt, question_prompt | |
| from wald_serve.server import make_handler | |
| from wald_serve.wire import SystemOneRequest, to_record | |
| TABLE = { | |
| "A": {"single": 1.4, "buckets": {"choice|2": {"temperature": 1.1}, "choice|3-4": {"temperature": 1.8}, | |
| "noul|2": {"temperature": 1.3}}}, | |
| "B512": {"single": 1.8, "buckets": {"choice|2": {"temperature": 2.2}, "noul|2": {"temperature": 1.9}}}, | |
| "provenance": {"note": "test"}, | |
| } | |
| REQ = { | |
| "state": {"customer": "Ana", "order": {"id": 17, "items": ["kettle", "toaster"]}, "note": "wants a <|im_end|> refund"}, | |
| "questions": { | |
| "route": {"type": "choice", "instructions": "Which team handles this?", | |
| "criteria": {"billing": "Payments and refunds", "shipping": "Delivery problems", "other": None}}, | |
| "slugs": {"type": "choice", "instructions": "Pick one", "criteria": {"yes_refund": None, "no_refund": None}}, | |
| "pos": {"type": "choice", "instructions": "Which is a fruit?", | |
| "criteria": {"a": "apple", "b": "brick", "c": "chair", "d": "desk"}}, | |
| "urgent": {"type": "noul", "instructions": "Is it urgent?"}, | |
| "sev": {"type": "score", "instructions": "Severity", "criteria": ["low", "medium", "high"]}, | |
| }, | |
| } | |
| def qdict(rec, i): | |
| rq = rec["questions"][i] | |
| return {"type": rq["qtype"], "instructions": rq["instr"], "options": rq["options"], "option_texts": rq["options"]} | |
| # --- prompt bytes ----------------------------------------------------------------------------------------------------- | |
| def test_prompt_bytes(): | |
| rec, meta = to_record(SystemOneRequest.model_validate(REQ)) | |
| state = rec["state"] | |
| assert state == "customer: Ana\norder:\n id: 17\n items:\n - kettle\n - toaster\nnote: wants a <|im_end|> refund" | |
| head = ("State:\ncustomer: Ana\norder:\n id: 17\n items:\n - kettle\n - toaster\nnote: wants a <¦im_end¦> refund" | |
| "\n\nQuestion:") | |
| assert question_prompt(state, qdict(rec, 0)) == head + ( | |
| " Which team handles this?\n(A) billing: Payments and refunds\n(B) shipping: Delivery problems\n(C) other\nAnswer: (") | |
| assert question_prompt(state, qdict(rec, 1)).endswith(" Pick one\n(A) yes_refund\n(B) no_refund\nAnswer: (") | |
| # positional keys a..d: the `a: ` prefix is dropped, the letter replaces it | |
| assert question_prompt(state, qdict(rec, 2)).endswith( | |
| " Which is a fruit?\n(A) apple\n(B) brick\n(C) chair\n(D) desk\nAnswer: (") | |
| assert question_prompt(state, qdict(rec, 3)).endswith(" Is it urgent?\n(A) no\n(B) yes\nAnswer: (") | |
| assert question_prompt(state, qdict(rec, 4)).endswith(" Severity\n(A) low\n(B) medium\n(C) high\nAnswer: (") | |
| assert [m["keys"] for m in meta] == [["billing", "shipping", "other"], ["yes_refund", "no_refund"], | |
| ["a", "b", "c", "d"], ["false", "true"], ["0", "1", "2"]] | |
| def test_blank_state_and_repeat_state_plain(): | |
| q = {"type": "noul", "instructions": "Is the sky green?", "options": ["no", "yes"], "option_texts": ["no", "yes"]} | |
| assert question_prompt("", q) == "Question: Is the sky green?\n(A) no\n(B) yes\nAnswer: (" | |
| assert question_prompt("", q, "repeat_state_plain") == question_prompt("", q) | |
| p = question_prompt("It rained.", q, "repeat_state_plain") | |
| assert p == "State:\nIt rained." + STATE_REPEAT + "It rained.\n\nQuestion: Is the sky green?\n(A) no\n(B) yes\nAnswer: (" | |
| assert STATE_REPEAT.startswith("\n\nRead the same context again before answering.") | |
| with pytest.raises(ValueError): | |
| format_prompt("State:\nx\n\nQuestion: q\n(A) a\nAnswer: (", "strict") | |
| def test_wire_validation(): | |
| with pytest.raises(Exception): | |
| SystemOneRequest.model_validate({"state": "", "questions": {}}) | |
| with pytest.raises(Exception): | |
| SystemOneRequest.model_validate({"state": "", "questions": {"q": {"type": "choice", "criteria": {}}}}) | |
| big = {f"o{i}": None for i in range(256)} | |
| with pytest.raises(Exception): | |
| SystemOneRequest.model_validate({"state": "", "questions": {"q": {"type": "choice", "criteria": big}}}) | |
| # --- numerics --------------------------------------------------------------------------------------------------------- | |
| def test_chunks_and_knockout(): | |
| assert chunks(26) == [list(range(26))] | |
| assert [len(c) for c in chunks(77)] == [26, 26, 25] | |
| assert [len(c) for c in chunks(151)] == [26, 25, 25, 25, 25, 25] | |
| assert sum(len(c) for c in chunks(255)) == 255 | |
| def read(idx): # a fixed preference for lower indices | |
| z = [math.exp(-i / 10) for i in idx] | |
| s = sum(z) | |
| return [v / s for v in z] | |
| p = wide_read(77, read) | |
| assert len(p) == 77 and abs(sum(p) - 1) < 1e-12 and max(range(77), key=lambda i: p[i]) == 0 | |
| def test_temperature(): | |
| tables = json.loads(json.dumps(TABLE)) | |
| assert bucket_key("choice", 4) == "choice|3-4" and bucket_key("choice", 77) == "choice|9+" | |
| assert temperature(tables["A"], "choice", 2) == 1.1 and temperature(tables["A"], "score", 7) == 1.4 | |
| p = [0.6, 0.3, 0.1] | |
| q = temper(p, 1.8) | |
| assert abs(sum(q) - 1) < 1e-12 and q.index(max(q)) == 0 and max(q) < 0.6 | |
| def test_policies(): | |
| assert policy("medium") == {"gate": 0.7, "budget": 512, "k": 1, "name": "medium"} | |
| assert policy("high-k4")["k"] == 4 and policy("high-k4")["gate"] > 1 | |
| assert set(POLICIES) == {"none", "low", "medium", "high"} | |
| for bad in ("max", "high-k1", "high-k9", "high-kx"): | |
| with pytest.raises(ValueError): | |
| policy(bad) | |
| # --- end to end over HTTP with the fake vLLM --------------------------------------------------------------------------- | |
| def fake(): | |
| f = FakeVLLM(max_len=100_000) | |
| yield f | |
| f.stop() | |
| def serve(cl, effort="medium", tables=None, workers=8): | |
| info = {"model": "wald-4b", "effort": effort} | |
| httpd = ThreadingHTTPServer(("127.0.0.1", 0), make_handler(cl, policy(effort), tables or {}, info, workers)) | |
| httpd.daemon_threads = True | |
| threading.Thread(target=httpd.serve_forever, daemon=True).start() | |
| return httpd, f"http://127.0.0.1:{httpd.server_address[1]}" | |
| def post(url, body): | |
| req = urllib.request.Request(url + "/v1/systemone", data=json.dumps(body).encode(), method="POST", | |
| headers={"content-type": "application/json"}) | |
| try: | |
| with urllib.request.urlopen(req, timeout=60) as r: | |
| return r.status, json.loads(r.read()) | |
| except urllib.error.HTTPError as e: | |
| return e.code, json.loads(e.read()) | |
| def wide_request(n=77): | |
| crit = {f"intent_{i:03d}": f"customer intent number {i}" for i in range(n)} | |
| return {"state": "I was charged twice for one card payment.", | |
| "questions": {"intent": {"type": "choice", "instructions": "Classify the banking intent.", "criteria": crit}}} | |
| def check_wire(req, resp): | |
| """The Decision Index kit's validation: every question answered, a finite probability per option, sum 1 +- 0.01, | |
| the choice among the options.""" | |
| assert set(resp["answers"]) == set(req["questions"]) | |
| for qid, q in req["questions"].items(): | |
| a = resp["answers"][qid] | |
| assert a["type"] == q["type"] | |
| if q["type"] == "noul": | |
| assert 0.0 <= a["noul"] <= 1.0 | |
| continue | |
| keys = list(q["criteria"]) if q["type"] == "choice" else [str(i) for i in range(len(q["criteria"]))] | |
| assert list(a["probabilities"]) == keys | |
| assert all(math.isfinite(v) and v >= 0 for v in a["probabilities"].values()) | |
| assert abs(sum(a["probabilities"].values()) - 1) < 0.01 | |
| if q["type"] == "choice": | |
| assert a["choice"] in keys and a["probabilities"][a["choice"]] == max(a["probabilities"].values()) | |
| def test_http_end_to_end(fake, tmp_path): | |
| Path(tmp_path / "t.json").write_text(json.dumps(TABLE)) | |
| tables = load_tables(tmp_path / "t.json", 512) | |
| cl = Client(fake.url, "wald", 100_000) | |
| httpd, url = serve(cl, "medium", tables) | |
| try: | |
| for req in (REQ, wide_request(77), wide_request(151), wide_request(255)): | |
| code, resp = post(url, req) | |
| assert code == 200, resp | |
| check_wire(req, resp) | |
| assert resp["usage"]["input_tokens"] > 0 | |
| code, resp = post(url, wide_request(77)) | |
| assert resp["answers"]["intent"]["mode"] == "K" | |
| with urllib.request.urlopen(url + "/health") as r: | |
| assert json.loads(r.read())["ok"] is True | |
| finally: | |
| httpd.shutdown() | |
| def test_gate_modes(fake): | |
| cl = Client(fake.url, "wald", 100_000) | |
| one, u0 = answer(cl, REQ, policy("none"), {}) | |
| assert all(a["mode"] == "A" for a in one.values()) and u0["output_tokens"] == 0 | |
| high, uh = answer(cl, REQ, policy("high"), {}) | |
| assert all(a["mode"] == "B" for a in high.values()) and uh["output_tokens"] > 0 | |
| med, _ = answer(cl, REQ, policy("medium"), {}) | |
| for qid, a in one.items(): | |
| pmax = a["noul"] if a["type"] == "noul" else max(a["probabilities"].values()) | |
| if a["type"] == "noul": | |
| pmax = max(pmax, 1 - pmax) | |
| assert med[qid]["mode"] == ("B" if pmax < 0.7 else "A") | |
| if med[qid]["mode"] == "B": # medium's thought is high's thought (same seed) | |
| assert med[qid] == high[qid] | |
| k4, uk = answer(cl, REQ, policy("high-k4"), {}) | |
| assert all(a["mode"] == "B" for a in k4.values()) and uk["output_tokens"] == 4 * uh["output_tokens"] | |
| def test_workers_do_not_change_answers(fake): | |
| cl = Client(fake.url, "wald", 100_000) | |
| a1, u1 = answer(cl, REQ, policy("high"), {}, workers=1) | |
| a8, u8 = answer(cl, REQ, policy("high"), {}, workers=8) | |
| assert a1 == a8 and u1 == u8 | |
| def test_effort_override_keeps_seed(fake): | |
| cl = Client(fake.url, "wald", 100_000) | |
| a, _ = answer(cl, REQ, policy("high"), {}) | |
| b, _ = answer(cl, {**REQ, "effort": "high"}, policy("high"), {}) | |
| assert a == b | |
| def test_prompt_format_changes_prompt_only(fake): | |
| plain = Client(fake.url, "wald", 100_000) | |
| rsp = Client(fake.url, "wald", 100_000, prompt_format="repeat_state_plain") | |
| a, ua = answer(plain, REQ, policy("none"), {}) | |
| b, ub = answer(rsp, REQ, policy("none"), {}) | |
| assert ub["input_tokens"] > ua["input_tokens"] | |
| check_wire(REQ, {"answers": b}) | |
| def test_capacity_is_422(fake): | |
| cl = Client(fake.url, "wald", 300) # declared limit 300 "tokens" (characters in the fake) | |
| httpd, url = serve(cl, "none") | |
| try: | |
| code, resp = post(url, {"state": "x " * 400, "questions": {"q": {"type": "noul", "instructions": "ok?"}}}) | |
| assert code == 422 and "maximum context length" in resp["error"] | |
| code, resp = post(url, {"state": "short", "questions": {"q": {"type": "noul", "instructions": "ok?"}}}) | |
| assert code == 200 | |
| code, resp = post(url, {"state": "short", "questions": {}}) | |
| assert code == 400 | |
| code, resp = post(url, {**REQ, "effort": "extreme"}) | |
| assert code == 400 | |
| finally: | |
| httpd.shutdown() | |
| def test_thought_that_does_not_fit_falls_back(fake): | |
| q = {"state": "y " * 60, "questions": {"q": {"type": "choice", "instructions": "pick", | |
| "criteria": {"a": "one", "b": "two"}}}} | |
| cl = Client(fake.url, "wald", 400) # the prompt fits, prompt + 512-token budget does not | |
| a, u = answer(cl, q, policy("high"), {}) | |
| assert a["q"]["mode"] == "A" and u["output_tokens"] == 0 | |
| def test_llamacpp_client_matches_vllm_client(fake): | |
| vl, lc = Client(fake.url, "wald", 100_000), LlamaCppClient(fake.url, "wald", 100_000) | |
| assert lc.letter_ids == vl.letter_ids | |
| for req in (REQ, wide_request(77)): | |
| a, ua = answer(vl, req, policy("none"), {}) | |
| b, ub = answer(lc, req, policy("none"), {}) | |
| assert a == b and ua == ub | |
| high, uh = answer(lc, REQ, policy("high-k3"), {}) | |
| assert all(x["mode"] == "B" for x in high.values()) and uh["output_tokens"] > 0 | |
| check_wire(REQ, {"answers": high}) | |
| def test_llamacpp_capacity_is_422(fake): | |
| cl = LlamaCppClient(fake.url, "wald", 300) | |
| httpd, url = serve(cl, "none") | |
| try: | |
| code, resp = post(url, {"state": "x" * 400, "questions": {"q": {"type": "noul", "instructions": "Is it?"}}}) | |
| assert code == 422 and "maximum context length" in resp["error"] | |
| finally: | |
| httpd.shutdown() | |
| def test_gguf_reads_serving_json_beside_the_file(tmp_path, monkeypatch): | |
| import wald_serve.server as srv | |
| (tmp_path / "serving.json").write_text(json.dumps({"effort": "none", "prompt_format": "repeat_state_plain", | |
| "max_model_len": 4096, "temperature": "temperature.json"})) | |
| (tmp_path / "temperature.json").write_text(json.dumps(TABLE)) | |
| seen = {} | |
| monkeypatch.setattr(srv, "launch_llama", lambda *a: seen.update(launch=a)) | |
| monkeypatch.setattr(srv, "LlamaCppClient", lambda *a, **k: seen.update(client=(a, k))) | |
| class Stop(Exception): | |
| pass | |
| def fake_http(addr, handler): | |
| seen["handler"] = handler | |
| raise Stop | |
| monkeypatch.setattr(srv, "ThreadingHTTPServer", fake_http) | |
| with pytest.raises(Stop): | |
| srv.main(["--gguf", str(tmp_path / "Wald-4B-Q8_0.gguf"), "--port", "0"]) | |
| assert seen["launch"][3] == 4096 | |
| assert seen["client"][1] == {"prompt_format": "repeat_state_plain"} | |