vera-core

vera-core

vera-core is a non-autoregressive System 1 decision model. Give it a state (text, email, ticket, or JSON) and typed questions; it returns typed answers with calibrated probabilities in a single forward pass, over an 8K (8,192-token) context. It never generates text, so there is nothing to parse and nothing to hallucinate. It speaks the same wire format as Jev (POST /v1/systemone).

Base model

vera-core is derived from Qwen3.5-2B (Apache-2.0). We take its text backbone, add a purpose-built decision head that scores the options of a question in one forward pass (the model reads the question, the option menu and the state, and no text is generated), and fully fine-tune the whole stack on decision data.

vera-core is a generalist base: fine-tune it for your use case

Fine-tuning on your own decisions will yield better results.

The released checkpoint works zero-shot across many kinds of decisions, but it is built to be a starting point. Fine-tune it on decisions from your own domain and it will fit your labels, your wording and your calibration far better than the generalist base.

vera-core is the larger sibling of vera-spark. See vera-core and vera-spark.

Use cases

vera-core is a fast decision layer for anything that can be phrased as "given this state, pick from these options". Fine-tune it on your own labels for:

  • Routing and triage: send a ticket, email or request to the right team, queue or priority.
  • Intent and topic classification: map messages to your own intent set, with calibrated confidence for escalation.
  • Tool and function selection: choose which tool or API an agent should call next.
  • Policy and compliance checks: decide whether a case follows a rule or breaks it.
  • Scoring and severity: rate on an ordinal scale (urgency, sentiment, risk) with a full probability distribution.
  • Fact and claim verification: check a statement against a document.
  • Judging and ranking: compare candidate responses and pick the better one.
  • Confidence-gated workflows: act automatically when confident, hand off to a slower model or a human when not.

Benchmarks

JevBench (accuracy %, higher is better; ECE lower is better)

model params easy original hard overall ECE (overall)
vera-core (this repo) ~2B 100 93.1 58.6 77.9 0.098
vera-spark ~630M 100 83.3 46.0 68.8 0.141
Jev (reference) – 100 99 74.1 – –
Laya (public leaderboard, not our run) 421M 94 73 34 – –
Julia-1 (our run) 144M 75.0 47.2 35.1 47.2 –
reflex-s1 (our run, fast profile) 23M/83M 81.2 45.8 39.6 50.2 –

ECE for vera-core: 0.098 overall (all 231 items pooled, 10 bins). Brier: 0.007 / 0.089 / 0.589. Median latency per question on a GPU (MI300X, batch of one): 41-44 ms.

Per-family results (correct / total):

split family score family score
easy extraction 12/12 intent 12/12
easy fact 12/12 tool_selection 12/12
original extraction 12/12 ordinal 12/12
original intent 12/12 routing 12/12
original policy 10/12 adequacy 9/12
hard adversarial 6/6 routing_hard 5/5
hard multi_hop 12/18 judge_hard 10/17
hard trap 5/8 probability 5/10
hard tradeoff 4/6 long_policy 8/19
hard ambiguous 4/7 temporal_numeric 6/15

Held-out decisions (3.6k items, decontaminated): 81.4.

Limitations

  • The released checkpoint is not tuned for browser or computer-use agents. For those, fine-tune it on traces from your own agent first.

vera-core and vera-spark

The two models share the same System One interface, wire format and decision format, so you can swap one for the other.

  • vera-core (this repo): the larger model (~2B parameters). Pick it when accuracy matters most, especially on hard, multi-step decisions.
  • vera-spark: the compact model (~630M parameters). Pick it when you want the smallest footprint.

Both are generalist bases meant to be fine-tuned for your use case.

Quickstart

vera-core uses the same vera package (vera-s1) as vera-spark:

pip install vera-s1

Python 3.10 or newer. The weights (3.8 GB) are downloaded once, on first use, and cached. A GPU is recommended: on CPU a request takes several seconds. For the fastest GPU path, use pip install "vera-s1[fast]".

import vera

agent = vera.load("scar-ai/vera-core")

state = "Hi, we were billed twice for March. Please refund the duplicate today or we will cancel our plan."
questions = {
    "department": {"type": "choice", "instructions": "Which department should handle this?",
                   "criteria": {"billing": "invoices, payments, refunds",
                                "technical": "bugs, outages, system errors",
                                "other": "everything else"}},
    "urgency": {"type": "score", "instructions": "How urgent is this?",
                "criteria": ["not urgent", "soon", "blocking"]},
    "churn_risk": {"type": "noul", "instructions": "Does the user threaten to cancel or leave?"},
}

result = agent.predict(state, questions)
print(result["answers"]["department"]["choice"])       # billing
print(result["answers"]["department"]["probabilities"])
print(result["answers"]["churn_risk"]["noul"])         # probability the answer is yes

All questions in a call are answered together. The state can be plain text or any JSON-serialisable object. Question types: choice (2 to 255 options), score (2 to 10 ordered levels, returns an expected score and the full distribution) and noul (yes/no probability).

Weights: bf16 (default) and fp32

main holds the weights in bf16 (3.8 GB). The full-precision fp32 weights (7.6 GB) are on the fp32 branch:

agent = vera.load("scar-ai/vera-core", revision="fp32")                                        # vera package
model = AutoModel.from_pretrained("scar-ai/vera-core", revision="fp32", trust_remote_code=True)  # transformers

For vera-serve, download the branch first and point --model at the folder:

hf download scar-ai/vera-core --revision fp32 --include "model.pt" --include "*.json" --include "*.jinja" --local-dir vera-core-fp32
vera-serve --model ./vera-core-fp32 --port 8200

Self-hosting: Jev-compatible HTTP server

vera-serve exposes the model on the same POST /v1/systemone request and response shape as Jev, so existing clients work by changing their base URL:

vera-serve --model scar-ai/vera-core --port 8200
curl -s localhost:8200/v1/systemone -H 'Content-Type: application/json' -d '{
  "state": {"document": "I was charged twice. Please fix this ASAP."},
  "questions": {"billing": {"type": "noul", "instructions": "Is this ticket about billing?"}}
}'

It binds 127.0.0.1 by default (use --host 0.0.0.0 to expose it) and has no authentication, so put it behind your own gateway if you expose it. Concurrent requests are batched.

Use with transformers

The checkpoint also loads through transformers with AutoModel (custom code, so trust_remote_code=True). Tested with transformers 5.18.0 (vera-core needs a transformers release that includes Qwen3.5); outputs match the vera package exactly (max probability difference 0.0 on a three-question test).

from transformers import AutoModel

model = AutoModel.from_pretrained("scar-ai/vera-core", trust_remote_code=True)   # add .to("cuda") if you have a GPU
state = "Hi, we were billed twice for March. Please refund the duplicate today."
questions = {"department": {"type": "choice", "instructions": "Which department should handle this?",
                            "criteria": {"billing": "invoices, payments, refunds", "other": "everything else"}}}
print(model.predict(state, questions)["answers"]["department"])    # same shape as POST /v1/systemone

For tensor-level use, model.encode_batch(state, questions) builds the inputs and model(**batch).logits returns the calibrated option scores (softmax(logits, -1) gives probabilities). The model is non-autoregressive and has no generate().

Inference engine and decision API status

vera answers decisions through the Jev-compatible POST /v1/systemone API, not through text generation. That decides which engines can serve it.

engine / API status notes
vera-serve (POST /v1/systemone, /api/v1/decide) Supported The reference decision API server; batches concurrent requests.
vera.load(...).predict(...) (Python) Supported Same result shape as the HTTP API.
transformers (AutoModel, trust_remote_code=True) Supported See above. predict() mirrors the decision API.
vLLM Supported via engines/ adapter See Run with vLLM or SGLang.
SGLang Supported via engines/ adapter See Run with vLLM or SGLang.
Ollama (and llama.cpp / GGUF) Not supported There is no GGUF architecture for the decision head, and Ollama serves text generation only.

An OpenAI-style chat or completions client will not work with vera-core on any engine: the model scores options, it does not generate text.

Run with vLLM or SGLang

vllm serve / sglang.launch_server on this repo do not work directly. Use the adapter in engines/: the engine runs the backbone and the vera head runs on top. A GPU is required.

pip install vera-s1
pip install vllm            # or: pip install "sglang[all]"
hf download scar-ai/vera-core --include "engines/*" --local-dir .
python engines/export_backbone.py --model scar-ai/vera-core --out ./vera-core-backbone
import sys; sys.path.insert(0, "engines")
from vera_engine import from_vllm   # or from_sglang (call it inside `if __name__ == "__main__":`)

agent = from_vllm("scar-ai/vera-core", "./vera-core-backbone")
print(agent.predict("I was charged twice. Please fix this ASAP.",
                    {"billing": {"type": "noul", "instructions": "Is this ticket about billing?"}})["answers"])

Tested on an MI300X with vLLM 0.29.0 and SGLang 0.5.21: same answers as the vera package.

Files

file description
model.pt weights (state dict, bf16), used by the vera package
config.json, model.safetensors* transformers-format config and weights (bf16) (AutoModel, trust_remote_code=True)
configuration_vera.py, modeling_vera.py, vera_model.py, vera_schema.py custom code for the transformers loader
dist/vera_s1-0.1.2-py3-none-any.whl the vera-s1 Python package, also on PyPI (source: vera-spark/package)
engines/export_backbone.py, engines/vera_engine.py vLLM / SGLang adapter (see Run with vLLM or SGLang)
arch.json architecture flags for the loader
tokenizer.json, tokenizer_config.json, chat_template.jinja tokenizer and prompt template
Downloads last month
61
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Evaluation results