Kodiak large — research preview v2
Research preview. The quality tier of Kodiak, shared while we build in public. It's more accurate than the small model, especially on tasks it has never seen, and less reliable at saying "I can't tell" (see Known weaknesses). Code, docs and evaluation: https://github.com/grizzlypeaksoftware/kodiak · Fast tier: kodiak-small-v2-preview · Feedback and failure cases welcome as GitHub issues.
Kodiak is an encoder-only "System One" decision model by Cortex Agent LLC: a state (text, list of texts, or JSON) plus typed questions in, calibrated answers out, in one forward pass. Choice answers are always one of your labels, scores stay inside your range, and every question can come back as "not answerable from this state" with its own probability.
# pip install "kodiak-s1[infer] @ git+https://github.com/grizzlypeaksoftware/kodiak"
from kodiak_s1.hub import Kodiak
kodiak = Kodiak.from_pretrained("cortex-agent-llc/kodiak-large-v2-preview") # GPU if available, else CPU
kodiak.decide("Hi, I ordered the walnut desk two weeks ago. Tracking has said 'label created' for 10 days. "
"If it can't arrive by Monday, please cancel and refund me.",
[{"type": "choice", "id": "intent", "text": "What does the customer want?",
"labels": ["delivery status or expedite", "cancel and refund", "product question"]}])
Large vs. small (eval set v0.2, averaged over three training runs each)
| small v2 (152M) | large v2 (~400M) | |
|---|---|---|
| Never-seen tasks (12 tasks), forced accuracy, choice questions | 55.3% | 60.9% |
| Best open zero-shot classifier on the same questions | 57.9% | 57.9% |
| Familiar tasks | 81.7% | 85.5% |
| Occupation from a biography (never seen) | 63.3% | 77.2% |
| When it abstains, how often it's right | 91% | 84% |
| Latency, GPU / 8 ARM CPU cores | 8 ms / ~80 ms | 16 ms / ~250 ms |
The large model beats every open zero-shot classifier we tested on never-seen tasks (release criterion 1; decision D33). Details: STRATEGY.md §4.
This checkpoint
- Run
b-base-s1-v2-s1(seed 1 of three, chosen by validation loss, not the eval set), step 6000. ModernBERT-large backbone (Apache-2.0). Calibration temperatures baked in; default abstain threshold fromcalibration.json(override per request withnull_threshold). - Eval v0.2: overall 67.2%, familiar tasks 85.4%, never-seen tasks 56.8% (61.0% forced), ECE 0.088. (Eval v0.1: overall 81.8%, ECE 0.049.)
- Includes
handler.pyfor Hugging Face Inference Endpoints. - Includes
model.onnx(fp32, calibration baked in) for self-hosting without Python: the Node.js + Docker server inserver/answers in ~80 ms per request on 8 CPU threads. Download this repo into a folder and pointKODIAK_MODELat it.
Known weaknesses (read before using)
- Abstention is weaker than the small model's. On a message that never names a carrier, it answered "FedEx" (0.46) where the small model
abstains. If a trustworthy "I can't tell" matters more than accuracy, use the small model, or require a higher
min_confidence. - Confidence is less reliable on unfamiliar inputs (calibration error ~0.13 on never-seen tasks; true of both sizes). A long-horizon test (text adventures) showed Kodiak can be confidently wrong on inputs unlike its training data; escalate low-confidence answers.
- Judgment scores on new scales are weak for both sizes (e.g. it rates "I was charged twice and nobody answers!" around 4/10 for urgency).
- Some intent calls are wrong where the small model is right (e.g. "box arrived crushed, lamp broken" → "delivery status").
- States longer than 512 tokens are truncated; English only.
- Not for high-stakes decisions about people without human review. Bias in Bios (occupation) is a held-out task with known gender bias.
Model tree for cortex-agent-llc/kodiak-large-v2-preview
Base model
answerdotai/ModernBERT-large