Kodiak large — research preview v2

Research preview. The quality tier of Kodiak, shared while we build in public. It's more accurate than the small model, especially on tasks it has never seen, and less reliable at saying "I can't tell" (see Known weaknesses). Code, docs and evaluation: https://github.com/grizzlypeaksoftware/kodiak · Fast tier: kodiak-small-v2-preview · Feedback and failure cases welcome as GitHub issues.

Kodiak is an encoder-only "System One" decision model by Cortex Agent LLC: a state (text, list of texts, or JSON) plus typed questions in, calibrated answers out, in one forward pass. Choice answers are always one of your labels, scores stay inside your range, and every question can come back as "not answerable from this state" with its own probability.

# pip install "kodiak-s1[infer] @ git+https://github.com/grizzlypeaksoftware/kodiak"
from kodiak_s1.hub import Kodiak

kodiak = Kodiak.from_pretrained("cortex-agent-llc/kodiak-large-v2-preview")   # GPU if available, else CPU
kodiak.decide("Hi, I ordered the walnut desk two weeks ago. Tracking has said 'label created' for 10 days. "
              "If it can't arrive by Monday, please cancel and refund me.",
              [{"type": "choice", "id": "intent", "text": "What does the customer want?",
                "labels": ["delivery status or expedite", "cancel and refund", "product question"]}])

Large vs. small (eval set v0.2, averaged over three training runs each)

small v2 (152M) large v2 (~400M)
Never-seen tasks (12 tasks), forced accuracy, choice questions 55.3% 60.9%
Best open zero-shot classifier on the same questions 57.9% 57.9%
Familiar tasks 81.7% 85.5%
Occupation from a biography (never seen) 63.3% 77.2%
When it abstains, how often it's right 91% 84%
Latency, GPU / 8 ARM CPU cores 8 ms / ~80 ms 16 ms / ~250 ms

The large model beats every open zero-shot classifier we tested on never-seen tasks (release criterion 1; decision D33). Details: STRATEGY.md §4.

This checkpoint

  • Run b-base-s1-v2-s1 (seed 1 of three, chosen by validation loss, not the eval set), step 6000. ModernBERT-large backbone (Apache-2.0). Calibration temperatures baked in; default abstain threshold from calibration.json (override per request with null_threshold).
  • Eval v0.2: overall 67.2%, familiar tasks 85.4%, never-seen tasks 56.8% (61.0% forced), ECE 0.088. (Eval v0.1: overall 81.8%, ECE 0.049.)
  • Includes handler.py for Hugging Face Inference Endpoints.
  • Includes model.onnx (fp32, calibration baked in) for self-hosting without Python: the Node.js + Docker server in server/ answers in ~80 ms per request on 8 CPU threads. Download this repo into a folder and point KODIAK_MODEL at it.

Known weaknesses (read before using)

  • Abstention is weaker than the small model's. On a message that never names a carrier, it answered "FedEx" (0.46) where the small model abstains. If a trustworthy "I can't tell" matters more than accuracy, use the small model, or require a higher min_confidence.
  • Confidence is less reliable on unfamiliar inputs (calibration error ~0.13 on never-seen tasks; true of both sizes). A long-horizon test (text adventures) showed Kodiak can be confidently wrong on inputs unlike its training data; escalate low-confidence answers.
  • Judgment scores on new scales are weak for both sizes (e.g. it rates "I was charged twice and nobody answers!" around 4/10 for urgency).
  • Some intent calls are wrong where the small model is right (e.g. "box arrived crushed, lamp broken" → "delivery status").
  • States longer than 512 tokens are truncated; English only.
  • Not for high-stakes decisions about people without human review. Bias in Bios (occupation) is a held-out task with known gender bias.
Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
0.4B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for cortex-agent-llc/kodiak-large-v2-preview

Quantized
(23)
this model