Kodiak large v2, accuracy mode (research preview)

The recommended large Kodiak. Three independently trained Kodiak large v2 models (ModernBERT-large, ~400M parameters each) answer every question, and their calibrated answers are averaged. Each model is overconfident in different places, so averaging cancels much of it: more accurate, better calibrated, and a more trustworthy "can't tell." The cost is about 3× the compute of one large model.

Kodiak is an open decision model by Cortex Agent LLC: a state (text, a list of texts, or JSON) and typed questions in; calibrated choice, score or "can't tell" answers out, in one forward pass per model. Code, docs and the full build log: https://github.com/grizzlypeaksoftware/kodiak

# pip install "kodiak-s1[infer] @ git+https://github.com/grizzlypeaksoftware/kodiak"
from kodiak_s1.hub import Kodiak

kodiak = Kodiak.from_pretrained("cortex-agent-llc/kodiak-large-v2-ensemble-preview")   # loads all three members
kodiak.decide(
    "Hi, I ordered the walnut desk two weeks ago. Tracking has said 'label created' for 10 days.",
    [{"type": "choice", "id": "intent", "text": "What does the customer want?",
      "labels": ["delivery status or expedite", "cancel and refund", "product question"]},
     {"type": "choice", "id": "carrier", "text": "Which carrier is shipping it?", "labels": ["UPS", "FedEx", "USPS"]}],
)
# -> intent: delivery status or expedite (0.81); carrier: can't tell (0.98)

Accuracy mode vs. one large model (eval set v0.2, choice questions)

One large model (mean of 3 runs) Accuracy mode
Never-seen tasks, forced accuracy 0.609 0.623
Familiar tasks 0.855 0.869
Calibration error (ECE), overall 0.087 0.059
Calibration error, never-seen tasks 0.128 0.098
When it says "can't tell", it's right 0.84 0.94
Latency, GPU / 8 CPU cores ~16 ms / ~80 ms ~48 ms / ~250 ms

The default abstain threshold (0.75) was chosen on validation data only, by a rule fixed before scoring: the most accurate threshold whose validation abstain precision is at least 90% (decision D45). Override per request with null_threshold.

Files

ensemble.json lists the members (m0, m1, m2: the three large v2 training runs, seeds 0-2; m1 is the same weights as kodiak-large-v2-preview) and the threshold. Each member folder is a normal Kodiak model folder with its own calibration.

Known weaknesses

  • World knowledge. On never-seen tasks that need outside knowledge (academic fields, legal holdings), an 8B LLM still leads by several points.
  • Ratings ("how urgent is this?") are weak; use the interval and treat scores as rough.
  • Wording traps. When a message repeats an option's exact words in a conditional ("if it can't arrive, cancel and refund me"), Kodiak can over-weight those words. Being fixed with simulator data in v0.2.
  • Research preview: validate on your own data; don't use for decisions about people without human review.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for cortex-agent-llc/kodiak-large-v2-ensemble-preview

Finetuned
(386)
this model