Kev: a 400M-parameter text decision model

Kev scores user-supplied candidates for a text input. It supports categorical choices, yes/no probabilities, and ordered rubric scores through a shared encoder and a single-output scoring head. It is fine-tuned from Ettin Encoder 400M using Halo.

Source code · Inference guide · Training guide

Model details

Property Value
Maintainer skundu42
Parameters 395,832,321 (approximately 400M)
Architecture ModernBertForSequenceClassification, 28 layers, hidden size 1,024
Head One scalar per candidate; softmax across the supplied candidates
Base revision 7662476d60abb071a5bd319c9f3074f3072c062d
Export Safetensors, FP32 weights
Supported decision length 1,024 tokens per encoded prompt/candidate pair
Candidate count 2–16
Calibration temperature 1.3734692667114208

The prompt includes the decision type, instructions, state, and full criteria list. Each candidate is encoded with that prompt and receives a scalar score. The runtime computes softmax(scores / temperature). The backbone's positional capacity does not override Kev's 1,024-token application limit. Overlength inputs are rejected without silent truncation.

Supported decisions

Type Criteria Result
choice Mapping of candidate names to descriptions Selected name and probabilities
noul Exactly false and true, or omitted for No/Yes Probability of true
score Ordered list of rubric levels Distribution and expected zero-based index

A five-level rubric produces a score between 0 and 4. For choice/score, the confidence summary is (K * max_probability - 1) / (K - 1); it is not a validated probability of correctness. Kev scores alternatives; it does not generate text.

Quickstart

Use a Linux CUDA GPU environment configured as described in the training guide. The project runtime expects Python 3.12, PyTorch 2.11.x and Transformers >=5.16.1,<5.17. The supplied loader requires CUDA and uses BF16 autocast when supported. Halo is used for training; inference uses the Kev wrapper and Transformers.

Clone the source and download the export. If repository access requires authentication, run hf auth login with an account that has access.

git clone https://github.com/skundu42/kev.git
cd kev
hf download skundu42/kev \
  --revision e62becd4fda8a69e9494c5e22ab60a4a1371ba69 \
  --local-dir ./models/kev \
  --include model.safetensors config.json kev_config.json tokenizer.json tokenizer_config.json

Run this Python example from the cloned source directory:

import json
from kev.inference import DecisionModel

model = DecisionModel.from_pretrained("./models/kev")
result = model.predict(
    state="I was charged twice for my subscription.",
    questions={
        "route": {
            "type": "choice",
            "instructions": "Choose the team that should handle this request.",
            "criteria": {
                "billing": "Payments, invoices, and refunds",
                "technical": "Bugs and product support",
                "sales": "Plans and purchasing",
            },
        },
        "refund_needed": {
            "type": "noul",
            "instructions": "Does this request describe a billing error that may need a refund?",
            "criteria": {"false": "No", "true": "Yes"},
        },
        "urgency": {
            "type": "score",
            "instructions": "Rate the urgency expressed by the customer.",
            "criteria": ["low", "medium", "high"],
        },
    },
)
print(json.dumps(result, indent=2))

DecisionModel.from_pretrained accepts a local export directory, not a Hub ID. A generic text-classification pipeline does not implement Kev's prompt formatting, candidate grouping, or temperature calibration; use the wrapper for the supported decision API.

For JSON CLI inference:

python -m kev.inference --model ./models/kev --input examples/request.json

See the HTTP API guide for authenticated serving and request limits.

Training and data

The published training configuration records one epoch, learning rate 2e-5, weight decay 0.01, warmup ratio 0.03, seed 42, per-device batch size 1, gradient accumulation 32, BF16 training, and gradient checkpointing. Training minimizes cross-entropy across candidate scores against the target distribution, including soft targets. Validation loss selects the checkpoint.

The prepared dataset is skundu42/kev-prepared at revision 5cea4539b6676edfb4fd1fb60878e9aee602b0cd. The published provenance records source revisions, filtering, sampling, exclusions, and retained counts.

Split Decisions
Training 393,549
Validation 16,947
Calibration 16,267
Test 17,476

Sources include:

Native labeled test partitions are retained. Native validation is split 50/50 into validation/calibration when a labeled test partition exists, otherwise 50/25/25 into validation/calibration/test. Normalized content and related groups are deduplicated across partitions, with holdouts taking priority. Exact/group deduplication does not detect paraphrases or establish absence of pretraining overlap. Source and task caps mean this is a sampled mixture, not the full upstream corpora.

Evaluation

The following values come from the published evaluation report, covering 17,476 held-out decisions. These are recorded results, not a new evaluation performed for this model card.

Metric Uncalibrated Calibrated
Accuracy 0.798581 0.798581
Log loss 0.492302 0.474615
Brier score 0.280210 0.276855
ECE (15 bins) 0.038409 0.014262
Ordinal MAE (997 score decisions) 0.372761 0.392342

Accuracy accepts any candidate with positive target mass. Log loss uses the complete target distribution; Brier score sums squared probability errors across candidates. ECE compares maximum predicted probability to target mass at the winning candidate. Ordinal MAE measures error in expected zero-based rubric indices. Lower is better for the error metrics.

A single positive temperature was fitted on 16,267 separate calibration decisions, searching [0.05, 20] with temperature 1 included as a baseline. Calibration log loss improved from 0.469281 to 0.452630; see calibration.json. Temperature scaling preserves the winning candidate but can change rubric expectations; the test ordinal MAE increased after calibration.

A separate paired Laya benchmark is also included. It measures native deployed behavior on Kev's own test mixture, with in-domain calibration for Kev and unknown Laya training overlap. Laya's input limits altered 1,597 inputs; the report separately includes a shared untruncated subset. Runtime, precision, and probability rounding differ. This comparison is not evidence of general superiority on unseen task families.

Intended uses and limitations

Kev is intended for experimentation with text routing, candidate selection, binary decisions, and rubric-based scoring. Evaluate on representative application data before deploying it, and choose thresholds using a separate validation set.

  • Probabilities are calibrated on this mixture, not guaranteed reliable under domain shift, new rubrics, or arbitrary yes/no questions.
  • Results depend on instructions, candidate wording, and the alternatives provided. The model cannot select an answer absent from the candidate set.
  • Training data can carry social biases and annotation errors. No comprehensive fairness, safety, or multilingual evaluation is reported here.
  • Aggregate mixture results do not establish performance on new task families. Consult source/task breakdowns before drawing conclusions.
  • The model has not been validated as an autonomous decision maker for medical, legal, financial, employment, or other high-impact decisions.
  • Training duration, actual training hardware, and energy/emissions measurements are not established by the published training configuration. No estimates are claimed here.

Repository files

File Purpose
model.safetensors Encoder and single-score head weights
config.json Transformers architecture configuration
tokenizer.json, tokenizer_config.json Tokenizer export
kev_config.json Decision limits, base revision, and inference temperature
training_config.json, training_args.bin Training settings; training_args.bin is not needed for inference
provenance.json Prepared dataset identity and preparation records
calibration.json Temperature fitting results
evaluation.json Uncalibrated and calibrated held-out metrics
benchmarks/laya-full/comparison.json Paired comparison report and methodology notes

License and attribution

No standalone license for the Kev fine-tuned weights is declared in this repository at the time of writing. This card does not grant a new license. Consult the base model's license and each upstream dataset's terms before use or redistribution; their terms continue to apply. See the linked source repository for code and its applicable licensing information.

Kev uses supervised decision training with temperature calibration. It does not claim to reproduce Jev's proprietary training recipe. Credit goes to the Ettin, Halo, and upstream dataset authors for their respective work.

Downloads last month
5
Safetensors
Model size
0.4B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for skundu42/kev

Finetuned
(11)
this model

Dataset used to train skundu42/kev