Instructions to use skundu42/kev with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use skundu42/kev with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="skundu42/kev")# Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("skundu42/kev") model = AutoModelForSequenceClassification.from_pretrained("skundu42/kev", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Kev: a 400M-parameter text decision model
Kev scores user-supplied candidates for a text input. It supports categorical choices, yes/no probabilities, and ordered rubric scores through a shared encoder and a single-output scoring head. It is fine-tuned from Ettin Encoder 400M using Halo.
Source code · Inference guide · Training guide
Model details
| Property | Value |
|---|---|
| Maintainer | skundu42 |
| Parameters | 395,832,321 (approximately 400M) |
| Architecture | ModernBertForSequenceClassification, 28 layers, hidden size 1,024 |
| Head | One scalar per candidate; softmax across the supplied candidates |
| Base revision | 7662476d60abb071a5bd319c9f3074f3072c062d |
| Export | Safetensors, FP32 weights |
| Supported decision length | 1,024 tokens per encoded prompt/candidate pair |
| Candidate count | 2–16 |
| Calibration temperature | 1.3734692667114208 |
The prompt includes the decision type, instructions, state, and full criteria list. Each candidate is encoded with that prompt and receives a scalar score. The runtime computes softmax(scores / temperature). The backbone's positional capacity does not override Kev's 1,024-token application limit. Overlength inputs are rejected without silent truncation.
Supported decisions
| Type | Criteria | Result |
|---|---|---|
choice |
Mapping of candidate names to descriptions | Selected name and probabilities |
noul |
Exactly false and true, or omitted for No/Yes |
Probability of true |
score |
Ordered list of rubric levels | Distribution and expected zero-based index |
A five-level rubric produces a score between 0 and 4. For choice/score, the confidence summary is (K * max_probability - 1) / (K - 1); it is not a validated probability of correctness. Kev scores alternatives; it does not generate text.
Quickstart
Use a Linux CUDA GPU environment configured as described in the training guide. The project runtime expects Python 3.12, PyTorch 2.11.x and Transformers >=5.16.1,<5.17. The supplied loader requires CUDA and uses BF16 autocast when supported. Halo is used for training; inference uses the Kev wrapper and Transformers.
Clone the source and download the export. If repository access requires authentication, run hf auth login with an account that has access.
git clone https://github.com/skundu42/kev.git
cd kev
hf download skundu42/kev \
--revision e62becd4fda8a69e9494c5e22ab60a4a1371ba69 \
--local-dir ./models/kev \
--include model.safetensors config.json kev_config.json tokenizer.json tokenizer_config.json
Run this Python example from the cloned source directory:
import json
from kev.inference import DecisionModel
model = DecisionModel.from_pretrained("./models/kev")
result = model.predict(
state="I was charged twice for my subscription.",
questions={
"route": {
"type": "choice",
"instructions": "Choose the team that should handle this request.",
"criteria": {
"billing": "Payments, invoices, and refunds",
"technical": "Bugs and product support",
"sales": "Plans and purchasing",
},
},
"refund_needed": {
"type": "noul",
"instructions": "Does this request describe a billing error that may need a refund?",
"criteria": {"false": "No", "true": "Yes"},
},
"urgency": {
"type": "score",
"instructions": "Rate the urgency expressed by the customer.",
"criteria": ["low", "medium", "high"],
},
},
)
print(json.dumps(result, indent=2))
DecisionModel.from_pretrained accepts a local export directory, not a Hub ID. A generic text-classification pipeline does not implement Kev's prompt formatting, candidate grouping, or temperature calibration; use the wrapper for the supported decision API.
For JSON CLI inference:
python -m kev.inference --model ./models/kev --input examples/request.json
See the HTTP API guide for authenticated serving and request limits.
Training and data
The published training configuration records one epoch, learning rate 2e-5, weight decay 0.01, warmup ratio 0.03, seed 42, per-device batch size 1, gradient accumulation 32, BF16 training, and gradient checkpointing. Training minimizes cross-entropy across candidate scores against the target distribution, including soft targets. Validation loss selects the checkpoint.
The prepared dataset is skundu42/kev-prepared at revision 5cea4539b6676edfb4fd1fb60878e9aee602b0cd. The published provenance records source revisions, filtering, sampling, exclusions, and retained counts.
| Split | Decisions |
|---|---|
| Training | 393,549 |
| Validation | 16,947 |
| Calibration | 16,267 |
| Test | 17,476 |
Sources include:
- tasksource/zero-shot-label-nli
- tasksource/tasksource-instruct-v0, restricted to Yelp review ratings and TweetEval sentiment
- nyu-mll/multi_nli
- facebook/anli
- tasksource/defeasible-nli
- tasksource/FOL-nli
- tasksource/doc-nli
- tasksource/bigbench, using eligible structured multiple-choice tasks
- Rowan/hellaswag
- google/boolq
Native labeled test partitions are retained. Native validation is split 50/50 into validation/calibration when a labeled test partition exists, otherwise 50/25/25 into validation/calibration/test. Normalized content and related groups are deduplicated across partitions, with holdouts taking priority. Exact/group deduplication does not detect paraphrases or establish absence of pretraining overlap. Source and task caps mean this is a sampled mixture, not the full upstream corpora.
Evaluation
The following values come from the published evaluation report, covering 17,476 held-out decisions. These are recorded results, not a new evaluation performed for this model card.
| Metric | Uncalibrated | Calibrated |
|---|---|---|
| Accuracy | 0.798581 | 0.798581 |
| Log loss | 0.492302 | 0.474615 |
| Brier score | 0.280210 | 0.276855 |
| ECE (15 bins) | 0.038409 | 0.014262 |
| Ordinal MAE (997 score decisions) | 0.372761 | 0.392342 |
Accuracy accepts any candidate with positive target mass. Log loss uses the complete target distribution; Brier score sums squared probability errors across candidates. ECE compares maximum predicted probability to target mass at the winning candidate. Ordinal MAE measures error in expected zero-based rubric indices. Lower is better for the error metrics.
A single positive temperature was fitted on 16,267 separate calibration decisions, searching [0.05, 20] with temperature 1 included as a baseline. Calibration log loss improved from 0.469281 to 0.452630; see calibration.json. Temperature scaling preserves the winning candidate but can change rubric expectations; the test ordinal MAE increased after calibration.
A separate paired Laya benchmark is also included. It measures native deployed behavior on Kev's own test mixture, with in-domain calibration for Kev and unknown Laya training overlap. Laya's input limits altered 1,597 inputs; the report separately includes a shared untruncated subset. Runtime, precision, and probability rounding differ. This comparison is not evidence of general superiority on unseen task families.
Intended uses and limitations
Kev is intended for experimentation with text routing, candidate selection, binary decisions, and rubric-based scoring. Evaluate on representative application data before deploying it, and choose thresholds using a separate validation set.
- Probabilities are calibrated on this mixture, not guaranteed reliable under domain shift, new rubrics, or arbitrary yes/no questions.
- Results depend on instructions, candidate wording, and the alternatives provided. The model cannot select an answer absent from the candidate set.
- Training data can carry social biases and annotation errors. No comprehensive fairness, safety, or multilingual evaluation is reported here.
- Aggregate mixture results do not establish performance on new task families. Consult source/task breakdowns before drawing conclusions.
- The model has not been validated as an autonomous decision maker for medical, legal, financial, employment, or other high-impact decisions.
- Training duration, actual training hardware, and energy/emissions measurements are not established by the published training configuration. No estimates are claimed here.
Repository files
| File | Purpose |
|---|---|
model.safetensors |
Encoder and single-score head weights |
config.json |
Transformers architecture configuration |
tokenizer.json, tokenizer_config.json |
Tokenizer export |
kev_config.json |
Decision limits, base revision, and inference temperature |
training_config.json, training_args.bin |
Training settings; training_args.bin is not needed for inference |
provenance.json |
Prepared dataset identity and preparation records |
calibration.json |
Temperature fitting results |
evaluation.json |
Uncalibrated and calibrated held-out metrics |
benchmarks/laya-full/comparison.json |
Paired comparison report and methodology notes |
License and attribution
No standalone license for the Kev fine-tuned weights is declared in this repository at the time of writing. This card does not grant a new license. Consult the base model's license and each upstream dataset's terms before use or redistribution; their terms continue to apply. See the linked source repository for code and its applicable licensing information.
Kev uses supervised decision training with temperature calibration. It does not claim to reproduce Jev's proprietary training recipe. Credit goes to the Ettin, Halo, and upstream dataset authors for their respective work.
- Downloads last month
- 5
Model tree for skundu42/kev
Base model
jhu-clsp/ettin-encoder-400m