Decisioncraft Kai 0.6B

Decisioncraft is a developer-request classifier adapted from Kai. It decides whether to suggest a command, ask for missing information, mark a request unsupported, or defer under a separately fitted confidence policy. It does not generate or execute commands.

This package contains the locally verified selected adapter, trained decision head, fitted routing policy, source and evaluation evidence. Use the full 40-character snapshot commit from this repository’s Files and versions / history for reproducible downloads. The installed-wheel CUDA reference check passed locally; that check alone does not establish a fresh Hub download.

Known limitation: despite perfect raw classification on the controlled synthetic holdout, a separate illustrative policy example asking to find the literal text docker run in a local file, explicitly without launching Docker, was confidently misrouted as unsupported. Do not treat the synthetic score as real-user reliability.

Intended behavior

  • suggest: a GNU local file/text or local Git operation is sufficiently specified to attempt a command suggestion.
  • clarify: a supported operation lacks material information or has an unresolved conflict.
  • unsupported: at least one positively requested action is outside that scope, including mixed local/external workflows.
  • defer: the calibrated maximum probability falls below a fitted confidence threshold. This is an operating-policy outcome, not a gold semantic label.

Quoted text and explicitly prohibited activities are not positive requested actions. Context can supply filenames, targets and preferences. The classifier is not a safety filter, and suggest does not imply downstream command correctness.

The integration emits a suggestion_request containing request/context for a downstream generator; it does not itself call Bashcraft. Its clarification message is fixed: “Please supply or resolve the missing details needed for this local file, text or Git request. This is a fixed prompt, not a generated diagnosis of which detail is missing.” Other routes emit a scope message or request human review. Outputs retain raw logits, raw probabilities, calibrated probabilities, raw argmax, policy and workflow.

Model and training

The base is Decision-2.0-Kai-0.6B, pinned at cd49ea3813fd8ba0928a9a23ef6c9a0f2f0cd764, with 597,103,104 parameters including its decision head. The released candidate is the predetermined primary seed 42: LoRA rank 4, alpha 8, dropout 0 plus the trained decision head, totaling 3,576,320 trainable parameters. Training used 1,008 examples, two epochs, AdamW learning rate 1e-4, zero weight decay, microbatch two, effective batch 32 and gradient clipping 1. The decision loss is cross-entropy over candidate logits, not next-token text generation.

A bounded validation study compared rules, frozen MiniLM embeddings with a linear head, two unchanged Kai prompts, head-only training and LoRA at 252/1,008 examples. The compact prompt and LoRA recipe were selected on validation before calibration or final scoring. Seeds 43 and 44 repeat the selected recipe without replacing primary seed 42. The adaptation starts from original Kai, not a continued head-only checkpoint.

The runtime keeps parameters FP32 with BF16 autocast on CUDA, uses the inspected lower-level decision model and enforces a 512-token encoded-input limit. The pinned upstream tokenizer behavior and warnings were retained. The base's advertised context limit is not this project's tested input contract.

Data and review

Decisioncraft data contains 1,800 synthetic views from 150 original scenario groups, 12 related views per group and 600 rows per label. Splits are training 1,008, selection-validation 252, calibration 180 and final test 360. The final test has 30 groups and two 180-row cohorts: familiar-family scenarios and whole-family holdouts for archive/compression and database operations.

An AI coding agent authored scenarios and a deterministic renderer. A separate agent reviewed all source semantics, generated derivations and final-case labels. The user approved 12 policy examples outside the corpus and the disclosed synthetic study. This is not human annotation of 1,800 rows or a natural request sample. Only request and context are model inputs.

Frozen final evaluation

The final protocol was committed before test inference at 3c1431b5ec2313394f83ffdc00543a9913b4ee71; all eight contenders completed with zero technical prediction failures. Each metric below uses the same 360 final cases, including 120 ambiguous and 120 complete supported cases.

Method Correct clarification Unnecessary clarification Coverage Routed macro-F1
Rules 2/120 0/120 100.0% 0.4275
MiniLM seed 42 50/120 5/120 83.3% 0.6886
Unchanged compact Kai 62/120 3/120 87.8% 0.5170
Adapted Kai seed 42 120/120 0/120 92.2% 0.9580
Adapted Kai seed 43 113/120 0/120 87.8% 0.9342
Adapted Kai seed 44 119/120 0/120 86.9% 0.9274

The primary recall gain is +48.3 percentage points, paired whole-scenario bootstrap 95% interval [+40.2, +56.8] points (2,000 replicates, seed 101). This met the predefined positive criterion plus final false clarification at most 10%. It is a claim about this synthetic study. Family results are descriptive; two held-out families do not establish general developer-domain transfer.

Primary raw argmax classified all 360 cases correctly. The frozen operating policy nonetheless deferred 28: 21 complete supported and seven unsupported requests. Routed accuracy is 332/360 (92.2%), not 100%. Temperature scaling worsened primary final log loss from 0.001395 to 0.021902 despite helping on the calibration split. The 90% calibration coverage target did not transfer to several other contenders. These policies were not changed after the test.

The retained policy uses temperature 3.4673685045253166, clarify threshold 0.9496928485360457 and defer threshold 0.9479888848264989. It defers first when maximum calibrated probability is strictly below the defer threshold; otherwise it clarifies if the clarify threshold is met, then selects the greater suggest/unsupported probability with suggest on ties.

The complete study report, numerical summary, independent review, latency record and all-example demonstration retain unfavorable findings as well as the primary result. The demonstration correctly routed eight of 12 known policy examples, deferred three and confidently misrouted one. These examples are illustrative, not a new unbiased benchmark.

Cost and limitations

On a shared RTX 5090, adapted single-request routing took median 14.744 ms and p95 15.016 ms over 60 fixed validation requests after three warmups. This includes encoding, inference, routing, output validation and JSON serialization; model loading was 2.375 seconds separately. PyTorch allocator peak was 2.28 GiB, not total device memory. CPU MiniLM median/p95 was 6.410/8.598 ms with four threads; different devices and shared-machine conditions prevent a universal hardware ranking.

Shared templates, explicit missing-information cues, 150 authored groups, limited human review and only two withheld families limit generalization. The known literal-text/negation failure is a concrete warning against inferring robustness from perfect synthetic raw accuracy. Long inputs, natural user requests and other domains remain unvalidated. Any tuning based on these final errors needs a new holdout.

Install and run

The development GitHub repository is private; these public assets provide the software and evidence without requiring access to it. Use Linux, uv, and a CUDA-capable GPU with sufficient free memory for the verified inference path. The measured environment used Python 3.13, CUDA 13.0 and RTX 5090; other environments can produce numerical differences. CPU inference is available but is not the exact CUDA reference check.

Choose a fresh working directory. Replace REPLACE_WITH_FULL_40_CHARACTER_MODEL_COMMIT with a commit from this repository's Files/history; do not use a moving main revision for a reproducibility claim. The Hugging Face downloader is bootstrapped separately from the locked model runtime:

MODEL_REVISION=REPLACE_WITH_FULL_40_CHARACTER_MODEL_COMMIT
uvx --from huggingface-hub==1.33.0 hf download \
  nima1/decisioncraft-kai-0.6b \
  --revision "$MODEL_REVISION" \
  --local-dir decisioncraft-model --exclude .gitattributes

RELEASE_DIR="$(pwd)/decisioncraft-model"
RUNTIME_ENV="$(pwd)/decisioncraft-runtime"
BASE_DIR="$(pwd)/decisioncraft-base"

The download excludes the Hub-generated .gitattributes; every runtime/document asset remains hash-bound. The snapshot includes release-manifest.json. Before extracting the source, check its archive hash and the wheel hash against that manifest:

python3 - <<'PY'
import hashlib
import json
from pathlib import Path

release = Path("decisioncraft-model")
manifest = json.loads((release / "release-manifest.json").read_text())
for name in ("software/decisioncraft-source.tar.gz",
             "software/decisioncraft-0.1.0-py3-none-any.whl"):
    assert hashlib.sha256((release / name).read_bytes()).hexdigest() == manifest["files_sha256"][name]
print("source archive and wheel hashes match the pinned release")
PY

tar -xzf "$RELEASE_DIR/software/decisioncraft-source.tar.gz"
cd decisioncraft-source
UV_PROJECT_ENVIRONMENT="$RUNTIME_ENV" \
  uv sync --locked --extra ml --no-dev --no-install-project
uv pip install --python "$RUNTIME_ENV/bin/python" --no-deps \
  "$RELEASE_DIR/software/decisioncraft-0.1.0-py3-none-any.whl"

The exact uv.lock preserves package-source and artifact identities, including the CUDA package index. A generic merged-index requirements installation is not the verified installation route. The command installs dependencies from the lockfile, then installs the distributed wheel without resolving them again.

Verify every package file, download the pinned original Kai base to a new directory, and run the saved-reference check:

"$RUNTIME_ENV/bin/decisioncraft" verify-release --release "$RELEASE_DIR"
"$RUNTIME_ENV/bin/decisioncraft" download-base \
  --release "$RELEASE_DIR" --output "$BASE_DIR"
"$RUNTIME_ENV/bin/decisioncraft" check-release \
  --release "$RELEASE_DIR" --base "$BASE_DIR" --device cuda:0
"$RUNTIME_ENV/bin/decisioncraft" predict \
  --release "$RELEASE_DIR" --base "$BASE_DIR" --device cuda:0 \
  --request 'List the entries here.' \
  --context 'current directory is `/workspace/manuals`.'

Local installed-wheel testing matched all six saved reference logits with maximum absolute difference 0. check-release checks checkpoint/runtime consistency, not real-user quality. The example emits a validated JSON routing result, with command_generated: false and command_executed: false. Repeated downloads require new destination directories.

Kai uses custom code. The runtime retrieves the exact original base and verifies reviewed source/file hashes before importing its Python. The upstream Transformers wrapper is inference-only; generic training or saving through that wrapper is not this integration path.

The release also includes evidence/decisioncraft-evidence.tar.gz, the full reports and figures, configurations, tutorials and license notices. Extract the evidence archive next to decisioncraft-source to follow the public saved-prediction audit. That CPU-only audit verifies reported numbers from retained predictions; it does not rerun training, historical latency or inference provenance. A full new experiment follows the milestone tutorials and records its own Git commits and artifact identities.

Attribution and license

Decisioncraft code, authored data and adaptation use Apache-2.0. Kai is provided by the vLLM Semantic Router Team; its direct weight origin is Qwen/Qwen3-0.6B-Base at da87bfb608c14b7cf20ba1ce41287e8de496c0cd. These are upstream contributions, not authored by Decisioncraft. LICENSE and NOTICE retain attribution; no upstream-author endorsement is claimed.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for nima1/decisioncraft-kai-0.6b

Adapter
(1)
this model