Rook-V1

A self-hosted decision model for bounded choices, routing and candidate scoring.

Rook-V1 combines a trained LoRA adapter and decision head with Decision 2.0 Lux 9B. This repository distributes one complete saved Rook model, with its tokenizer, configuration and native implementation. It is an uncalibrated research preview; the frozen base weights are downloaded separately.

Rook-V1 and Jev

Rook-V1 is a downloadable adaptation for self-hosting. Jev is TypeSafe's hosted decision model with typed probabilistic outputs. The following workflow comparison was supplied by the model owner on October 6, 2026.

Owner-reported benchmark results. The owner states that both models ran matched examples under identical scoring rules. Raw runs, sample counts, dataset revision, tested Jev version and the tested Rook artifact were not supplied with the card. These figures have not been independently verified or bound to the downloadable Rook-V1 artifact. They are separate from the saved evaluation below and from TypeSafe's published benchmark.

Owner-reported workflow Jev Rook Rook minus Jev
Security incident handling 72.00% 71.00% -1.00 pp
Agent trace observability 62.00% 61.00% -1.00 pp
Invoice processing 62.00% 62.50% +0.50 pp
Customer-service routing 76.00% 76.50% +0.50 pp
Equal-weight mean 68.00% 67.75% -0.25 pp

The reported comparison places Rook 0.50 percentage points ahead on invoice processing and customer-service routing, and 0.25 percentage points behind Jev on the mean. Sample counts and uncertainty were not reported.

Owner-reported Rook and Jev workflow scores

See owner-reported-benchmarks.json for provenance and the missing run metadata. TypeSafe's publisher benchmark separately reports Jev at 67.8% across four workflows. Its other model results are retained in the source snapshot; they cannot establish a matched Rook ranking.

Recorded evaluation of this saved model

This package uses saved candidate 207, which had the highest recorded fixed utility and lowest row NLL among the three Rook candidates, tied on overall accuracy. This is a distribution choice. The original evaluation retained pristine Lux after every Rook candidate failed the predeclared non-regression gates.

Recorded metric Rook-V1 Lux control Difference
Validation accuracy 81.67% (196/240) 81.25% (195/240) +0.42 pp
Choice accuracy (58 cases) 77.59% 81.03% -3.45 pp
Noul accuracy (100 cases) 96.00% 98.00% -2.00 pp
Score accuracy (82 cases) 67.07% 60.98% +6.10 pp
Fixed utility 0.887708 0.871319 +0.016389
Row NLL (lower is better) 0.737619 0.633134 +0.104485

Recorded Rook-V1 validation accuracy

The overall gain is one additional correct case. Choice/noul and deadline-date/reservation-allocation family regressions, plus increased NLL, prevented qualification. No calibration or sealed-test inference followed. Labels came from InferHub cx/gpt-6.1-sol, with blind review and same-model reference-visible adjudication; they are AI-adjudicated, not independent human gold. This custom evaluation is not the official BANKING77 or CLINC150 benchmark. Jev was not run in this saved validation evaluation.

See recorded-evaluation.json, the evaluation report and the frozen protocol.

Download and reconstruct

from huggingface_hub import snapshot_download

rook_path = snapshot_download("aryanbains/Rook-V1")

Private repository access requires local Hugging Face authentication. Reconstruct with base vllm-sr/Decision-2.0-Lux-9B, pinned revision 78bf3c03d9147aeb30b641edfe0e30ed04887ca5, plus this package's adapter/ and decision_head.safetensors. Loading the base alone loads pristine Lux. No merged 9B weight bundle is included.

The native decision model uses a custom head and verified source bundle. Use the saved implementation/rook/native/full_checkpoint.py and checkpoint.py reload functions with the original pinned bundle, audit/config/data identity and compatible GPU environment. They retain their strict integrity and runtime guards. This is not a generic text-generation AutoModel repository; a one-line inference API is not provided. publication-manifest.json lists every distributed file and its SHA256.

Package component Location
Trained LoRA adapter adapter/adapter_model.safetensors
Trained decision head decision_head.safetensors
Tokenizer and template tokenizer.json, tokenizer_config.json, chat_template.jinja
Saved model configuration decision_config.json, full-config.json
Native implementation implementation/rook/
Original integrity record rook-checkpoint.json
Historical optimizer/scheduler/RNG state trainer-state.pt

The original manifest and trainer state are retained to preserve the historical reload contract; they include training diagnostics and row identifiers. Raw examples and held-out packets are excluded. The other saved candidates are not distributed in this repository. Existing local recovery artifacts remain unchanged.

Training and runtime

The adaptation run used 1,652 accepted English training rows, rank-32 LoRA (alpha 64, dropout 0.05), FP32 masked candidate cross-entropy, adapter/head learning rates 2e-5/5e-5, microbatch 1, accumulation 8 and seed 17. This saved model completed one pass, 207 updates; the full run continued to three passes and 621 updates. Data combines adapted public intent datasets and original decision tasks; see data attribution.

Recorded runtime: Python 3.12.15, PyTorch 2.12.0+cu130, Transformers 5.17.0, PEFT 0.21.2, Accelerate 1.15.0, Tokenizers 0.23.2, Safetensors 0.8.0, Triton 3.7.0, CUDA 13.0, NVIDIA L40S, driver 580.126.09. Full dependencies and deterministic settings are recorded in the original manifest. CPU inference, merged weights and other serving engines have not been verified.

Input-only pricing proposal

Offering USD / million input tokens USD / million output tokens Status
Jev 1.13.0 $0.042 $0 Published vendor tariff, checked October 6, 2026
Rook-V1 target $0.030 $0 Proposed tariff; no hosted service or validated serving economics

Jev published input price and Rook-V1 proposed target

The proposed input tariff is 28.6% below Jev's published price. At 200 equally billable input tokens per request, one million requests would cost $6.00 under the Rook proposal versus $8.40 for Jev, with zero output charge. Tokenizers can count differently, so this is tariff arithmetic, not measured savings.

The saved model's unbatched evaluation processed 40,488 native input tokens in 40.7319 seconds. At the historical $1.866/hour compute-plus-IPv4 rate, that implies about $0.521 per million native processed tokens, excluding loading, idle time, storage, networking and operations. A $0.030 tariff would require approximately 17,278 billable tokens/second to cover that hourly compute cost. Current measurements do not establish those economics. See pricing-scenario.json.

License and model status

Apache-2.0 Rook code, with upstream Lux/Decision/Qwen notices retained. Tokenizer/source terms remain upstream terms. BANKING77 is CC-BY-4.0 and CLINC150 is CC-BY-3.0; model/code licensing does not relicense data.

The original checkpoint flags remain calibrated: false and release: false. Rook-V1 names this distributed research model; it does not change the recorded qualification decision. It is intended for reproducibility and research, with no established production-readiness or general superiority claim.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for aryanbains/Rook-V1

Finetuned
Qwen/Qwen3.5-9B
Adapter
(1)
this model