SelfJev-4B: read the document once, branch across questions, and score their candidate answers.

SelfJev-4B

Structured decisions from text. A 4B backbone. Your own GPU.

Route a customer message. Check a claim. Review an AI response. SelfJev answers questions over the options you supply and returns probabilities through a Python SDK or HTTP API. Its default engine shares the document computation across questions and scores candidates directly, without generating an answer token by token.

Quickstart · Results · Public datasets · Model size · Architecture · Latency · Research · Source code

Ask for Get back Use it for
Yes / no Probability of yes Verification, guardrails, detection
One choice Selected option and probabilities Routing, intent, categorization
An ordered score Score and probabilities over your levels Review and grading
Every matching option Selected set and per-option probabilities Tags, checklists, multiple issues

This repository contains the 230 MB LoRA adapter, not the entire model. The Qwen3.5-4B base weights download separately. The 4B model size refers to the backbone; the adapter stores the learned update.

Results

SelfJev and Jev accuracy on two authored test suites and the public-source development tasks; exact values below.

Evaluation Questions SelfJev-4B Jev API SelfJev evidence
Text decisions 1,991 95.7% 97.2% report
AI response review 946 93.1% 92.5% report
Public-source tasks 3,300 83.4% 82.1% report

What these measure. Text decisions cover authored routing, policy, evidence and difficult language cases. AI response review covers verification, judging, scoring, guardrails and jailbreak detection. Public-source tasks cover 11 adapted datasets. Each score counts a question as correct only when the expected answer matches; for multiple selections, the entire set must match.

These are recorded results from the current TreeServer engine and matched Jev predictions, not scores from a newly merged vLLM release. The authored sets use AI-checked labels, not human ground truth. Both have informed research direction. The public-source set is a repeatedly reused development benchmark. A fresh independent evaluation remains necessary.

How does it compare with another open model?

On the same 1,991 text-decision questions, the recorded Eikos-4B run scored 92.8% using our task adapter and its letter-based readout. This is a comparison under this project's protocol, not a general ranking of open models. Eikos report

There is no SelfJev S1Bench result yet. Lev's published S1Bench scores cannot be compared with our authored-test scores. Comparing SelfJev, Lev, Reflex and other decision models requires the same dataset revision, item IDs, candidate sets, scoring rules and coverage.

Public datasets

Accuracy for all 11 public-source datasets. Orange circles are SelfJev; green squares are Jev. Values are also available in the table below.

83.4% across 3,300 questions. There are 300 questions per source, so the macro average over sources and the question-weighted average are equal. The 171 authored development questions are excluded here; the full 3,471-question development benchmark remains available in the reports.

Source dataset Task in our protocol Training exposure Questions SelfJev Jev
CLINC150 Intent routing held-out source 300 95.0% 94.3%
DBpedia Topic classification held-out source 300 98.0% 98.0%
TREC Question type held-out source 300 91.3% 94.0%
Emotion Emotion classification held-out source 300 57.3% 57.7%
BoolQ Reading comprehension held-out source 300 87.7% 90.7%
SST-2 Sentiment held-out source 300 91.3% 96.7%
Banking77 Banking intent seen source 300 96.3% 95.7%
AG News News topic seen source 300 90.3% 87.3%
TweetEval Tweet sentiment seen source 300 67.3% 64.3%
MNLI Entailment (binary) seen source 300 95.0% 92.7%
GoEmotions Emotion tags (exact set) seen source 300 47.7% 31.7%

“Held-out source” means that dataset was excluded from SelfJev's task-specific training. “Seen source” means other rows from that source were used in training. Neither establishes absence from Qwen's pretraining.

These are adapted tasks, not official dataset leaderboard scores. For example, Banking77 uses eight candidate options per question, CLINC uses seven or eight, and DBpedia uses six. MNLI is recast as binary entailment. GoEmotions scores exact matching over six candidate tags. Our preprocessing, sampling and answer sets are part of the evaluation and must accompany the numbers.

Known development-data issues include shared MNLI premises across splits and near-duplicate authored policy texts. See data methodology. All 11 slices are shown, including weaker emotion and sentiment results.

Download chart data · SelfJev predictions · Matched Jev predictions

Accuracy and model size

Accuracy against nominal backbone parameters for seven measured systems on identical public-source questions. Full results below.

Model / recipe Backbone size Training Accuracy Evidence
Qwen3 reranker 0.6B 0.6B stock 62.2% report
Qwen3 reranker 4B 4B stock 63.8% report
Qwen3 reranker 8B 8B stock 67.0% report
Qwen3 + LoRA 0.6B 0.6B early fine-tune 74.5% report
Qwen3 + LoRA 4B 4B early fine-tune 80.8% report
Qwen3 + LoRA 8B 8B early fine-tune 81.0% report
SelfJev-4B 4B current release 83.4% report

All points use the same 3,300 question IDs, labels and candidate sets. The stock models are rerankers mapped to our decision tasks. The earlier LoRA runs use older training recipes; SelfJev uses a different backbone generation and more training data. This shows the measured systems' trade-off between size and accuracy. It does not isolate the effect of parameter count or establish that a 4B model outperforms larger models generally.

Sizes are nominal backbone parameter counts, not checkpoint megabytes or runtime memory. Jev appears as a horizontal reference because its parameter count is undisclosed. We do not assign a guessed size to proprietary models.

AI response review

Breakdown of the 946-question authored suite, using the same current-engine predictions as the overview:

Task Questions SelfJev Jev
Factual verification 165 87.3% 86.7%
Response judging 222 94.1% 94.6%
Quality scoring 200 92.5% 93.0%
Policy guardrails 170 96.5% 92.9%
Jailbreak detection 189 94.7% 94.2%

Labels were authored and checked with AI judges. Questions about one document are correlated; small differences should not be read as proof of superiority. This suite measures our defined answer choices and criteria, not unrestricted judgment quality. Dataset design · Predictions

Quickstart

On Linux with an NVIDIA GPU, Git and uv installed:

GIT_LFS_SKIP_SMUDGE=1 git clone https://github.com/Jwuthri/SelfJev.git
cd SelfJev
uv sync --frozen --no-dev --extra serve --extra gpu
uv run --no-sync python - <<'PYTHON'
from huggingface_hub import snapshot_download
snapshot_download(
    "Jwuthrich/selfjev-4b",
    local_dir="weights/selfjev_4b_hf",
    allow_patterns=["adapter_model.safetensors", "adapter_config.json", "model.json"],
)
PYTHON

# Choose a secret for your own server.
export SELFJEV_API_KEYS="replace-with-your-long-random-secret"
uv run --no-sync selfjev serve \
  --adapter weights/selfjev_4b_hf --host 127.0.0.1 --port 8000

In another terminal, use the same secret:

from selfjev import SelfJev, Noul, Choice

client = SelfJev(
    base_url="http://127.0.0.1:8000",
    api_key="replace-with-your-long-random-secret",
)
result = client.system_one(
    state="I was charged twice. Please refund the duplicate payment.",
    questions={
        "refund": Noul("Does the customer want a refund?"),
        "team": Choice("Which team should handle this?", {
            "billing": "payments and refunds",
            "support": "technical issues",
        }),
    },
)
print(result.nouls["refund"].noul)
print(result.choices["team"].choice)

The example shows the interface; it does not present fabricated model output. A generic Transformers text-generation or classification pipeline does not reproduce SelfJev's prompts, scoring rules or API. The server key is a secret you choose for your deployment, unrelated to Hugging Face credentials.

HTTP API · Self-hosting and AWS · Fine-tuning

Architecture

The shared-prefix tree: one document feeds independent question branches; each question feeds its candidate branches; scores become option probabilities. Attention sees ancestors only, and DeltaNet inherits parent recurrent states.

The default engine builds a shared-prefix tree: compute the document once, branch for each question, then branch for candidate answers. It reads yes/no evidence from the logits and converts those scores into the requested answer type. The server includes every candidate in the question, matching training.

The tree saves repeated computation at two levels: every question reuses the document, and every candidate reuses its question. In full-attention layers, the tree mask lets a branch read its ancestors but not sibling branches. In DeltaNet layers, branches inherit the recurrent state of their parent. Each candidate therefore sees its own document → question → candidate path, up to numerical differences between kernels. The diagram describes TreeServer; vLLM uses its own prefix-cache execution.

Component Current release
Backbone Qwen3.5-4B
Base revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a
Fine-tuning Rank-64 LoRA on attention and DeltaNet projections
Default serving TreeServer, shared document and question computation
Alternative serving vLLM with a merged checkpoint and prefix caching
Training examples 79,943 non-test questions; texts up to 16K tokens
Training target 50% checked labels + 50% stored Jev probabilities
Training schedule One epoch, learning rate 2e-4

Jev probabilities served as a teacher signal; they did not decide the authored training labels. Authored labels were checked by an AI judge. Full lineage is recorded in model.json, with architecture detail in the research documentation.

The research path

Recorded accuracy for selected research checkpoints, from the first shared-prefix tree to the current serving engine.

The improvements came from several changes: verified difficult examples, the instruct backbone, explicit option lists, broader training coverage, tree-aware training and soft teacher targets. The chart is a history of complete systems, not an ablation assigning each gain to one change.

The research also includes smaller stock rerankers, 8B models, Jina, T5Gemma, custom architectures, distillation and RLCD. The full experiment ledger records results and dead ends; earlier implementations remain at the archive tag documented there. The checkpoint's original engine recorded a slightly different text-decision score; this card consistently uses the current engine for SelfJev's headline results.

Hardware and speed

A 24 GB NVIDIA GPU is a practical starting recommendation, not a measured minimum. Plan for 16–32 GB host RAM and at least 50 GB free disk. The adapter download size does not describe inference memory: the full backbone and runtime state must fit too.

The default engine avoids token-by-token answer generation and shares work across questions. Latency still depends on input length, question count, candidate count, GPU and concurrency. We have two sets of measurements below, with different models and engines. Current SelfJev-4B has no controlled NVIDIA latency sweep yet. Its CPU minimum is also unmeasured.

Earlier prototype on NVIDIA GPUs

Median server processing time on A10G, L40S and H100, for one or sixteen questions across four text lengths. This is the archived Qwen3 model on vLLM, not the current SelfJev checkpoint.

Text tokens Questions A10G L40S H100
8 1 55 ms 36 ms 22 ms
512 1 136 ms 55 ms 30 ms
2,048 1 361 ms 121 ms 47 ms
4,096 1 698 ms 228 ms 82 ms
8 16 401 ms 135 ms 58 ms
512 16 496 ms 163 ms 76 ms
2,048 16 808 ms 268 ms 120 ms
4,096 16 1,263 ms 424 ms 189 ms

These are server-side medians, excluding network time: ten warmed, successful requests per point, one request at a time, three answer options per question. All three sweeps use the archived tree_4b_combo Qwen3-4B model on vLLM in bf16. They were recorded in separate runs, not a simultaneous controlled hardware experiment. The plot uses a labeled logarithmic time axis; the table gives the actual milliseconds.

These measurements show what the earlier system achieved; they are not a latency promise for the current Qwen3.5 model or the merged release. A10G samples · L40S samples · H100 samples

Current SelfJev-4B on Apple Silicon

Current SelfJev-4B on Apple M5 Pro: local call latency for eight or 512 text tokens and one or sixteen questions. Exact measurements below.

Text tokens Questions M5 Pro / 48 GB
8 1 688 ms
8 16 6,786 ms
512 1 2,018 ms
512 16 8,599 ms

This uses the current adapter merged into Qwen3.5-4B, the TreeServer engine, and PyTorch MPS in bf16 on a 20-core Apple M5 Pro GPU with 48 GB unified memory. Each cell is the median of ten timed calls after two warmups, with MPS synchronized before and after each call. Inputs are synthetic repeated text, with three answer options per question. Timing includes tokenization and model work; it excludes HTTP/network.

The MPS run uses the slower PyTorch recurrent fallback. It is a lower-level engine experiment, not a validated Mac server. Model and engine differences prevent a hardware-only comparison against the NVIDIA chart. The merged download also has not been benchmarked through vLLM using these workloads.

Download latency samples, configuration and medians · Latency CSV · Full speed research

Limits and reproducibility

  • Probability is not a guarantee. The reported current-engine runs have no fitted calibration attached. Validate thresholds on application data before acting on confidence.
  • Some tasks remain weak. Emotion labeling, fine-grained sentiment and selecting an exact set of labels are visibly harder than topic and intent classification.
  • Benchmarks have a scope. AI-authored tests, reused development sets, sampled options and known overlap issues limit the conclusions. We make no S1Bench or general state-of-the-art claim.
  • Engine versions matter. These accuracy results refer to the recorded current TreeServer runs. A merged checkpoint, quantization or another serving backend needs its own verification.

Every plotted value is regenerated from saved predictions or raw timing samples. The generator checks matched IDs, targets, question types and candidate sets for report-to-report comparisons, and recomputes latency medians from ten timed samples per cell. Chart data and source SHA-256 hashes · Generator · Architecture and latency plotting module · Card template

Place the downloadable generator, plotting module and template in scripts/docs/ of a SelfJev source checkout, then run:

uv run --no-project --with matplotlib==3.11.1 python scripts/docs/build_model_card.py

This builds figures from reports and the canonical data/all.jsonl.gz; it performs no model inference. Rebuild the canonical file using the data instructions if absent.

File Purpose
adapter_model.safetensors Trained LoRA weights
adapter_config.json PEFT configuration
model.json Base revision, training provenance and original recorded scores
assets/*.png, assets/*.svg Model-card figures, raster and vector
assets/chart-data.json, assets/public-datasets.csv Underlying results and source provenance
assets/latency-data.json, assets/latency-data.csv Selected raw timing samples, methodology and recalculated medians

Adapter SHA-256: dfbf2834d883987893ec305a6093a345fd79c77f903f60f2f914cbe3b6058d1b.

The adapter uses the Qwen base model linked above. No separate adapter license is declared in this repository. Independent project; not affiliated with TypeSafe or Qwen.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Jwuthrich/selfjev-4b

Finetuned
Qwen/Qwen3.5-4B
Adapter
(659)
this model