Granite–Jev 1B: Selective Verification
A runnable research release of trained residual experts and selective Jev verification for IBM Granite 4.0 1B. It includes two expert checkpoints, a local router, and a Python inference runner. IBM's frozen backbone downloads separately from its original repository. Jev runs through TypeSafe's hosted API when the two generated answers disagree.
This release implements the R38 post-proposal policy, bank 3701, for English
four-choice reasoning. It is not a general chat release or a standalone merged model.
The experts are trained weights; Jev itself is not included. No vLLM extension is required.
AutoModelForCausalLM.from_pretrained() on this repository alone does not implement
the system: use the supplied runner.
- Maintainer: Jubba Smail, OVRLab — jubba@ovrlab.io
- Research and training code: OVRLab/jev-guided-decoding
- Paper source (working manuscript): Granite–Jev
- Evaluation report: completed R38 study
- Base: ibm-granite/granite-4.0-1b,
revision
6a7381ba1f54d684ff508d991aeb7dc580157103.
Run it
Use Python 3.12, a fresh virtual environment, and sufficient memory for a FP32
backbone (about 6.53 GB of weights, plus activations, cache and runtime overhead).
The runner supports cuda, mps and cpu; CPU inference can be slow.
CUDA users need a compatible NVIDIA driver. First-run downloads require internet.
python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install huggingface-hub==0.36.2
python -c "from huggingface_hub import snapshot_download; snapshot_download('OVRLab/granite-jev-1b-research', local_dir='granite-jev')"
cd granite-jev
python -m pip install -r requirements.txt
python run_granite_jev.py --verify
Inspect the downloaded Python source before executing it. For reproducibility, pass
the desired immutable Hugging Face commit as revision= to snapshot_download.
Create a private text file containing your TypeSafe/Jev API key, outside this download directory, and restrict its permissions. Supply its path:
python run_granite_jev.py --input example.json --device cuda \
--key-file /absolute/path/to/private/jev-key --output result.json
On Apple Silicon use --device mps; for CPU use --device cpu. Alternatively, the
runner reads the TYPESAFE_API_KEY environment variable or ~/.typesafe.ai/jev.
The key is read locally and used only as authorization to https://api.typesafe.ai;
never put it in a model config, input JSON, notebook output, or a public repository.
The key belongs to each person running the model and is distinct from a Hugging Face token.
The baseline works without Jev credentials:
python run_granite_jev.py --input example.json --device cuda \
--mode native --output native-result.json
The default is greedy FP32 generation with a 1,024-token ceiling for each candidate,
a 4,096-token context limit, and a 900-second generation deadline. --max-new-tokens
may lower the ceiling for a quick smoke check; those settings do not reproduce the
reported benchmark protocol. The one optional API request has its own 90-second
timeout. Oversized inputs fail; they are not silently truncated. Existing output
files are never overwritten.
Input and output
example.json is an original, fictional example. Inputs require context, question
and exactly four distinct choices; an optional stable id controls reproducible
anonymous candidate order. Reference answers and all other fields are rejected.
{
"context": "Every blue key is made of metal. The small key is blue.",
"question": "Which conclusion follows?",
"choices": ["The small key is metal.", "The small key is wood.",
"Every metal key is blue.", "The small key is not blue."]
}
The result contains the chosen text, its exact generated_token_ids, finish_reason,
action, jev_called, candidate traces and, if called, Jev scores and token usage.
ANSWER: 1 through ANSWER: 4 refer to the input choice order. Generation is
unrestricted over the model vocabulary; a valid answer format is not guaranteed.
The input is mapped to the frozen R38 four-choice prompt contract internally; using
that contract does not turn an arbitrary user input into a benchmark example.
How it works
Context + question + four choices
|
v
Frozen Granite -> native answer
|
Local feature extraction + router
|
Propose ONE trained expert
(correction or preservation)
|
Granite + residual expert after block 19
|
v
expert-generated answer
|
Native and expert answers agree?
/ \
yes no
| |
Keep native Jev scores anonymous pair
|
expert score - native score > 0.05?
/ \
yes no
| |
Keep expert tokens Keep native
The router uses contextual Granite features without live Jev features. Each expert
is a rank-32 residual module after zero-based block 19 and affects the final
prompt position and subsequent generation positions. Its update is
h + 0.25 * rms(h) * tanh(W_up * SiLU(W_down * h / rms(h))).
The two experts add 262,144 trained parameters in total; the router is stored
separately. The base marketed as "1B" has approximately 1.632 billion parameters.
This is not tokenwise MoE routing or access to Jev's hidden states.
Both branches run with independent caches, and Granite owns every final output
token. An extra unmodified forward pass extracts router features. Agreement means
equal parsed choices or exactly identical text; agreement retains native without
contacting Jev. Both candidates are still generated. A disagreement sends the
question, context, choices and both answer texts to the hosted provider. The runner
pins and checks returned model jev-1.13.0 and requires an advantage strictly
greater than 0.05 to replace native. Jev does not generate the final answer.
The runner processes one request at a time. API failures terminate that request with an error; they are not scored as a successful native fallback. There is no automatic paid-request replay. If the provider retires the pinned version, this release needs a separately validated update. It does not provide a hosted Hugging Face Inference Endpoint, chat widget, or vLLM integration automatically.
Recorded evaluation of this exact bank
These are historical full-cohort generated-answer results under our own prompt and parser, not official leaderboard scores or new release-validation benchmarks. Bank 3701 was packaged by ascending registered seed, not by a best-score rule. The paper's main estimates average banks 3701/3702/3703; this download contains only 3701 and must not be assigned those averages.
| System | ReClor validation, 500 | LogiQA English Eval, 651 |
|---|---|---|
| Frozen Granite, native | 45.80% (229/500) | 31.95% (208/651) |
| This bank + selective Jev | 57.80% (289/500) | 37.94% (247/651) |
| Three-bank selective mean, paper reference | 55.67% | 37.53% |
| Majority-choice reference | 25.80% | 41.47% |
For bank 3701, Jev was called on 194/500 ReClor cases and 312/651 LogiQA cases. Relative to native, selection fixed 61 answers and damaged 1 on ReClor; it fixed 45 and damaged 6 on LogiQA. LogiQA performance remains below its 41.47% majority reference. These results establish a bounded improvement over this native baseline, not strong general intelligence or superiority over larger models.
The full report includes no-Jev local commit, unconditional proposal, cross-context donor verification, fixed experts, shared-adapter and full-bank controls, along with costs, statistical uncertainty and limitations. More candidates can improve accuracy at additional cost. A single native answer uses less compute than this two-generation method. Agreement saves API calls, not candidate generation. Hosted Jev cost, latency and undisclosed model size belong in system-level comparisons. No results from an unfinished broader benchmark campaign are claimed here.
Training and provenance
The two experts were trained in the earlier R37 bank study on 437 MuSR cases, and the routers were fitted on a separate 145 cases. Original Granite and Jev weights remained frozen. R38 reused those artifacts without further training or threshold tuning. See the training study.
artifact_manifest.json binds the two safetensors, no-Jev router, configuration,
vendored runtime and transport to their SHA-256 digests. SHA256SUMS covers the
complete public release inventory except itself. The _frozen source files retain
their original bytes from source revision 366b078c07fa6573753d8f78df82b930bdf7d3b1;
some shared modules include historical helpers unused by this runner. Source
provenance is not a claim that every experiment in those modules is released here.
See VALIDATION.md for packaging checks, SECURITY.md for known advisories in the historical pinned dependencies and the supported loading scope, and NOTICE.md for external licenses and attribution. OVRLab code and expert artifacts are released under MIT. The separately downloaded IBM model remains under Apache 2.0; TypeSafe service terms and API charges apply independently. No IBM backbone weights, Jev weights, benchmark prose, or API credentials are redistributed in this repository.
Limitations and intended use
Use this release to study selective verification on English four-choice reasoning. Open-ended chat, multilingual reasoning, coding, long-context tasks and serving under concurrency are not validated here. Training-data exposure in upstream models is unknown. Generated explanations are not independently validated proofs. Three-bank research replication is not independent third-party reproduction. GPU/CPU numerical differences and future hosted-provider changes can affect outputs. FP32 precision and the supplied pinned dependencies are part of the tested contract; quantization or a different engine requires new validation.
Please cite this repository at its immutable revision and the linked manuscript when reporting results; an arXiv identifier has not been assigned in this release.
Model tree for OVRLab/granite-jev-1b-research
Base model
ibm-granite/granite-4.0-1b-base