Granite–Jev 1B: Selective Verification

A runnable research release of trained residual experts and selective Jev verification for IBM Granite 4.0 1B. It includes two expert checkpoints, a local router, and a Python inference runner. IBM's frozen backbone downloads separately from its original repository. Jev runs through TypeSafe's hosted API when the two generated answers disagree.

This release implements the R38 post-proposal policy, bank 3701, for English four-choice reasoning. It is not a general chat release or a standalone merged model. The experts are trained weights; Jev itself is not included. No vLLM extension is required. AutoModelForCausalLM.from_pretrained() on this repository alone does not implement the system: use the supplied runner.

Run it

Use Python 3.12, a fresh virtual environment, and sufficient memory for a FP32 backbone (about 6.53 GB of weights, plus activations, cache and runtime overhead). The runner supports cuda, mps and cpu; CPU inference can be slow. CUDA users need a compatible NVIDIA driver. First-run downloads require internet.

python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install huggingface-hub==0.36.2
python -c "from huggingface_hub import snapshot_download; snapshot_download('OVRLab/granite-jev-1b-research', local_dir='granite-jev')"
cd granite-jev
python -m pip install -r requirements.txt
python run_granite_jev.py --verify

Inspect the downloaded Python source before executing it. For reproducibility, pass the desired immutable Hugging Face commit as revision= to snapshot_download.

Create a private text file containing your TypeSafe/Jev API key, outside this download directory, and restrict its permissions. Supply its path:

python run_granite_jev.py --input example.json --device cuda \
  --key-file /absolute/path/to/private/jev-key --output result.json

On Apple Silicon use --device mps; for CPU use --device cpu. Alternatively, the runner reads the TYPESAFE_API_KEY environment variable or ~/.typesafe.ai/jev. The key is read locally and used only as authorization to https://api.typesafe.ai; never put it in a model config, input JSON, notebook output, or a public repository. The key belongs to each person running the model and is distinct from a Hugging Face token.

The baseline works without Jev credentials:

python run_granite_jev.py --input example.json --device cuda \
  --mode native --output native-result.json

The default is greedy FP32 generation with a 1,024-token ceiling for each candidate, a 4,096-token context limit, and a 900-second generation deadline. --max-new-tokens may lower the ceiling for a quick smoke check; those settings do not reproduce the reported benchmark protocol. The one optional API request has its own 90-second timeout. Oversized inputs fail; they are not silently truncated. Existing output files are never overwritten.

Input and output

example.json is an original, fictional example. Inputs require context, question and exactly four distinct choices; an optional stable id controls reproducible anonymous candidate order. Reference answers and all other fields are rejected.

{
  "context": "Every blue key is made of metal. The small key is blue.",
  "question": "Which conclusion follows?",
  "choices": ["The small key is metal.", "The small key is wood.",
              "Every metal key is blue.", "The small key is not blue."]
}

The result contains the chosen text, its exact generated_token_ids, finish_reason, action, jev_called, candidate traces and, if called, Jev scores and token usage. ANSWER: 1 through ANSWER: 4 refer to the input choice order. Generation is unrestricted over the model vocabulary; a valid answer format is not guaranteed. The input is mapped to the frozen R38 four-choice prompt contract internally; using that contract does not turn an arbitrary user input into a benchmark example.

How it works

Context + question + four choices
             |
             v
    Frozen Granite -> native answer
             |
    Local feature extraction + router
             |
      Propose ONE trained expert
       (correction or preservation)
             |
    Granite + residual expert after block 19
             |
             v
       expert-generated answer
             |
   Native and expert answers agree?
       /                      \
     yes                       no
      |                         |
 Keep native          Jev scores anonymous pair
                                |
                    expert score - native score > 0.05?
                         /                   \
                       yes                    no
                        |                      |
                 Keep expert tokens       Keep native

The router uses contextual Granite features without live Jev features. Each expert is a rank-32 residual module after zero-based block 19 and affects the final prompt position and subsequent generation positions. Its update is h + 0.25 * rms(h) * tanh(W_up * SiLU(W_down * h / rms(h))). The two experts add 262,144 trained parameters in total; the router is stored separately. The base marketed as "1B" has approximately 1.632 billion parameters. This is not tokenwise MoE routing or access to Jev's hidden states.

Both branches run with independent caches, and Granite owns every final output token. An extra unmodified forward pass extracts router features. Agreement means equal parsed choices or exactly identical text; agreement retains native without contacting Jev. Both candidates are still generated. A disagreement sends the question, context, choices and both answer texts to the hosted provider. The runner pins and checks returned model jev-1.13.0 and requires an advantage strictly greater than 0.05 to replace native. Jev does not generate the final answer.

The runner processes one request at a time. API failures terminate that request with an error; they are not scored as a successful native fallback. There is no automatic paid-request replay. If the provider retires the pinned version, this release needs a separately validated update. It does not provide a hosted Hugging Face Inference Endpoint, chat widget, or vLLM integration automatically.

Recorded evaluation of this exact bank

These are historical full-cohort generated-answer results under our own prompt and parser, not official leaderboard scores or new release-validation benchmarks. Bank 3701 was packaged by ascending registered seed, not by a best-score rule. The paper's main estimates average banks 3701/3702/3703; this download contains only 3701 and must not be assigned those averages.

System ReClor validation, 500 LogiQA English Eval, 651
Frozen Granite, native 45.80% (229/500) 31.95% (208/651)
This bank + selective Jev 57.80% (289/500) 37.94% (247/651)
Three-bank selective mean, paper reference 55.67% 37.53%
Majority-choice reference 25.80% 41.47%

For bank 3701, Jev was called on 194/500 ReClor cases and 312/651 LogiQA cases. Relative to native, selection fixed 61 answers and damaged 1 on ReClor; it fixed 45 and damaged 6 on LogiQA. LogiQA performance remains below its 41.47% majority reference. These results establish a bounded improvement over this native baseline, not strong general intelligence or superiority over larger models.

The full report includes no-Jev local commit, unconditional proposal, cross-context donor verification, fixed experts, shared-adapter and full-bank controls, along with costs, statistical uncertainty and limitations. More candidates can improve accuracy at additional cost. A single native answer uses less compute than this two-generation method. Agreement saves API calls, not candidate generation. Hosted Jev cost, latency and undisclosed model size belong in system-level comparisons. No results from an unfinished broader benchmark campaign are claimed here.

Training and provenance

The two experts were trained in the earlier R37 bank study on 437 MuSR cases, and the routers were fitted on a separate 145 cases. Original Granite and Jev weights remained frozen. R38 reused those artifacts without further training or threshold tuning. See the training study.

artifact_manifest.json binds the two safetensors, no-Jev router, configuration, vendored runtime and transport to their SHA-256 digests. SHA256SUMS covers the complete public release inventory except itself. The _frozen source files retain their original bytes from source revision 366b078c07fa6573753d8f78df82b930bdf7d3b1; some shared modules include historical helpers unused by this runner. Source provenance is not a claim that every experiment in those modules is released here.

See VALIDATION.md for packaging checks, SECURITY.md for known advisories in the historical pinned dependencies and the supported loading scope, and NOTICE.md for external licenses and attribution. OVRLab code and expert artifacts are released under MIT. The separately downloaded IBM model remains under Apache 2.0; TypeSafe service terms and API charges apply independently. No IBM backbone weights, Jev weights, benchmark prose, or API credentials are redistributed in this repository.

Limitations and intended use

Use this release to study selective verification on English four-choice reasoning. Open-ended chat, multilingual reasoning, coding, long-context tasks and serving under concurrency are not validated here. Training-data exposure in upstream models is unknown. Generated explanations are not independently validated proofs. Three-bank research replication is not independent third-party reproduction. GPU/CPU numerical differences and future hosted-provider changes can affect outputs. FP32 precision and the supplied pinned dependencies are part of the tested contract; quantization or a different engine requires new validation.

Please cite this repository at its immutable revision and the linked manuscript when reporting results; an arXiv identifier has not been assigned in this release.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for OVRLab/granite-jev-1b-research

Adapter
(1)
this model

Dataset used to train OVRLab/granite-jev-1b-research