Glim 12B

Benchmark snapshot

Glim compared with open Jev-Omni, Gemma and Qwen

Public JevBench comparison

Glim 12B is unified v4. Jev-Omni and Glim use a trained decision head. Gemma and Qwen are unmodified base-model decision-scoring baselines: one forward pass, thinking disabled, candidate answer-token logits, without generated explanations or agent loops. They were not fine-tuned into dedicated decision models for these results. This measures constrained decision accuracy; it does not establish equal architectures, runtime efficiency or a native-generation leaderboard rank.

TypeSafe AI's hosted Jev is separate from the original open Jev-Omni checkpoint shown here. A matching TypeSafe score is pending; no score is inferred from other benchmark scopes. The planned Qwen LoRA candidate has no verified result yet.

Complete protocol, counts and limitations.

One complete merged Gemma 4 12B checkpoint, one general decision head, one processor and a Choice/Noul/Score adapter. No separate base-model or head repository needs assembling. This is a research release independent of TypeSafe AI. The supplied state/media and bounded answer options are the decision input; no agent loop or generated explanation is used.

Continued head training used 59,000 examples: all 47,680 general v2 examples, 7,962 GUI examples and 3,358 direct-visual-category SoccerNet examples. Our backbone stayed frozen; the upstream backbone already contains merged LoRA training. Training sources retain mixed terms. Source media, benchmark examples and labels are not redistributed. The upstream model's Apache-2.0 declaration applies to its artifacts; inspect source terms before commercial use. Wrapper code comes from the MIT-licensed project.

Measured results

Internal general test: v3 and v4 both answered 8,051/9,284 (86.72%). Labeled visual SoccerNet development: 414/847 to 419/847 (48.88% to 49.47%). That five-answer gain is on the epoch-selection set and is not a demonstrated broad visual gain. Usable public SoccerNet test, all categories: 235/496 to 237/496 (47.38% to 47.78%). Four malformed source-test labels are excluded; this is not an official 500-row challenge submission. Clean derived Mind2Web test: v3 852/1,084, v4 858/1,084. The correct control is guaranteed among four preselected named candidates; this is not official Mind2Web task success.

External benchmark table uses all 1,901 labeled BLINK validation questions and all 1,580 TempCompass MC / 2,453 yes-no test questions. Decider rows cover only eligible single-image BLINK questions, not the full benchmark; compare the same IDs, not different denominators. Original option order is preserved. Batch one, H100, BF16, SDPA, thinking disabled, one decision forward; ordinary VLMs score allowed letter-token logits. FLA 0.5.2 is installed for Qwen, but causal convolution uses the reference fallback. Latency includes preprocessing and first-use kernel compilation; these are eager-harness results, not each vendor's best production server.

Model Source suite Video representation Correct / total Micro accuracy Task macro
gemma12b blink_val sequence 1162/1901 61.13% 61.19%
gemma12b tempcompass_multi-choice sequence 1065/1580 67.41% 67.50%
gemma12b tempcompass_yes_no sequence 1773/2453 72.28% 71.98%
qwen4b blink_val sequence 1196/1901 62.91% 63.13%
qwen4b tempcompass_multi-choice sequence 1055/1580 66.77% 66.72%
qwen4b tempcompass_yes_no sequence 1751/2453 71.38% 70.95%
qwen9b blink_val sequence 1253/1901 65.91% 65.89%
qwen9b tempcompass_multi-choice sequence 1132/1580 71.65% 71.74%
qwen9b tempcompass_yes_no sequence 1842/2453 75.09% 74.87%
decider2b blink_val sequence 481/793 60.66% 61.64%
jev_general_v3 blink_val sequence 1028/1901 54.08% 54.37%
jev_upstream blink_val sequence 995/1901 52.34% 52.89%
jev_unified_v4 blink_val sequence 1026/1901 53.97% 54.31%
jev_general_v3 tempcompass_multi-choice sequence 1081/1580 68.42% 68.48%
jev_upstream tempcompass_multi-choice sequence 1075/1580 68.04% 68.11%
jev_unified_v4 tempcompass_multi-choice sequence 1082/1580 68.48% 68.53%
jev_general_v3 tempcompass_yes_no sequence 1751/2453 71.38% 71.07%
jev_upstream tempcompass_yes_no sequence 1722/2453 70.20% 70.06%
jev_unified_v4 tempcompass_yes_no sequence 1745/2453 71.14% 70.86%
jev_general_v3 blink_val native 1028/1901 54.08% 54.37%
jev_upstream blink_val native 995/1901 52.34% 52.89%
jev_unified_v4 blink_val native 1026/1901 53.97% 54.31%
jev_general_v3 tempcompass_multi-choice native 1057/1580 66.90% 67.06%
jev_upstream tempcompass_multi-choice native 1056/1580 66.84% 66.96%
jev_unified_v4 tempcompass_multi-choice native 1051/1580 66.52% 66.69%
jev_general_v3 tempcompass_yes_no native 1683/2453 68.61% 68.63%
jev_upstream tempcompass_yes_no native 1796/2453 73.22% 73.19%
jev_unified_v4 tempcompass_yes_no native 1673/2453 68.20% 68.20%
jev_general_v3 blink_val timeline 1028/1901 54.08% 54.37%
jev_upstream blink_val timeline 995/1901 52.34% 52.89%
jev_unified_v4 blink_val timeline 1026/1901 53.97% 54.31%
jev_general_v3 tempcompass_multi-choice timeline 1107/1580 70.06% 70.15%
jev_upstream tempcompass_multi-choice timeline 1108/1580 70.13% 70.21%
jev_unified_v4 tempcompass_multi-choice timeline 1094/1580 69.24% 69.34%
jev_general_v3 tempcompass_yes_no timeline 1765/2453 71.95% 71.90%
jev_upstream tempcompass_yes_no timeline 1781/2453 72.60% 72.65%
jev_unified_v4 tempcompass_yes_no timeline 1742/2453 71.02% 70.92%

Sequence and timestamp modes use the same 16 midpoint frames. Native video mode uses the processor's sampling and 70 soft tokens per frame; its sample indices and token budget can differ. It also incurs the harness's initial midpoint decoding, so use a separate production latency profile before claiming native-video speed. These scores are under the declared fast protocol and are not comparable without qualification to native-generation model-card/leaderboard scores. No SOTA or hidden-test rank is claimed. BLINK hidden test, full MVBench, Video-MME, captioning and current sealed JevBench composites were not evaluated in this release.

Three repeated GUI development screenshots were excluded from v4 selection, and two training screenshots were excluded from the primary test diagnostic. Exact-byte benchmark overlap with recorded general/GUI training media was zero; crops, transcoding, near duplicates and unknown upstream pretraining remain limitations. See dataset_audit.json, metrics.json and release.lock.json for provenance, per-task scores, failures and latency scopes.

Load on a remote CUDA GPU

Download unified_jev.py at the chosen release revision, inspect it, then use:

from unified_jev import load_unified_jev
model = load_unified_jev("ferdinandl007/glim-12b", revision="<exact release revision>")
result = model.decide_typed({
    "state": "The meeting starts at 10:00; it is 09:00.",
    "questions": {"started": {"type": "noul", "instructions": "Has it started?"}}
})

The classifier's predict(..., modality="image"|"video"|"audio", media=...) uses the upstream loader. Multiple typed questions currently require one classifier forward per question. Public loading defaults to anonymous read. Scores/probabilities are raw, uncalibrated head outputs. A 256-slot structure does not establish quality for 256-option decisions; upstream quality evidence is bounded to smaller option sets.

Training/benchmark code: vl-jev-modal, especially docs/UNIFIED_RELEASE_AND_BENCHMARKS.md. All artifacts were built and tested on Modal.

Paired comparisons and typed text diagnostics

Paired accuracy differences use 2,000 source-cluster bootstrap resamples; 95% intervals are exploratory and are not a leaderboard rank.

Comparison Suite N Accuracy difference 95% interval
jev_unified_v4/sequence versus jev_upstream/sequence/blink_val blink_val 1901 +1.63 pp [-0.42, +3.67] pp
jev_unified_v4/sequence versus jev_upstream/sequence/tempcompass_multi-choice tempcompass_multi-choice 1580 +0.44 pp [-1.23, +2.16] pp
jev_unified_v4/sequence versus jev_upstream/sequence/tempcompass_yes_no tempcompass_yes_no 2453 +0.94 pp [-0.62, +2.57] pp
jev_unified_v4/sequence versus jev_general_v3/sequence/blink_val blink_val 1901 -0.11 pp [-1.00, +0.74] pp
jev_unified_v4/sequence versus jev_general_v3/sequence/tempcompass_multi-choice tempcompass_multi-choice 1580 +0.06 pp [-0.76, +0.90] pp
jev_unified_v4/sequence versus jev_general_v3/sequence/tempcompass_yes_no tempcompass_yes_no 2453 -0.24 pp [-0.87, +0.35] pp
jev_unified_v4/sequence versus gemma12b/sequence/blink_val blink_val 1901 -7.15 pp [-9.63, -4.64] pp
jev_unified_v4/sequence versus gemma12b/sequence/tempcompass_multi-choice tempcompass_multi-choice 1580 +1.08 pp [-0.86, +2.94] pp
jev_unified_v4/sequence versus gemma12b/sequence/tempcompass_yes_no tempcompass_yes_no 2453 -1.14 pp [-2.80, +0.61] pp
jev_unified_v4/sequence versus qwen4b/sequence/blink_val blink_val 1901 -8.94 pp [-12.17, -5.95] pp
jev_unified_v4/sequence versus qwen4b/sequence/tempcompass_multi-choice tempcompass_multi-choice 1580 +1.71 pp [-0.93, +4.41] pp
jev_unified_v4/sequence versus qwen4b/sequence/tempcompass_yes_no tempcompass_yes_no 2453 -0.24 pp [-2.44, +2.02] pp
jev_unified_v4/sequence versus qwen9b/sequence/blink_val blink_val 1901 -11.94 pp [-14.63, -9.36] pp
jev_unified_v4/sequence versus qwen9b/sequence/tempcompass_multi-choice tempcompass_multi-choice 1580 -3.16 pp [-5.92, -0.52] pp
jev_unified_v4/sequence versus qwen9b/sequence/tempcompass_yes_no tempcompass_yes_no 2453 -3.95 pp [-6.06, -1.85] pp
jev_unified_v4/native versus jev_unified_v4/sequence/blink_val blink_val 1901 +0.00 pp [+0.00, +0.00] pp
jev_unified_v4/native versus jev_unified_v4/sequence/tempcompass_multi-choice tempcompass_multi-choice 1580 -1.96 pp [-4.18, +0.33] pp
jev_unified_v4/native versus jev_unified_v4/sequence/tempcompass_yes_no tempcompass_yes_no 2453 -2.94 pp [-4.80, -1.05] pp
jev_unified_v4/timeline versus jev_unified_v4/sequence/blink_val blink_val 1901 +0.00 pp [+0.00, +0.00] pp
jev_unified_v4/timeline versus jev_unified_v4/sequence/tempcompass_multi-choice tempcompass_multi-choice 1580 +0.76 pp [-0.77, +2.27] pp
jev_unified_v4/timeline versus jev_unified_v4/sequence/tempcompass_yes_no tempcompass_yes_no 2453 -0.12 pp [-1.64, +1.36] pp
jev_unified_v4/sequence versus decider2b/sequence/blink_val/eligible_single_image eligible_single_image 793 +3.03 pp [-0.63, +7.05] pp

Public JevBench uses the original 231 public tasks at source revision bb05a335bc809e61b20c0f745d25499a82b326fc, unmodified source correctness scoring, original label order and our native typed prompt. It does not evaluate the current sealed composite. No public benchmark tasks were used in our continued training.

Head Correct / total Micro accuracy Group macro Choice Noul Score exact class
upstream 197/231 85.28% 82.82% 85.61% 86.49% 77.78%
v3 206/231 89.18% 87.18% 88.49% 93.24% 77.78%
v4 205/231 88.74% 86.67% 88.49% 91.89% 77.78%

The inherited audio interface has not been accuracy-benchmarked in this release. Multi-question typed requests still require one forward per question. Raw Score expectations and Noul/Choice probabilities are not calibrated.

Reproducible source: vl-jev-modal at a7f737133d66. Public loading of the complete checkpoint and all three typed interfaces was verified on Modal; the artifact head SHA-256 is e0bfd32690d4656238e29e402b01772ca93e97602c12875b84f61793eef169c6.

Ollama compatibility

Glim 12B is the public name of this audited unified v4 checkpoint; weights and benchmark scores are unchanged. See OLLAMA_COMPATIBILITY.md for the verified runtime differences and publication path. systemone_adapter.py preserves the native head behind the shared request/answer shape, but is not a stock Ollama runner. No ollama pull or quantized parity claim is made. The new 117,330-example text corpus is not yet included in this checkpoint.

Downloads last month
57
Safetensors
Model size
12B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ferdinandl007/glim-12b

Finetuned
(3)
this model
Quantizations
1 model