Glim 12B
Benchmark snapshot
Glim 12B is unified v4. Jev-Omni and Glim use a trained decision head. Gemma and Qwen are unmodified base-model decision-scoring baselines: one forward pass, thinking disabled, candidate answer-token logits, without generated explanations or agent loops. They were not fine-tuned into dedicated decision models for these results. This measures constrained decision accuracy; it does not establish equal architectures, runtime efficiency or a native-generation leaderboard rank.
TypeSafe AI's hosted Jev is separate from the original open Jev-Omni checkpoint shown here. A matching TypeSafe score is pending; no score is inferred from other benchmark scopes. The planned Qwen LoRA candidate has no verified result yet.
Complete protocol, counts and limitations.
One complete merged Gemma 4 12B checkpoint, one general decision head, one processor and a Choice/Noul/Score adapter. No separate base-model or head repository needs assembling. This is a research release independent of TypeSafe AI. The supplied state/media and bounded answer options are the decision input; no agent loop or generated explanation is used.
Continued head training used 59,000 examples: all 47,680 general v2 examples, 7,962 GUI examples and 3,358 direct-visual-category SoccerNet examples. Our backbone stayed frozen; the upstream backbone already contains merged LoRA training. Training sources retain mixed terms. Source media, benchmark examples and labels are not redistributed. The upstream model's Apache-2.0 declaration applies to its artifacts; inspect source terms before commercial use. Wrapper code comes from the MIT-licensed project.
Measured results
Internal general test: v3 and v4 both answered 8,051/9,284 (86.72%). Labeled visual SoccerNet development: 414/847 to 419/847 (48.88% to 49.47%). That five-answer gain is on the epoch-selection set and is not a demonstrated broad visual gain. Usable public SoccerNet test, all categories: 235/496 to 237/496 (47.38% to 47.78%). Four malformed source-test labels are excluded; this is not an official 500-row challenge submission. Clean derived Mind2Web test: v3 852/1,084, v4 858/1,084. The correct control is guaranteed among four preselected named candidates; this is not official Mind2Web task success.
External benchmark table uses all 1,901 labeled BLINK validation questions and all 1,580 TempCompass MC / 2,453 yes-no test questions. Decider rows cover only eligible single-image BLINK questions, not the full benchmark; compare the same IDs, not different denominators. Original option order is preserved. Batch one, H100, BF16, SDPA, thinking disabled, one decision forward; ordinary VLMs score allowed letter-token logits. FLA 0.5.2 is installed for Qwen, but causal convolution uses the reference fallback. Latency includes preprocessing and first-use kernel compilation; these are eager-harness results, not each vendor's best production server.
| Model | Source suite | Video representation | Correct / total | Micro accuracy | Task macro |
|---|---|---|---|---|---|
| gemma12b | blink_val | sequence | 1162/1901 | 61.13% | 61.19% |
| gemma12b | tempcompass_multi-choice | sequence | 1065/1580 | 67.41% | 67.50% |
| gemma12b | tempcompass_yes_no | sequence | 1773/2453 | 72.28% | 71.98% |
| qwen4b | blink_val | sequence | 1196/1901 | 62.91% | 63.13% |
| qwen4b | tempcompass_multi-choice | sequence | 1055/1580 | 66.77% | 66.72% |
| qwen4b | tempcompass_yes_no | sequence | 1751/2453 | 71.38% | 70.95% |
| qwen9b | blink_val | sequence | 1253/1901 | 65.91% | 65.89% |
| qwen9b | tempcompass_multi-choice | sequence | 1132/1580 | 71.65% | 71.74% |
| qwen9b | tempcompass_yes_no | sequence | 1842/2453 | 75.09% | 74.87% |
| decider2b | blink_val | sequence | 481/793 | 60.66% | 61.64% |
| jev_general_v3 | blink_val | sequence | 1028/1901 | 54.08% | 54.37% |
| jev_upstream | blink_val | sequence | 995/1901 | 52.34% | 52.89% |
| jev_unified_v4 | blink_val | sequence | 1026/1901 | 53.97% | 54.31% |
| jev_general_v3 | tempcompass_multi-choice | sequence | 1081/1580 | 68.42% | 68.48% |
| jev_upstream | tempcompass_multi-choice | sequence | 1075/1580 | 68.04% | 68.11% |
| jev_unified_v4 | tempcompass_multi-choice | sequence | 1082/1580 | 68.48% | 68.53% |
| jev_general_v3 | tempcompass_yes_no | sequence | 1751/2453 | 71.38% | 71.07% |
| jev_upstream | tempcompass_yes_no | sequence | 1722/2453 | 70.20% | 70.06% |
| jev_unified_v4 | tempcompass_yes_no | sequence | 1745/2453 | 71.14% | 70.86% |
| jev_general_v3 | blink_val | native | 1028/1901 | 54.08% | 54.37% |
| jev_upstream | blink_val | native | 995/1901 | 52.34% | 52.89% |
| jev_unified_v4 | blink_val | native | 1026/1901 | 53.97% | 54.31% |
| jev_general_v3 | tempcompass_multi-choice | native | 1057/1580 | 66.90% | 67.06% |
| jev_upstream | tempcompass_multi-choice | native | 1056/1580 | 66.84% | 66.96% |
| jev_unified_v4 | tempcompass_multi-choice | native | 1051/1580 | 66.52% | 66.69% |
| jev_general_v3 | tempcompass_yes_no | native | 1683/2453 | 68.61% | 68.63% |
| jev_upstream | tempcompass_yes_no | native | 1796/2453 | 73.22% | 73.19% |
| jev_unified_v4 | tempcompass_yes_no | native | 1673/2453 | 68.20% | 68.20% |
| jev_general_v3 | blink_val | timeline | 1028/1901 | 54.08% | 54.37% |
| jev_upstream | blink_val | timeline | 995/1901 | 52.34% | 52.89% |
| jev_unified_v4 | blink_val | timeline | 1026/1901 | 53.97% | 54.31% |
| jev_general_v3 | tempcompass_multi-choice | timeline | 1107/1580 | 70.06% | 70.15% |
| jev_upstream | tempcompass_multi-choice | timeline | 1108/1580 | 70.13% | 70.21% |
| jev_unified_v4 | tempcompass_multi-choice | timeline | 1094/1580 | 69.24% | 69.34% |
| jev_general_v3 | tempcompass_yes_no | timeline | 1765/2453 | 71.95% | 71.90% |
| jev_upstream | tempcompass_yes_no | timeline | 1781/2453 | 72.60% | 72.65% |
| jev_unified_v4 | tempcompass_yes_no | timeline | 1742/2453 | 71.02% | 70.92% |
Sequence and timestamp modes use the same 16 midpoint frames. Native video mode uses the processor's sampling and 70 soft tokens per frame; its sample indices and token budget can differ. It also incurs the harness's initial midpoint decoding, so use a separate production latency profile before claiming native-video speed. These scores are under the declared fast protocol and are not comparable without qualification to native-generation model-card/leaderboard scores. No SOTA or hidden-test rank is claimed. BLINK hidden test, full MVBench, Video-MME, captioning and current sealed JevBench composites were not evaluated in this release.
Three repeated GUI development screenshots were excluded from v4 selection, and two training screenshots were excluded from the primary test diagnostic. Exact-byte benchmark overlap with recorded general/GUI training media was zero; crops, transcoding, near duplicates and unknown upstream pretraining remain limitations. See dataset_audit.json, metrics.json and release.lock.json for provenance, per-task scores, failures and latency scopes.
Load on a remote CUDA GPU
Download unified_jev.py at the chosen release revision, inspect it, then use:
from unified_jev import load_unified_jev
model = load_unified_jev("ferdinandl007/glim-12b", revision="<exact release revision>")
result = model.decide_typed({
"state": "The meeting starts at 10:00; it is 09:00.",
"questions": {"started": {"type": "noul", "instructions": "Has it started?"}}
})
The classifier's predict(..., modality="image"|"video"|"audio", media=...) uses the upstream loader. Multiple typed questions currently require one classifier forward per question. Public loading defaults to anonymous read. Scores/probabilities are raw, uncalibrated head outputs. A 256-slot structure does not establish quality for 256-option decisions; upstream quality evidence is bounded to smaller option sets.
Training/benchmark code: vl-jev-modal, especially docs/UNIFIED_RELEASE_AND_BENCHMARKS.md. All artifacts were built and tested on Modal.
Paired comparisons and typed text diagnostics
Paired accuracy differences use 2,000 source-cluster bootstrap resamples; 95% intervals are exploratory and are not a leaderboard rank.
| Comparison | Suite | N | Accuracy difference | 95% interval |
|---|---|---|---|---|
| jev_unified_v4/sequence versus jev_upstream/sequence/blink_val | blink_val | 1901 | +1.63 pp | [-0.42, +3.67] pp |
| jev_unified_v4/sequence versus jev_upstream/sequence/tempcompass_multi-choice | tempcompass_multi-choice | 1580 | +0.44 pp | [-1.23, +2.16] pp |
| jev_unified_v4/sequence versus jev_upstream/sequence/tempcompass_yes_no | tempcompass_yes_no | 2453 | +0.94 pp | [-0.62, +2.57] pp |
| jev_unified_v4/sequence versus jev_general_v3/sequence/blink_val | blink_val | 1901 | -0.11 pp | [-1.00, +0.74] pp |
| jev_unified_v4/sequence versus jev_general_v3/sequence/tempcompass_multi-choice | tempcompass_multi-choice | 1580 | +0.06 pp | [-0.76, +0.90] pp |
| jev_unified_v4/sequence versus jev_general_v3/sequence/tempcompass_yes_no | tempcompass_yes_no | 2453 | -0.24 pp | [-0.87, +0.35] pp |
| jev_unified_v4/sequence versus gemma12b/sequence/blink_val | blink_val | 1901 | -7.15 pp | [-9.63, -4.64] pp |
| jev_unified_v4/sequence versus gemma12b/sequence/tempcompass_multi-choice | tempcompass_multi-choice | 1580 | +1.08 pp | [-0.86, +2.94] pp |
| jev_unified_v4/sequence versus gemma12b/sequence/tempcompass_yes_no | tempcompass_yes_no | 2453 | -1.14 pp | [-2.80, +0.61] pp |
| jev_unified_v4/sequence versus qwen4b/sequence/blink_val | blink_val | 1901 | -8.94 pp | [-12.17, -5.95] pp |
| jev_unified_v4/sequence versus qwen4b/sequence/tempcompass_multi-choice | tempcompass_multi-choice | 1580 | +1.71 pp | [-0.93, +4.41] pp |
| jev_unified_v4/sequence versus qwen4b/sequence/tempcompass_yes_no | tempcompass_yes_no | 2453 | -0.24 pp | [-2.44, +2.02] pp |
| jev_unified_v4/sequence versus qwen9b/sequence/blink_val | blink_val | 1901 | -11.94 pp | [-14.63, -9.36] pp |
| jev_unified_v4/sequence versus qwen9b/sequence/tempcompass_multi-choice | tempcompass_multi-choice | 1580 | -3.16 pp | [-5.92, -0.52] pp |
| jev_unified_v4/sequence versus qwen9b/sequence/tempcompass_yes_no | tempcompass_yes_no | 2453 | -3.95 pp | [-6.06, -1.85] pp |
| jev_unified_v4/native versus jev_unified_v4/sequence/blink_val | blink_val | 1901 | +0.00 pp | [+0.00, +0.00] pp |
| jev_unified_v4/native versus jev_unified_v4/sequence/tempcompass_multi-choice | tempcompass_multi-choice | 1580 | -1.96 pp | [-4.18, +0.33] pp |
| jev_unified_v4/native versus jev_unified_v4/sequence/tempcompass_yes_no | tempcompass_yes_no | 2453 | -2.94 pp | [-4.80, -1.05] pp |
| jev_unified_v4/timeline versus jev_unified_v4/sequence/blink_val | blink_val | 1901 | +0.00 pp | [+0.00, +0.00] pp |
| jev_unified_v4/timeline versus jev_unified_v4/sequence/tempcompass_multi-choice | tempcompass_multi-choice | 1580 | +0.76 pp | [-0.77, +2.27] pp |
| jev_unified_v4/timeline versus jev_unified_v4/sequence/tempcompass_yes_no | tempcompass_yes_no | 2453 | -0.12 pp | [-1.64, +1.36] pp |
| jev_unified_v4/sequence versus decider2b/sequence/blink_val/eligible_single_image | eligible_single_image | 793 | +3.03 pp | [-0.63, +7.05] pp |
Public JevBench uses the original 231 public tasks at source revision bb05a335bc809e61b20c0f745d25499a82b326fc, unmodified source correctness scoring, original label order and our native typed prompt. It does not evaluate the current sealed composite. No public benchmark tasks were used in our continued training.
| Head | Correct / total | Micro accuracy | Group macro | Choice | Noul | Score exact class |
|---|---|---|---|---|---|---|
| upstream | 197/231 | 85.28% | 82.82% | 85.61% | 86.49% | 77.78% |
| v3 | 206/231 | 89.18% | 87.18% | 88.49% | 93.24% | 77.78% |
| v4 | 205/231 | 88.74% | 86.67% | 88.49% | 91.89% | 77.78% |
The inherited audio interface has not been accuracy-benchmarked in this release. Multi-question typed requests still require one forward per question. Raw Score expectations and Noul/Choice probabilities are not calibrated.
Reproducible source: vl-jev-modal at a7f737133d66. Public loading of the complete checkpoint and all three typed interfaces was verified on Modal; the artifact head SHA-256 is e0bfd32690d4656238e29e402b01772ca93e97602c12875b84f61793eef169c6.
Ollama compatibility
Glim 12B is the public name of this audited unified v4 checkpoint; weights and benchmark scores are unchanged. See OLLAMA_COMPATIBILITY.md for the verified runtime differences and publication path. systemone_adapter.py preserves the native head behind the shared request/answer shape, but is not a stock Ollama runner. No ollama pull or quantized parity claim is made. The new 117,330-example text corpus is not yet included in this checkpoint.
- Downloads last month
- 57