instinct-vision-san

A 500M vision-language model finetuned for live camera temporal reasoning β€” change detection, change direction, and temporal order binding β€” built on the laya/jev scoring architecture (single forward-pass option scoring, no autoregressive generation). Trained without any restricted-license data, making it usable as a commercially-distributable checkpoint, and small enough for edge/NPU deployment.

Highlights

  • 500M parameters (SmolVLM2-500M-Video-Instruct backbone, Apache 2.0) β€” ~40% smaller than instinct-vision-ni while matching its temporal competency
  • Single forward-pass option scoring β€” no autoregressive decode; scores answer options in one pass
  • Multiframe temporal reasoning: 6/6 held-out jitter suites @1.000, stable across seeds (999/42/43)
  • ScienceQA image-only 85.3% (N=143) β€” slightly above the 0.8B ni model
  • ~25-frame context fits max_len 2048 thanks to SmolVLM2's token-efficient video encoding
  • Fast: 0.052 s 1-frame / 0.081 s 2-frame (GB10, bf16)
  • Edge-validated: RKNN fp16 port on RK3588 showed 10/10 agreement with the bf16 host model

Results

Test Result
Multiframe temporal, held-out jitter N=50 Γ— 6 suites 1.000 each (6/6 @1.000), stable across seeds 999/42/43
ScienceQA image-only (N=143) 85.3%
Speed, 1-frame / 2-frame (GB10, bf16) 0.052 s / 0.081 s
COCO general VQA (519 auto-generated MCQs, eval only) 83.0% overall
β€” count 33.7%
β€” attribute 98.0%

Comparison with instinct-vision-ni (0.8B, same clean-data recipe):

Test ni (0.8B) san (this model, 500M)
Multiframe temporal N=50 Γ— 6 suites 1.000 each 1.000 each
ScienceQA image-only N=143 84.6% 85.3%
Speed 1-frame / 2-frame 0.056 / 0.097 s 0.052 / 0.081 s

san matches ni on every temporal suite, slightly beats it on the ScienceQA spot-check and on attribute/VQA-style general questions, at ~60% of the parameter count.

Honest caveat: counting/numeric QA is weak (~34% on COCO count questions). For counting tasks, pair this model with an object detector.

Training data (full disclosure)

Source Examples License Notes
ScienceQA (derek-thomas/ScienceQA mirror, train split) 5,000 CC BY-SA 4.0 general VQA retention
AI2D (lmms-lab-encoder/ai2d, test split, first 250 rows excluded) 2,821 CC BY-SA 4.0 (mirror card carries no explicit tag β€” caveat noted) diagram QA; ⚠️ future AI2D evals on this split are contaminated for this model
Synthetic temporal pairs from own camera snapshots (PIL jitter, 896px) 5,000 self-made photos, Apache-2.0 use delta/same/which Γ— pair-type Γ— polarity grid; private frames used at train time only, never shipped
Total 12,821

Excluded by design: A-OKVQA, VQAv2, any COCO-image training set (COCO used for evaluation only), iconqa (NC), chartqa (GPL), VLFeedback/RichHF/CrisisMMD (NC).

Architecture & training

  • Backbone: HuggingFaceTB/SmolVLM2-500M-Video-Instruct (Apache 2.0), natively video/multi-frame
  • laya-vision scoring head, full-model finetune (head-only freeze was tried first and fails temporal order binding on this backbone: which-reversed 0/50 β€” the backbone must be finetuned at low lr to bind frame order)
  • lr 1e-4 head / 1e-5 backbone, batch size 8, max_len 2048, AdamW8bit, 4,800 steps (~3,100s on GB10)
  • Post-loop per-type temperature calibration (filtered 3.5k mix)

Usage

Requires the laya-vision package. See example.py for a runnable snippet.

from laya import VLMAgent

model = VLMAgent("spele1100/instinct-vision-san", device="cuda", dtype="bf16")

# Single-frame QA
result = model.score({"images": ["frame.jpg"], "prompt": "What changed?"})

# Multi-frame temporal reasoning β€” frames in temporal order
result = model.score({"images": ["f1.jpg", "f2.jpg"], "prompt": "Did the object appear or disappear?"})

State dict accepts {"images": [...]} for multi-frame input; ~25 frames fit max_len 2048.

License

Apache 2.0 (this checkpoint + code). Training-data licenses disclosed above; ScienceQA and AI2D are CC BY-SA 4.0 (share-alike applies to those datasets). No COCO-derived images were used at train time.

Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
0.5B params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for spele1100/instinct-vision-san