instinct-vision-ni

A 0.8B vision-language model finetuned for live camera temporal reasoning โ€” change detection, change direction, and temporal order binding โ€” built on the laya/jev scoring architecture (single forward-pass option scoring, no autoregressive generation).

instinct-vision-ni is trained without any restricted-license data, making it usable as a commercially-distributable checkpoint.

Training data

Source Examples License Notes
ScienceQA (derek-thomas/ScienceQA mirror, train split) 5,000 CC BY-SA 4.0 (mirror card tag verified 2026-09-26; original AllenAI ScienceQA is CC BY-SA 4.0) general VQA retention
AI2D (lmms-lab-encoder/ai2d, test split โ€” only split published; first 250 rows excluded) 2,821 CC BY-SA 4.0 (AllenAI AI2D; HF mirror card carries no explicit tag โ€” caveat noted) diagram QA. โš ๏ธ future ai2d evals against this split are contaminated for this model
Synthetic temporal pairs built by the laya-vision temporal-pair builder + supplement (self-made photos, PIL jitter, 896px) 5,000 Apache-2.0-licensed use granted by owner delta/same/which ร— full pair-type ร— question-polarity grid
Total 12,821 (30 dropped by max_len filter)

Architecture & training

  • Backbone: Qwen/Qwen3.5-0.8B (Apache 2.0); laya-vision scoring head, freeze=full
  • lr 1e-4 head / 1e-5 backbone, bs 8, max_len 2048, AdamW8bit, 4,800 steps (~3 epochs), 10,200s on GB10
  • Post-loop per-type temperature calibration (filtered 3.5k mix)
  • Hardware: NVIDIA GB10 (SM121), 2026-09-26/27. driver script available in the laya-vision training repo

Results

Test Result
Multiframe temporal (8-task suite) 8/8 (all @1.0 confidence)
Held-out jitter N=50 ร— 6 tasks 1.000 each
ScienceQA image-only N=143 84.6%
Speed 1-frame / 2-frame 0.056 / 0.097 s

License

Apache 2.0 (this checkpoint + code). Training-data licenses are disclosed above; ScienceQA and AI2D are CC BY-SA 4.0 (share-alike applies to those datasets). No COCO-derived images were used.

Usage

VLMAgent("/path/to/instinct-vision-ni", device="cuda", dtype="bf16") via laya-vision. State dict {"images": [...]} for multi-frame; ~4 frames fit max_len 2048 at 896px.

Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
0.9B params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for spele1100/instinct-vision-ni

Finetuned
(490)
this model