Instructions to use spele1100/instinct-vision-san with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use spele1100/instinct-vision-san with Transformers:
# Use a pipeline as a high-level helper # Warning: Pipeline type "visual-question-answering" is no longer supported in transformers v5. # You must load the model directly (see below) or downgrade to v4.x with: # pip install "transformers<5.0.0" from transformers import pipeline pipe = pipeline("visual-question-answering", model="spele1100/instinct-vision-san")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("spele1100/instinct-vision-san", device_map="auto") - Laya
How to use spele1100/instinct-vision-san with Laya:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
instinct-vision-san
A 500M vision-language model finetuned for live camera temporal reasoning β change detection, change direction, and temporal order binding β built on the laya/jev scoring architecture (single forward-pass option scoring, no autoregressive generation). Trained without any restricted-license data, making it usable as a commercially-distributable checkpoint, and small enough for edge/NPU deployment.
Highlights
- 500M parameters (SmolVLM2-500M-Video-Instruct backbone, Apache 2.0) β ~40% smaller than instinct-vision-ni while matching its temporal competency
- Single forward-pass option scoring β no autoregressive decode; scores answer options in one pass
- Multiframe temporal reasoning: 6/6 held-out jitter suites @1.000, stable across seeds (999/42/43)
- ScienceQA image-only 85.3% (N=143) β slightly above the 0.8B ni model
- ~25-frame context fits
max_len2048 thanks to SmolVLM2's token-efficient video encoding - Fast: 0.052 s 1-frame / 0.081 s 2-frame (GB10, bf16)
- Edge-validated: RKNN fp16 port on RK3588 showed 10/10 agreement with the bf16 host model
Results
| Test | Result |
|---|---|
| Multiframe temporal, held-out jitter N=50 Γ 6 suites | 1.000 each (6/6 @1.000), stable across seeds 999/42/43 |
| ScienceQA image-only (N=143) | 85.3% |
| Speed, 1-frame / 2-frame (GB10, bf16) | 0.052 s / 0.081 s |
| COCO general VQA (519 auto-generated MCQs, eval only) | 83.0% overall |
| β count | 33.7% |
| β attribute | 98.0% |
Comparison with instinct-vision-ni (0.8B, same clean-data recipe):
| Test | ni (0.8B) | san (this model, 500M) |
|---|---|---|
| Multiframe temporal N=50 Γ 6 suites | 1.000 each | 1.000 each |
| ScienceQA image-only N=143 | 84.6% | 85.3% |
| Speed 1-frame / 2-frame | 0.056 / 0.097 s | 0.052 / 0.081 s |
san matches ni on every temporal suite, slightly beats it on the ScienceQA spot-check and on attribute/VQA-style general questions, at ~60% of the parameter count.
Honest caveat: counting/numeric QA is weak (~34% on COCO count questions). For counting tasks, pair this model with an object detector.
Training data (full disclosure)
| Source | Examples | License | Notes |
|---|---|---|---|
| ScienceQA (derek-thomas/ScienceQA mirror, train split) | 5,000 | CC BY-SA 4.0 | general VQA retention |
| AI2D (lmms-lab-encoder/ai2d, test split, first 250 rows excluded) | 2,821 | CC BY-SA 4.0 (mirror card carries no explicit tag β caveat noted) | diagram QA; β οΈ future AI2D evals on this split are contaminated for this model |
| Synthetic temporal pairs from own camera snapshots (PIL jitter, 896px) | 5,000 | self-made photos, Apache-2.0 use | delta/same/which Γ pair-type Γ polarity grid; private frames used at train time only, never shipped |
| Total | 12,821 |
Excluded by design: A-OKVQA, VQAv2, any COCO-image training set (COCO used for evaluation only), iconqa (NC), chartqa (GPL), VLFeedback/RichHF/CrisisMMD (NC).
Architecture & training
- Backbone: HuggingFaceTB/SmolVLM2-500M-Video-Instruct (Apache 2.0), natively video/multi-frame
- laya-vision scoring head, full-model finetune (head-only freeze was tried first and fails temporal order binding on this backbone: which-reversed 0/50 β the backbone must be finetuned at low lr to bind frame order)
- lr 1e-4 head / 1e-5 backbone, batch size 8,
max_len2048, AdamW8bit, 4,800 steps (~3,100s on GB10) - Post-loop per-type temperature calibration (filtered 3.5k mix)
Usage
Requires the laya-vision package. See example.py for a runnable snippet.
from laya import VLMAgent
model = VLMAgent("spele1100/instinct-vision-san", device="cuda", dtype="bf16")
# Single-frame QA
result = model.score({"images": ["frame.jpg"], "prompt": "What changed?"})
# Multi-frame temporal reasoning β frames in temporal order
result = model.score({"images": ["f1.jpg", "f2.jpg"], "prompt": "Did the object appear or disappear?"})
State dict accepts {"images": [...]} for multi-frame input; ~25 frames fit max_len 2048.
License
Apache 2.0 (this checkpoint + code). Training-data licenses disclosed above; ScienceQA and AI2D are CC BY-SA 4.0 (share-alike applies to those datasets). No COCO-derived images were used at train time.
Model tree for spele1100/instinct-vision-san
Base model
HuggingFaceTB/SmolLM2-360M