Instructions to use spele1100/instinct-vision-ni with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use spele1100/instinct-vision-ni with Transformers:
# Use a pipeline as a high-level helper # Warning: Pipeline type "visual-question-answering" is no longer supported in transformers v5. # You must load the model directly (see below) or downgrade to v4.x with: # pip install "transformers<5.0.0" from transformers import pipeline pipe = pipeline("visual-question-answering", model="spele1100/instinct-vision-ni")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("spele1100/instinct-vision-ni", device_map="auto") - Laya
How to use spele1100/instinct-vision-ni with Laya:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
instinct-vision-ni
A 0.8B vision-language model finetuned for live camera temporal reasoning โ change detection, change direction, and temporal order binding โ built on the laya/jev scoring architecture (single forward-pass option scoring, no autoregressive generation).
instinct-vision-ni is trained without any restricted-license data, making it usable
as a commercially-distributable checkpoint.
Training data
| Source | Examples | License | Notes |
|---|---|---|---|
| ScienceQA (derek-thomas/ScienceQA mirror, train split) | 5,000 | CC BY-SA 4.0 (mirror card tag verified 2026-09-26; original AllenAI ScienceQA is CC BY-SA 4.0) | general VQA retention |
| AI2D (lmms-lab-encoder/ai2d, test split โ only split published; first 250 rows excluded) | 2,821 | CC BY-SA 4.0 (AllenAI AI2D; HF mirror card carries no explicit tag โ caveat noted) | diagram QA. โ ๏ธ future ai2d evals against this split are contaminated for this model |
| Synthetic temporal pairs built by the laya-vision temporal-pair builder + supplement (self-made photos, PIL jitter, 896px) | 5,000 | Apache-2.0-licensed use granted by owner | delta/same/which ร full pair-type ร question-polarity grid |
| Total | 12,821 (30 dropped by max_len filter) |
Architecture & training
- Backbone: Qwen/Qwen3.5-0.8B (Apache 2.0); laya-vision scoring head, freeze=full
- lr 1e-4 head / 1e-5 backbone, bs 8, max_len 2048, AdamW8bit, 4,800 steps (~3 epochs), 10,200s on GB10
- Post-loop per-type temperature calibration (filtered 3.5k mix)
- Hardware: NVIDIA GB10 (SM121), 2026-09-26/27. driver script available in the laya-vision training repo
Results
| Test | Result |
|---|---|
| Multiframe temporal (8-task suite) | 8/8 (all @1.0 confidence) |
| Held-out jitter N=50 ร 6 tasks | 1.000 each |
| ScienceQA image-only N=143 | 84.6% |
| Speed 1-frame / 2-frame | 0.056 / 0.097 s |
License
Apache 2.0 (this checkpoint + code). Training-data licenses are disclosed above; ScienceQA and AI2D are CC BY-SA 4.0 (share-alike applies to those datasets). No COCO-derived images were used.
Usage
VLMAgent("/path/to/instinct-vision-ni", device="cuda", dtype="bf16") via laya-vision.
State dict {"images": [...]} for multi-frame; ~4 frames fit max_len 2048 at 896px.