Instructions to use eulogik/TinyDoc-VLM-256M with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use eulogik/TinyDoc-VLM-256M with Transformers:
# Use a pipeline as a high-level helper # Warning: Pipeline type "visual-question-answering" is no longer supported in transformers v5. # You must load the model directly (see below) or downgrade to v4.x with: # pip install "transformers<5.0.0" from transformers import pipeline pipe = pipeline("visual-question-answering", model="eulogik/TinyDoc-VLM-256M")# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForSeq2SeqLM model = AutoModelForSeq2SeqLM.from_pretrained("eulogik/TinyDoc-VLM-256M", device_map="auto") - PEFT
How to use eulogik/TinyDoc-VLM-256M with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
TinyDoc-VLM-256M
⚠️ RETIRED RESEARCH CHECKPOINT — no performance claims
Measured 0.0% on OCRBench (n=1,000, all 10 categories) in 2026-09. The model degenerates under every prompt tried, including the exact training format (
"Extract document information: <image>"), with or without repetition controls; the LoRA adapter is worse (literal degeneration). The checkpoint is kept for reproducibility and research only.Earlier marketing numbers for this model (DocVQA ~65%, OCRBench ~60%) were never measured and have been withdrawn by the project. Evidence & artifacts:
evaluation/phase0/results/ocrbench_256m_full.*anddocs/BENCHMARKS.mdin the GitHub repo.
What actually ships
The working product of this project is the local grounded-extraction SDK
(pip install tinydoc) — schema-valid JSON, evidence bounding boxes, confidence
scores — running on free ollama:qwen2.5vl:3b, no API key:
| System | SROIE field F1 (n=100, same scorer) |
|---|---|
| TinyDoc pipeline (qwen2.5vl:3b) | 0.870 |
| PP-OCR + heuristics (free competitor) | 0.376 |
| Tesseract + regex | 0.227 |
All numbers recompute from committed artifacts
(python3 evaluation/phase0/recompute_scores.py).
from tinydoc.pipeline import ReceiptPipeline
pipe = ReceiptPipeline("auto") # free Ollama engine
doc = pipe.extract("receipt.jpg")
print(doc.fields, doc.schema_valid, doc.confidence)
for fr in doc.field_results:
print(fr.name, fr.value, fr.evidence["quote"], fr.evidence["bbox"])
About this checkpoint (kept for research)
Experimental 256M-parameter document VLM (2026-06 run), legacy 384px pipeline:
Image (384×384) → SigLIP-B/16 (93M) → PixelShuffle compressor ×3 (→64 tokens)
→ SmolLM2-135M decoder (30 layers, GQA, 8192 ctx) → multi-task heads
Verified facts (not claims):
- Vision tower is alive (features std≈1, cross-image cosine −0.57) — the dead-tower defect was specific to the separate 768px run.
- Processor
image_token_idmismatch was fixed in evaluation code; it does not explain the failures. - A 768px retrain (
TinyDoc-VLM-768-checkpoints) was abandoned: vision tower dead at init.
Loading (research use)
from tinydoc_vlm import TinyDocVLMForConditionalGeneration, TinyDocVLMProcessor
model = TinyDocVLMForConditionalGeneration.from_pretrained("eulogik/TinyDoc-VLM-256M")
processor = TinyDocVLMProcessor()
Links
| Resource | URL |
|---|---|
| GitHub (eval artifacts, SDK, decision record) | github.com/eulogik/TinyDoc-VLM |
| PyPI SDK | pypi.org/project/tinydoc |
| LoRA adapter (also retired) | eulogik/TinyDoc-VLM-LoRA |
| Benchmark write-up (measured only) | docs/BENCHMARKS.md |
License
Apache 2.0. Free for commercial use.
Built by eulogik
- Downloads last month
- 610