TinyDoc-VLM-256M

⚠️ RETIRED RESEARCH CHECKPOINT — no performance claims

Measured 0.0% on OCRBench (n=1,000, all 10 categories) in 2026-09. The model degenerates under every prompt tried, including the exact training format ("Extract document information: <image>"), with or without repetition controls; the LoRA adapter is worse (literal degeneration). The checkpoint is kept for reproducibility and research only.

Earlier marketing numbers for this model (DocVQA ~65%, OCRBench ~60%) were never measured and have been withdrawn by the project. Evidence & artifacts: evaluation/phase0/results/ocrbench_256m_full.* and docs/BENCHMARKS.md in the GitHub repo.

What actually ships

The working product of this project is the local grounded-extraction SDK (pip install tinydoc) — schema-valid JSON, evidence bounding boxes, confidence scores — running on free ollama:qwen2.5vl:3b, no API key:

System SROIE field F1 (n=100, same scorer)
TinyDoc pipeline (qwen2.5vl:3b) 0.870
PP-OCR + heuristics (free competitor) 0.376
Tesseract + regex 0.227

All numbers recompute from committed artifacts (python3 evaluation/phase0/recompute_scores.py).

from tinydoc.pipeline import ReceiptPipeline

pipe = ReceiptPipeline("auto")           # free Ollama engine
doc = pipe.extract("receipt.jpg")
print(doc.fields, doc.schema_valid, doc.confidence)
for fr in doc.field_results:
    print(fr.name, fr.value, fr.evidence["quote"], fr.evidence["bbox"])

About this checkpoint (kept for research)

Experimental 256M-parameter document VLM (2026-06 run), legacy 384px pipeline:

Image (384×384) → SigLIP-B/16 (93M) → PixelShuffle compressor ×3 (→64 tokens)
→ SmolLM2-135M decoder (30 layers, GQA, 8192 ctx) → multi-task heads

Verified facts (not claims):

  • Vision tower is alive (features std≈1, cross-image cosine −0.57) — the dead-tower defect was specific to the separate 768px run.
  • Processor image_token_id mismatch was fixed in evaluation code; it does not explain the failures.
  • A 768px retrain (TinyDoc-VLM-768-checkpoints) was abandoned: vision tower dead at init.

Loading (research use)

from tinydoc_vlm import TinyDocVLMForConditionalGeneration, TinyDocVLMProcessor
model = TinyDocVLMForConditionalGeneration.from_pretrained("eulogik/TinyDoc-VLM-256M")
processor = TinyDocVLMProcessor()

Links

Resource URL
GitHub (eval artifacts, SDK, decision record) github.com/eulogik/TinyDoc-VLM
PyPI SDK pypi.org/project/tinydoc
LoRA adapter (also retired) eulogik/TinyDoc-VLM-LoRA
Benchmark write-up (measured only) docs/BENCHMARKS.md

License

Apache 2.0. Free for commercial use.


Built by eulogik

Downloads last month
610
Safetensors
Model size
0.3B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for eulogik/TinyDoc-VLM-256M

Adapters
1 model

Space using eulogik/TinyDoc-VLM-256M 1

Collections including eulogik/TinyDoc-VLM-256M