shrew-ocr-preview-lora

Final adapter (v0.3). The adapter in this repo is the v0.3 (generation E6) LoRA and remains valid for the v0.2 bucket base as documented below. shrew-ocr-preview v0.4 was trained on an intermediate merged base that is not published, so no standalone adapter is released for v0.4 and none will be for later versions β€” the released weights are now a full-weight fork of the Granite base (merged only: bf16, GPTQ-8bit, GGUF). Use the merged shrew-ocr-preview weights. See Lineage on the merged model card. This repo's weights and its v0.3 tag are unchanged.

LoRA adapter for shrew-ocr-preview β€” per-page document image β†’ structured JSON (metadata, summary, RAG-ready semantic chunks, figures/tables with bounding boxes and HTML), fine-tuned from ibm-granite/granite-vision-4.1-4b.

This repo contains only the adapter (generation E6 / release v0.3): r=256, Ξ±=512, dropout 0.05, language-model decoder only (q/k/v/o and gate/up/down projections across all 40 layers), trained 1 epoch / 1,343 steps on the dense/broadsheet-up-weighted mix, eval_loss 0.1408 (v0.2: 0.1421). The vision tower and projectors are untouched β€” all task adaptation lives in the language blocks.

Use this repo for adapter composition, inspecting the fine-tune delta, or continued training. For inference, use shrew-ocr-preview (merged bf16) or shrew-ocr-preview-GPTQ-8bit: adapter-path serving at rank 256 is slower (measured ~590 tok/s decode ceiling for base + adapter vs 850–1,400 tok/s for the merged INT8 variant on the same hardware). The reference serving pipeline for the merged variants is shrew-server (PDF in, structured JSON out; see the merged model card).

Applying the adapter

The base model's stock image tiling does not match the adapter's training. Set the bucket tile grids in both the model config and the image processor before inference; without this, output quality degrades severely:

from transformers import AutoModelForImageTextToText, AutoProcessor
from peft import PeftModel

PINPOINTS = [[1536, 1152], [2304, 1536], [3072, 2304], [1152, 1152]]

base = AutoModelForImageTextToText.from_pretrained(
    "ibm-granite/granite-vision-4.1-4b", dtype="bfloat16", trust_remote_code=True)
model = PeftModel.from_pretrained(base, "btbtyler09/shrew-ocr-preview-lora")

processor = AutoProcessor.from_pretrained(
    "ibm-granite/granite-vision-4.1-4b", trust_remote_code=True)
model.config.image_grid_pinpoints = PINPOINTS
processor.image_processor.image_grid_pinpoints = PINPOINTS

(The merged repos ship config.json/preprocessor_config.json with these values already set.)

The v0.3 model emits a 36-value section_type taxonomy (listed on the merged model card); consumers validating that field must accept it. All other usage requirements apply unchanged: system prompt set to the literal string structured_extraction, temperature 0, and glyph-routed page preparation into the bucket resolutions above. The prepare_page reference implementation, output schema, serving guidance, and streaming repetition guard are documented in the merged model card. Read it before deploying.

For vLLM adapter serving, pass --enable-lora --max-lora-rank 256; the config patch above still applies to the base checkpoint. The merged variants serve faster.

Results and limitations

Measured results on the OHR-Bench document-RAG corpus (8,561 pages; retrieval hit@5/MRR@10 vs human ground truth, MinerU, and PaddleOCR), together with known failure modes (dense broadsheet scans, non-Latin scripts, broadsheet reading order), are documented in the merged model card β€” the merged bf16 model is mathematically identical to base + this adapter.

Versions

variant precision size notes
shrew-ocr-preview bf16 7.5 GB reference quality (v0.4, merged only)
shrew-ocr-preview-GPTQ-8bit INT8 LM / bf16 vision 4.9 GB ~1.8Γ— serving throughput; v0.3 parity card vs bf16 PASS; serve with --dtype half
shrew-ocr-preview-GGUF Q8_0 or f16 LM / f16 vision 3.6–6.8 GB llama.cpp; full context per slot required (-c = N Γ— 32768)
shrew-ocr-preview-lora LoRA adapter (r=256, bf16) (this repo) 2.0 GB v0.3 adapter, final β€” applies only to the v0.2 bucket base; no v0.4 adapter

Changelog

v0.4 β€” no adapter released (see the note at the top); this repo is unchanged at its v0.3 tag.

v0.3 (2026-09-14, tag v0.3) β€” generation E6. New LoRA (same r=256 recipe, 1 epoch / 1,343 steps, eval_loss 0.1408 vs 0.1421) trained on the up-weighted dense/broadsheet slices with the open section_type taxonomy; merged bf16, GPTQ-8bit, GGUF and adapter pushed in lockstep. Promoted on six pre-registered criteria (retrieval paired PASS, product gates PASS, regions PASS, parity TIE, contract screens PASS, image surface flat). Loops on OmniDocBench newspapers 66 % β†’ 46 % in isolation; figure far-false-positives 156 β†’ 51. Requires shrew-server β‰₯ 0.3.14 for the 36-value section_type contract. Buckets / pinpoints unchanged from v0.2. Retrieval tables are now the exact-kNN read.

v0.2 β€” glyph-routed tile buckets (E4); first GPTQ-8bit and GGUF releases.

This is a preview: the merged repos update in place under their names as the model improves; this adapter repo is final at v0.3. Each push's commit message records the generation β€” pin a commit (revision=) for reproducibility.

Base model: ibm-granite/granite-vision-4.1-4b (Apache 2.0), trained with peft 0.18.1.

Downloads last month
25
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for btbtyler09/shrew-ocr-preview-lora

Adapter
(3)
this model