shrew-ocr-preview
shrew-ocr-preview converts one document page image per request into a single JSON object containing document metadata, a summary, self-contained semantic chunks (sized for RAG ingestion, not raw OCR lines), and figures/tables with bounding boxes and HTML. A text modality accepts HTML/markdown/plain-text input and produces the same output schema.
Fine-tuned from ibm-granite/granite-vision-4.1-4b (merged weights, bf16; v0.4 — see Changelog and Lineage). GPTQ-8bit and GGUF conversions are released alongside this model for faster serving.
Preview release. Works well on mainstream printed documents (papers, reports, filings, manuals). Known failure modes are listed under Limitations; measured results under Results. Weights are updated in place under this name — pin a commit (
revision=) or the release tag (v0.4) for reproducibility.
Output schema
One request = one page. The model returns exactly one JSON object, five keys always present:
{
"metadata": {"title", "authors" (list), "organization", "year", "doc_type"} — null where unknown,
"summary": str | null,
"semantic_chunks": [{"chunk_id", "title", "content",
"section_type" ∈ the 36-value taxonomy below}],
"figures": [{"bbox": [x0,y0,x1,y1] | null, "caption", "description"}],
"tables": [{"html": "<table>…", "bbox": [...] | null, "caption", "description"}]
}
Bounding boxes are xyxy on a 0–1000 normalized grid over the page image. In text modality, bboxes
are null.
section_type (v0.3) is one of 36 trained values — document structure (abstract, introduction,
methodology, results, discussion, conclusion, appendix, technical_content), generic page
text (body, table_text, list), newspaper/magazine regions (news_article, news_brief,
news_analysis, feature_article, opinion, commentary, editorial, preview, preview_list,
headline, header, masthead, title, footer, index, photo_caption, photo_feature,
stat_box, sidebar, advertisement, legal_notice, official_document, obituary,
weather_box) and the fallback other. Consumers that validate section_type must accept this
set (v0.2 emitted only the first eight); shrew-server ≥ 0.3.14 does, and folds any stray label
onto it instead of failing the page.
Usage
Recommended path: shrew-server (MIT, pin tag
0.4.0 or later — earlier servers do not ship the v0.4 decoding recipe), the reference server for this model. It implements the model's entire input contract
server-side — glyph-routed bucket preprocessing, the structured_extraction request shape, the v0.4
decoding recipe (sampled first pass, one un-enforced retry, then fallback), the streaming repetition guard, schema/coercion
gates, and multi-page assembly. POST a PDF, receive structured JSON. It does not serve the model
itself; point it at an OpenAI-compatible endpoint (vLLM, below):
vllm serve btbtyler09/shrew-ocr-preview --trust-remote-code \
--served-model-name shrew-ocr-preview \
--max-model-len 32768 --limit-mm-per-prompt '{"image":1}' --no-enable-prefix-caching
VLM_URL=http://localhost:8000 VLM_MODEL=shrew-ocr-preview shrew serve
curl -X POST localhost:8080/v1/convert -F file=@doc.pdf -F pipeline_mode=structured
Full instructions, including a Docker Compose quickstart, are in the repo README under "Using with shrew-ocr-preview (recommended)".
For direct integration without shrew-server, the requirements below define the input contract. Deviating from any of these degrades output quality:
1. System prompt. Set the system prompt to the literal string structured_extraction. Do not
send instruction text; the model was trained on this fixed prompt only.
2. Decoding (v0.4 recipe). The measured results below were produced with this recipe, and it is what the reference server ships:
- first pass:
temperature0.4,presence_penalty0.3,max_tokens20000, no top_p / top_k; - if the first pass fails (repetition-guard abort, truncation, invalid or off-schema JSON): one retry at the same settings with a fresh seed, without grammar enforcement;
- if the retry also fails: hand the page to a fallback extractor if one is configured; otherwise record a failed page. Do not loop further.
Greedy decoding (temperature 0, the v0.3 recommendation) still produces valid output but loses
on hard pages: on a 228-page holdout of dense Chinese newspaper scans the first-pass clean rate
was 91.7 % sampled at 0.4 (majority of three seeds) vs 75.0 % greedy for these weights, with text
fidelity unchanged (paired recall against an OCR reference, inside a pre-registered 0.05 bound).
Sampling makes borderline pages seed-dependent (about a quarter of hard pages flip across seeds);
the reference server derives a fixed per-page seed so reruns reproduce. The un-enforced retry was
chosen by measurement: on 65 first-pass failures it recovered 69 % of pages vs 15 % for a
grammar-enforced retry at the same temperature, because enforcement makes these pages loop under
the grammar instead of finishing. generation_config.json carries do_sample, temperature 0.4
and max_new_tokens 20000; presence_penalty is not a generate() field, so set it in the
serving request. Serve with context length ≥ 32768; dense pages need room for both the image
tokens and a long completion.
3. Input resolution ("buckets"). Resize each page image to one of three portrait tile grids, selected by the page's measured glyph height (target ~10 px after resize). Training used exactly this routing. Reference implementation:
import cv2, statistics
import numpy as np
from PIL import Image
BUCKETS = [("B1", (1152, 1536)), ("B2", (1536, 2304)), ("B3", (2304, 3072))]
SQUARE = ("B0", (1152, 1152)) # square-ish inputs only (e.g. table crops)
def glyph_height(img, max_side=2600):
"""Median connected-component height in native px — the routing signal."""
W, H = img.size
s = min(1.0, max_side / max(W, H))
im = img.convert("L")
if s < 1.0:
im = im.resize((int(W * s), int(H * s)), Image.BOX)
g = cv2.adaptiveThreshold(np.asarray(im), 255, cv2.ADAPTIVE_THRESH_GAUSSIAN_C,
cv2.THRESH_BINARY_INV, 31, 10)
n, _, stats, _ = cv2.connectedComponentsWithStats(g, connectivity=8)
hs = [stats[i][3] for i in range(1, n)
if 2 <= stats[i][3] <= 60 and 1 <= stats[i][2] <= 60 and stats[i][4] >= 4
and 0.08 <= stats[i][2] / max(stats[i][3], 1) <= 6.0]
return statistics.median(hs) / max(s, 1e-6) if len(hs) >= 50 else None
def prepare_page(img, target=10.0):
"""Route to the smallest bucket that reaches ~10px effective glyph height, then enhance."""
w, h = img.size
if h and 0.9 <= w / h <= 1.15:
bw, bh = SQUARE[1]
else:
g = glyph_height(img)
bw, bh = BUCKETS[1][1] # default when unmeasurable
if g:
for _, (cw, ch) in BUCKETS:
if g * min(cw / w, ch / h) >= target * 0.95:
bw, bh = cw, ch
break
else:
bw, bh = BUCKETS[-1][1]
s = min(bw / w, bh / h)
fit = img.resize((round(w * s), round(h * s)), Image.LANCZOS)
gray = np.asarray(fit.convert("L"))
e = cv2.createCLAHE(clipLimit=2.0, tileGridSize=(8, 8)).apply(gray) # CLAHE
blur = cv2.GaussianBlur(e, (0, 0), 1.2)
e = cv2.addWeighted(e, 1.8, blur, -0.8, 0) # unsharp
return Image.fromarray(e).convert("RGB")
The checkpoint's config.json and preprocessor_config.json carry the matching
image_grid_pinpoints. Do not remove or modify them; the tile packing must match training.
Serving with vLLM
vllm serve /path/to/shrew-ocr-preview \
--trust-remote-code --dtype bfloat16 \
--served-model-name shrew-ocr-preview \
--max-model-len 32768 --limit-mm-per-prompt '{"image":1}' \
--no-enable-prefix-caching
Scaling note: the model is small (2–5 GB weights). For batch serving on multi-GPU hosts,
data-parallel replicas (--data-parallel-size N) outperform tensor parallelism substantially
(+54% measured on a 4-GPU node) — prefer DP unless a single GPU cannot hold the weights. On
memory-constrained GPUs keep --max-num-batched-tokens at 2048 or below: the vision encoder
batches image tiles, and large prefill budgets can OOM the tower on high-tile pages.
Request shape (OpenAI-compatible):
{
"model": "shrew-ocr-preview",
"temperature": 0.4,
"presence_penalty": 0.3,
"max_tokens": 20000,
"messages": [
{"role": "system", "content": "structured_extraction"},
{"role": "user", "content": [
{"type": "image_url", "image_url": {"url": "data:image/png;base64,<prepare_page output>"}},
{"type": "text", "text": "Extract the structured representation of this document page."}]}
]
}
Recommended: streaming repetition guard. On hard pages the failure mode is degenerate
repetition, not plausible-but-wrong output. Stream the completion, compute
len(window)/len(zlib.compress(window)) over a trailing ~2,000-character window every ~800
characters, and abort after 2 consecutive windows above ~15. Clean pages measure ~2; looping
output exceeds 25.
Penalty parameters (measured: 900 first-pass runs, 150 stratified pages × 6 decode arms on the
production stack): presence_penalty 0.3 at first pass raised the first-pass success rate (valid
JSON passing all schema and degeneration gates, no retry) 0.880→0.887 with extraction precision
flat (0.949→0.949 median 10-gram precision vs ground truth), and is the reference server's
first-pass default; 0.3–0.6 both measured fidelity-neutral on healthy pages. In a separate rescue
experiment, a penalized retry (presence_penalty 0.6) recovered 12/12 sampled loop-failed
Latin-script broadsheet pages with fidelity flat; dense non-Latin broadsheets did not recover at
any penalty (see Limitations — that class is a training gap, not a decoding one). Do not use
frequency or repetition penalties or no_repeat_ngram_size: ngram blocking suppresses JSON
tokens that must repeat ("bbox": [); frequency penalties accumulate with each repeated
occurrence, and repetition penalties apply to every previously-seen token — both degrade required
schema tokens in long structured outputs. Do not use grammar-constrained (schema-enforced)
decoding at first pass: it degrades table transcription severely (table one-shot 0.975→0.225),
primarily by overrunning the token budget mid-table (grammar-constrained decoding transcribes
exhaustively), with degraded table fidelity (0.483→0.280) on the pages that do finish; the
v0.4 reference server does not enforce on the retry either — an enforced retry was measured to recover far fewer pages (see Decoding). One caveat: retry rescue was validated on Latin-script pages only — for pages whose text
the vision tower cannot resolve, a penalized retry can convert a detectable loop into plausible
hallucination (observed: a penalized retry on an unsupported-script page produced fluent output
with zero n-gram overlap with the page), so validate retry output with the same window check and
schema gates and mark it lower-confidence downstream.
Which borderline pages loop varies with serving configuration (kernel paths, compilation settings); the overall loop rate does not. Compare deployments by loop rate over a fixed page set, not by which pages failed.
Text modality
The same model accepts born-digital text — HTML (emails, filings), markdown, source code, plain
text — and returns the same 5-key output schema with figures[].bbox and tables[].bbox null.
Do not route OCR output of scanned pages here; scanned pages go through the image modality.
Request shape: same envelope, with the raw content as a single text part in place of the image. Send the content as-is — no wrapper, no instructions, no cleaning:
{
"model": "shrew-ocr-preview",
"temperature": 0.4,
"presence_penalty": 0.3,
"max_tokens": 12000,
"messages": [
{"role": "system", "content": "structured_extraction"},
{"role": "user", "content": [{"type": "text", "text": "<raw HTML / markdown / plain text>"}]}
]
}
Input sizing: send 2,000–9,000 characters (~500–2,500 tokens) per request; treat 13,000 characters as the ceiling (training inputs never exceeded it). Split longer documents at structural boundaries (headings, sections) and send one request per section. Each request in the recommended range yields roughly 2–6 semantic chunks (median chunk ~820 characters).
Results — OHR-Bench document RAG
Measured on the OHR-Bench corpus
(ICCV 2025): 1,261 PDFs / 8,561 pages across 7 domains (textbook, law, finance, newspaper, manual,
academic, administration). Every page runs through our full production path (rasterize → bucket
routing → model → schema gates → assembly); each system's structured output is chunked under the
same budget, embedded with nvidia/llama-nemotron-embed-vl-1b-v2 ("nemotron-vl"), and scored as
retrieval hit@5 / MRR@10 over OHR-Bench's ~8.5k
human-verified Q&A pairs. These are our own retrieval-harness measurements, not official OHR-Bench
generation (LCS/F1) numbers. gt is retrieval over OHR-Bench's human ground-truth structured
data; MinerU and PaddleOCR outputs were run through the identical chunking and indexing.
Text retrieval, hit@5 / MRR@10 by evidence type — exact k-nearest-neighbour search over the embeddings (v0.2's table was read through an approximate HNSW index; the exact read is the reference from v0.3 on, so the rows are not directly comparable to the v0.2 card). Best per row in bold:
| evidence type | gt (human) | MinerU | PaddleOCR | shrew v0.4 (bf16) | shrew v0.3 weights, same server |
|---|---|---|---|---|---|
| plain text | .976 / .896 | .898 / .813 | .956 / .878 | .959 / .873 | .960 / .874 |
| multi-evidence | .978 / .883 | .956 / .873 | .963 / .854 | .970 / .864 | .970 / .855 |
| table | .958 / .842 | .927 / .795 | .921 / .795 | .918 / .799 | .913 / .787 |
| formula | .969 / .886 | .925 / .835 | .933 / .858 | .949 / .859 | .944 / .848 |
| chart | .878 / .751 | .659 / .540 | .668 / .520 | .765 / .630 | .757 / .636 |
| vision | .794 / .580 | .634 / .486 | .744 / .583 | .769 / .617 | .761 / .607 |
| reading order† | .930 / .839 | .942 / .866 | .941 / .848 | .754 / .692 | .476 / .450 |
| all 8,694 queries | .908 / .809 | .884 / .785 |
Both shrew columns are the same reference server and harness; the v0.3 column is the v0.3 weights re-parsed with their shipped greedy recipe (the v0.3 card's own table was read through an earlier server, which scored the same weights .844 pooled and .098 on reading order — most of that difference is server assembly, not the model). Paired over all queries, v0.4 is better than v0.3 (394 queries hit in the top 5 for v0.4 only vs 180 for v0.3 only; two-sided sign test p = 2×10⁻¹⁹). The entire margin is reading order on dense Chinese newspaper scans (225 vs 25); on the other 8,003 queries the two are indistinguishable (197 vs 175, p = 0.28) and no evidence type regressed. Two more paired reads separate the ingredients: the v0.4 decoding recipe alone, applied to the v0.3 weights, lifts pooled hit@5 .884 → .896 (274 vs 169, p = 7×10⁻⁷); the v0.4 weights at the same recipe add .896 → .908 (328 vs 219, p = 4×10⁻⁶).
† OHR-Bench draws reading-order queries almost entirely from dense broadsheet newspaper scans (Chinese). These were a repetition-loop failure class through v0.3; v0.4 parses most of them (pages failed on these documents 3.8 % vs 11.8 % for the v0.3 weights on the same server), and the remaining gap to MinerU is character accuracy on dense CJK, not loops.
Figure/table localization vs our own frozen human-annotated gold subset — 551 corpus pages / 1,100 boxes, not an OHR-Bench artifact (greedy match at IoU ≥ 0.5):
| arm | figure recall@0.5 | figure mean IoU | table recall@0.5 | table mean IoU | far false positives |
|---|---|---|---|---|---|
| v0.3 bf16 | 0.644 | 0.803 | 0.603 | 0.806 | 51 |
| v0.2 bf16 | 0.617 | 0.801 | 0.601 | 0.805 | 156 |
Not re-measured for v0.4 (the v0.3 rows are the last measurement). The page-level image surface was: over the 238 vision queries, cropped-figure retrieval hit@5 .193 (v0.3 weights, same server: .231) and MRR@10 .102 (.096) — flat within the noise of a 238-query set.
Reliability (reference server, 8,561 pages, v0.4 recipe): 99.3 % of pages produced valid schema-complete JSON on the first pass, 32 more were rescued by the single un-enforced retry, and 27 pages (0.32 %) failed — these are the pages a deployment would hand to a fallback extractor. The v0.3 weights on the same server with their shipped greedy recipe: 86 failed pages (1.00 %), 28 rescued by a greedy re-roll and 33 by the grammar-enforced retry; the v0.3 card measured 2.53 % hard failures under an earlier server. Failures still terminate as repetition-guard aborts or schema rejections rather than silent bad output.
Limitations
- Difficult documents. Dense broadsheet scans (historical newspapers), low-resolution scans of dense layouts, and pages whose text the vision tower cannot resolve can produce repetition loops instead of output. The streaming guard above converts these into fast, detectable failures. Work on this class is ongoing.
- Non-Latin scripts. Chinese print improved substantially in v0.4 (dense newspaper scans now parse; see Results), but character accuracy on dense CJK remains below Latin print. Cyrillic, Arabic and handwriting were not evaluated for this release — treat them as unsupported.
- Bounding boxes are model-supervised. Figure/table geometry is trained from model-generated labels with automated repair; boxes are generally tight but can under- or over-shoot on unusual layouts. Pad boxes outward slightly when cropping; do not treat edges as pixel-exact.
- One page per request. The model has no cross-page state; feed multi-page documents page by page and assemble downstream.
- Reading order on dense broadsheets is no longer at the failure floor (.754 hit@5 vs .098 in v0.3) but still trails the ground-truth and MinerU columns (see Results).
Lineage
shrew-ocr-preview began as a LoRA adapter on ibm-granite/granite-vision-4.1-4b. Since v0.2 each release has been a merge on top of the previous one — the glyph-routed bucket base (v0.2), the E6 fine-tune merged onto it (v0.3), an intermediate full-epoch merge that was not published, and the reinforcement-learning stage merged onto that (v0.4) — so the released weights are now a full-weight fork of the Granite base rather than an adapter that can be applied to it. The vision tower is no longer identical to the base: the attention layers of its upper encoder blocks and the vision-to-language projectors were fine-tuned in the unpublished epoch, and v0.4 inherits them. As of v0.4, releases are merged weights only: bf16, GPTQ-8bit and GGUF. No further standalone LoRA adapters will be published; the v0.3 adapter in shrew-ocr-preview-lora is the last one and applies only to the v0.2 bucket base.
Versions
| variant | precision | size | notes |
|---|---|---|---|
| shrew-ocr-preview | bf16 (this repo) | 7.5 GB | reference quality (v0.4) |
| shrew-ocr-preview-GPTQ-8bit | INT8 LM / bf16 vision | 4.9 GB | ~1.8× serving throughput; v0.4 smoke through the reference server PASS; serve with --dtype half |
| shrew-ocr-preview-GGUF | Q8_0 or f16 LM / f16 vision | 3.6–6.8 GB | llama.cpp; full context per slot required (-c = N × 32768) |
| shrew-ocr-preview-lora | LoRA adapter (r=256, bf16) | 2.0 GB | v0.3 adapter, final — applies only to the v0.2 bucket base; no v0.4 adapter is released (see Lineage) |
Changelog
v0.4 (2026-09-28, tag v0.4) — reinforcement-learning stage merged onto an unpublished
full-epoch merge of v0.3 (see Lineage), shipped with a new decoding recipe (sampled first pass at
temperature 0.4 / presence 0.3, one un-enforced retry at the same settings, then fallback). Paired
OHR-Bench retrieval vs the v0.3 weights on the same server: better overall (p = 2×10⁻¹⁹), driven
entirely by reading order on dense Chinese newspapers (.476 → .754), with no evidence type worse;
failed pages 1.00 % → 0.32 %. Weights, GPTQ-8bit and GGUF pushed in lockstep; no adapter (see
Lineage). Requires shrew-server ≥ 0.4.0 (ships the recipe and reports which
decoding rung produced each page). Buckets / pinpoints and the 36-value section_type contract
unchanged from v0.3.
v0.3 (2026-09-14, tag v0.3) — generation E6. New LoRA (same r=256 recipe, 1 epoch /
1,343 steps, eval_loss 0.1408 vs 0.1421) trained on the up-weighted dense/broadsheet slices with
the open section_type taxonomy; merged bf16, GPTQ-8bit, GGUF and adapter pushed in lockstep.
Promoted on six pre-registered criteria (retrieval paired PASS, product gates PASS, regions PASS,
parity TIE, contract screens PASS, image surface flat). Loops on OmniDocBench newspapers
66 % → 46 % in isolation; figure far-false-positives 156 → 51. The INT8 variant was gated by a
215-page product-path parity card against bf16 (all criteria met; text similarity median 1.000 on
pages both arms parsed). Requires shrew-server ≥ 0.3.14 for the 36-value section_type
contract. Buckets / pinpoints unchanged from v0.2. Retrieval tables are now the exact-kNN read.
v0.2 — glyph-routed tile buckets (E4); first GPTQ-8bit and GGUF releases.
This is a preview: weights update in place under these names as the model improves. Each weight
push's commit message records the training and calibration generation — pin a commit
(revision=) for reproducibility.
Base model: ibm-granite/granite-vision-4.1-4b (Apache 2.0). The released weights are a full-weight fork of the base (see Lineage): the language model and the upper vision-encoder attention layers and projectors are fine-tuned; the rest of the vision tower is unchanged.
- Downloads last month
- 40