Instructions to use vvs184/deepseek-ocr-page-classifier-b04 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use vvs184/deepseek-ocr-page-classifier-b04 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-classification", model="vvs184/deepseek-ocr-page-classifier-b04", trust_remote_code=True) pipe("https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/hub/parrots.png")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("vvs184/deepseek-ocr-page-classifier-b04", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
DeepSeek-OCR page classifier (architecture B)
Variant: unmerged. Base: deepseek-ai/DeepSeek-OCR at revision 9f30c71f441d010e5429c532364a86705536c53a.
Intended use: classify single rendered document pages into 7 operational page types, routing low-confidence pages to UNKNOWN for human review. Not intended for OCR, text generation or any decision without human oversight.
Architecture
page image -> 640x640 mean-grey pad -> Normalize(0.5, 0.5, 0.5)
-> SAM ViT-B (+ neck) ---------------------------+
-> CLIP-L (SAM features as patch embeddings) ----+-> concat [2048]
-> trained linear projector [2048 -> 1280] -> 10x10 grid
-> one image_newline per row + one view_seperator = 111 page tokens
[BOS row][111 page tokens][learned CLASSIFY token]
-> 12-layer DeepSeek MoE decoder (Q/O LoRA r16 a32) -> final hidden state
-> LayerNorm -> Dropout -> Linear [1280 -> 7] -> logits / T -> softmax
-> confidence < threshold => UNKNOWN
Checkpoint stripping
| Upstream tensor | Disposition | Reason |
|---|---|---|
lm_head.weight |
dropped | vocabulary projection unused |
model.vision_model.embeddings.patch_embedding.weight |
dropped | CLIP patch conv bypassed (SAM features are the patch embeddings) |
model.projector.layers.{weight,bias} |
dropped | superseded by trained projector |
model.embed_tokens.weight |
sliced | only the BOS row 0 is used (vocab_size=1) |
model.layers.N.self_attn.{q,o}_proj.weight |
adapted | LoRA base or merged |
| every other tensor | kept | BF16 bytes unchanged |
Upstream tensors: 2,710; exported: 2,706; dropped: 4. Parameters: upstream 3,336,106,240, exported 3,001,925,888, removed 334,180,352. Every tensor is listed in stripped_inventory.json.
Trained parameters
trained.safetensors (F32) holds 55 tensors, 3,618,567 parameters.
| Group | Parameters |
|---|---|
| classification_token | 1,280 |
| head | 11,527 |
| lora | 983,040 |
| projector | 2,622,720 |
Inputs and decisions
Preprocessing: RGB, contain-resize to 640 px, mean-grey (127) square pad, scale to [0, 1], Normalize(mean 0.5, std 0.5). PDF rasterization is not included.
Labels: Inspection_Report, Job_Paperwork, Other, Pictures, Service_Order, Vendor_Invoice, Vendor_Proposal. Calibration: softmax(logits / T) with T = 1.5147499307352106. Pages whose top probability is below 0.8 are labeled UNKNOWN; the best known label is kept as the runner-up.
Loading
Tested with transformers==4.46.3 (hard requirement) and torch==2.14.0, BF16 on CUDA, eager attention, native ATen dispatch, one physical page per forward pass.
from huggingface_hub import snapshot_download
from transformers.dynamic_module_utils import get_class_from_dynamic_module
folder = snapshot_download("vvs184/deepseek-ocr-page-classifier-b04", revision="main")
model_class = get_class_from_dynamic_module(
"modeling_page_classifier.PageClassifierModel", folder
)
model = model_class.load_for_inference(folder, device="cuda")
predictions = model.classify([page_image]) # PIL RGB images
AutoModel.from_pretrained(folder, trust_remote_code=True) restores the same weights through the same strict loader (missing or unexpected tensors, other loading options and configuration overrides are refused); call model.prepare_for_inference("cuda") before use.
NOT PRODUCTION-QUALIFIED
This model has not passed its production quality gate. The blinded human gold set is unfilled, so the required floors (accuracy 0.99, macro-F1 0.985, supported critical recall 0.995) are unmet and unproven.
Pseudo-label measurements only (not human truth): accuracy 0.9663, macro-F1 0.8091, minimum critical recall 0.9706.
Limitations
- Near-threshold sensitivity: pages close to the rejection threshold can flip between a label and UNKNOWN under different batching, padding or kernels.
- Singleton requirement: serve one physical page per forward pass; grouped batches change MoE numerics and decisions.
- PDF rasterization (PyMuPDF), upload limits and serving controls are not included; inputs must be rendered page images.
- Pseudo-label metrics are engineering checks, not human quality evidence.
Parity
Repository serving path versus this export (content_sha256 in export_manifest.json), torch.bfloat16 on cuda. Unmerged accepted: True.
| Variant | Cohort | Pages | Decision changes | Max abs logit diff | Rejection changes | Verdict |
|---|---|---|---|---|---|---|
| merged | sample_first | 512 | 0 | 0.3125 | 0 | reported_only |
| merged | sample_second | 512 | 0 | 0.3125 | 0 | reported_only |
| merged | test | 4689 | 5 | 0.875 | 5 | reported_only |
| merged | val | 4848 | 9 | 0.75 | 9 | reported_only |
| unmerged | sample_first | 512 | 0 | 0 | 0 | True |
| unmerged | sample_second | 512 | 0 | 0 | 0 | True |
| unmerged | test | 4689 | 0 | 0 | 0 | True |
| unmerged | val | 4848 | 0 | 0 | 0 | True |
License
Upstream DeepSeek-OCR code and weights: MIT License, Copyright (c) 2023 DeepSeek (see LICENSE). Trained classifier weights and page-classifier code: license other, private, see NOTICE.
- Downloads last month
- 13
Model tree for vvs184/deepseek-ocr-page-classifier-b04
Base model
deepseek-ai/DeepSeek-OCR