Instructions to use ANZ-Innovation/VietSurya-v1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use ANZ-Innovation/VietSurya-v1 with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("datalab-to/surya-ocr-2") model = PeftModel.from_pretrained(base_model, "ANZ-Innovation/VietSurya-v1") - Notebooks
- Google Colab
- Kaggle
VietSurya-v1
A Vietnamese document OCR LoRA adapter for Datalab Surya OCR 2. This repository contains adapter weights and compatible tokenizer/processor assets; it does not contain the base model weights. Load it together with the base model.
Important evaluation note
This adapter is the best saved candidate by validation loss in its training run (step 22,925; loss 0.04034 on 512 stratified validation examples). Validation loss is not a direct OCR quality metric. In a prior small Vietnamese MDPBench comparison (18 images), the base model scored 72.1 and this fine-tune scored 48.1 (higher is better). Treat this release as experimental; that small comparison does not establish an improvement over the base model. Evaluate on your own representative documents before relying on it.
Load the adapter
import torch
from peft import PeftModel
from transformers import AutoModelForImageTextToText, AutoProcessor
base_id = "datalab-to/surya-ocr-2"
adapter_id = "ANZ-Innovation/VietSurya-v1"
processor = AutoProcessor.from_pretrained(adapter_id)
base = AutoModelForImageTextToText.from_pretrained(
base_id, torch_dtype="auto", device_map="auto"
)
model = PeftModel.from_pretrained(base, adapter_id)
Use the base model's documented Surya OCR 2 prompts and processor conventions for inference.
Training summary
- Method: LoRA fine-tuning; rank 32, alpha 64, dropout 0.05.
- Training completed at step 22,925, one epoch, BF16.
- Maximum training sequence length: 8,192 tokens.
- Spatial modules were frozen; this adapter targets selected language-model linear-attention and MLP modules.
- The training corpus combined Vietnamese document sources. The data itself is not included in this repository.
License and attribution
This is a derivative adapter for datalab-to/surya-ocr-2. It is distributed under the base model's modified AI Pubs Open Rail-M license included as LICENSE. That license carries use-based restrictions, attribution, and share-alike terms; read it before using or redistributing this adapter. No endorsement by Datalab is implied.
- Downloads last month
- 11
Model tree for ANZ-Innovation/VietSurya-v1
Base model
datalab-to/surya-ocr-2