Kiel-2-OCR

Kiel-2-OCR is a fine-tuned Vision-Language Model (VLM) designed for high-precision document parsing, multi-lingual OCR, table extraction, and complex visual text recognition. It is built by fine-tuning the powerhouse vision architecture Qwen/Qwen2.5-VL-3B-Instruct using low-rank adapters (LoRA) via Hugging Face's TRL framework.


Model Details

  • Developed by: KielTech
  • Model Type: Vision-Language Model (OCR & Document AI)
  • Base Model: Qwen/Qwen2.5-VL-3B-Instruct
  • Language(s): Multi-lingual (English and supported Qwen languages)
  • License: Apache 2.0
  • Fine-tuning Method: Parameter-Efficient Fine-Tuning (PEFT) / LoRA
  • Hugging Face Hub: kiel2/Kiel-2-OCR

Intended Uses & Limitations

Intended Uses

  • Automated text extraction from structured and unstructured documents (PDF screenshots, receipts, invoices, forms).
  • Reading complex layout structures, handwritten notes, and dense text blocks.
  • Table understanding and key-value data extraction.

Limitations

  • The model inherits the native constraints of the Qwen2.5-VL architecture.
  • Performance on highly dense technical schematics or low-resolution text depends heavily on the input resolution configured during inference.

How to Get Started with the Model

You can load and run Kiel-2-OCR directly using Hugging Face transformers:

import torch
from transformers import Qwen2_5_VLForConditionalGeneration, AutoProcessor
from PIL import Image
import requests

MODEL_ID = "kiel2/Kiel-2-OCR"

print("Loading Kiel-2-OCR processor and model...")
processor = AutoProcessor.from_pretrained(MODEL_ID, trust_remote_code=True)
model = Qwen2_5_VLForConditionalGeneration.from_pretrained(
    MODEL_ID,
    torch_dtype=torch.float16,
    device_map="auto",
    trust_remote_code=True
)
model.eval()

# Prepare an image containing text
image_url = "[https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/transformers/tasks/car.jpg](https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/transformers/tasks/car.jpg)"
image = Image.open(requests.get(image_url, stream=True).raw)

conversation = [
    {
        "role": "user",
        "content": [
            {"type": "image", "image": image},
            {"type": "text", "text": "Extract all readable text from this image accurately."},
        ],
    }
]

# Apply chat template
text_prompt = processor.apply_chat_template(
    conversation, tokenize=False, add_generation_prompt=True
)

# Process inputs
inputs = processor(
    text=[text_prompt],
    images=[image],
    padding=True,
    return_tensors="pt"
).to(model.device)

# Generate response
with torch.no_grad():
    output_ids = model.generate(**inputs, max_new_tokens=512)

generated_ids = [
    output_ids[len(input_ids):] for input_ids, output_ids in zip(inputs.input_ids, output_ids)
]
response = processor.batch_decode(
    generated_ids, skip_special_tokens=True, clean_up_tokenization_spaces=True
)[0]

print("Extracted Text:\n", response)
Downloads last month
57
GGUF
Model size
3B params
Architecture
qwen2vl
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for kiel2/Kiel-2-OCR

Adapter
(299)
this model