Instructions to use sionic-ai/PepperOCR-VL with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use sionic-ai/PepperOCR-VL with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="sionic-ai/PepperOCR-VL") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("sionic-ai/PepperOCR-VL") model = AutoModelForMultimodalLM.from_pretrained("sionic-ai/PepperOCR-VL", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use sionic-ai/PepperOCR-VL with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "sionic-ai/PepperOCR-VL" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "sionic-ai/PepperOCR-VL", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/sionic-ai/PepperOCR-VL
- SGLang
How to use sionic-ai/PepperOCR-VL with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "sionic-ai/PepperOCR-VL" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "sionic-ai/PepperOCR-VL", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "sionic-ai/PepperOCR-VL" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "sionic-ai/PepperOCR-VL", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use sionic-ai/PepperOCR-VL with Docker Model Runner:
docker model run hf.co/sionic-ai/PepperOCR-VL
Very Promising!
Being an OCR model i threw it through a couple of images that i had issues with before with Gemma4 and other models.
Damn near perfect from what i can see. I mean it didn't carry the formatting; But that's small potatoes compared to it getting spelling wrong, or just deciding to write something unrelated, drop paragraphs, or decide to spellcheck and NOT follow instructions.
Mind you til this point I've relied on Tesseract OCR which does probably... 95% accurate (with a lot of common problems, like | or 1 instead of I). Gemma4 i got decent enough results but a few pages borked; But this gets much higher accuracy so far than even that.
Yeah I've only ran about 15 pages through it so far from 1bit text image to full-blown color image with slightly angled pages. But damn if this doesn't actually OCR correctly as far as i can tell. At least for english.
K poured another ~300 something pages in and may do another 1000; Seeing a few pages (20?) that borked, usually repeating a phrase or paragraph til the max tokens, or sometimes listing the filenames that were uploaded (so highest filesize or smallest). mostly easy to identify.
Were the image captures screencaps or flat it would likely get a 99%-100% accuracy. But with camera captures via phone, it's a bit lower.
But for the most part, decent results.
Thank you so much for the detailed feedback! To help improve the issues you mentioned, we recommend trying the best-performing configuration we used in our own evaluations.
Thank you so much for the detailed feedback! To help improve the issues you mentioned, we recommend trying the best-performing configuration we used in our own evaluations.
hmm, i was using a cli script with instructions intended for general LLM's and a fixed seed. Still, getting good results on the first few pages is very promising when in say Tesseract i'd get a lot of iffy characters between words at times.
I'll look over the configuration and see if i can incorporate it.
Limitations:
One page per request; no multi-page or PDF input.
Hmm a lot of the output i was getting was from double page images... And was working fine. I suppose splitting to individual pages on bad outputs may result in a better output...
Same testing page i was using which got with good results...
