Very Promising!

#1
by yano2mch - opened

Being an OCR model i threw it through a couple of images that i had issues with before with Gemma4 and other models.

Damn near perfect from what i can see. I mean it didn't carry the formatting; But that's small potatoes compared to it getting spelling wrong, or just deciding to write something unrelated, drop paragraphs, or decide to spellcheck and NOT follow instructions.

Mind you til this point I've relied on Tesseract OCR which does probably... 95% accurate (with a lot of common problems, like | or 1 instead of I). Gemma4 i got decent enough results but a few pages borked; But this gets much higher accuracy so far than even that.

Yeah I've only ran about 15 pages through it so far from 1bit text image to full-blown color image with slightly angled pages. But damn if this doesn't actually OCR correctly as far as i can tell. At least for english.

K poured another ~300 something pages in and may do another 1000; Seeing a few pages (20?) that borked, usually repeating a phrase or paragraph til the max tokens, or sometimes listing the filenames that were uploaded (so highest filesize or smallest). mostly easy to identify.

Were the image captures screencaps or flat it would likely get a 99%-100% accuracy. But with camera captures via phone, it's a bit lower.

But for the most part, decent results.

sionic-ai org

Thank you so much for the detailed feedback! To help improve the issues you mentioned, we recommend trying the best-performing configuration we used in our own evaluations.

Thank you so much for the detailed feedback! To help improve the issues you mentioned, we recommend trying the best-performing configuration we used in our own evaluations.

hmm, i was using a cli script with instructions intended for general LLM's and a fixed seed. Still, getting good results on the first few pages is very promising when in say Tesseract i'd get a lot of iffy characters between words at times.

I'll look over the configuration and see if i can incorporate it.

Limitations:
One page per request; no multi-page or PDF input.

Hmm a lot of the output i was getting was from double page images... And was working fine. I suppose splitting to individual pages on bad outputs may result in a better output...

Same testing page i was using which got with good results...

reckless-pg9

Sign up or log in to comment