ITA-OCR — GGUF weights

A fine-tune of GLM-OCR for Italian handwriting, packaged as GGUF for llama.cpp. These are the weights used by the ITA-OCR desktop application, which runs recognition entirely on the user's machine — no cloud OCR service, no API key, no telemetry.

Files

File Role Size
glm-ocr-ocr-ita-v8-native-3ep-clean-q8_0.gguf fine-tuned language model 686 MiB
glm-ocr-base-mmproj-q8_0.gguf vision projector (multimodal) 462 MiB

Both are required: the projector alone cannot transcribe, the model alone cannot see the page. Quantised to Q8_0.

95c6a7b7f318293c4a5276a5dab773d0ef55d49ad357e24f89f9d4a701a5be8a  glm-ocr-ocr-ita-v8-native-3ep-clean-q8_0.gguf
2f83f69e7e5268474c8e257606ace1c1c28ff3546e1308a63fb15d3dc2ebf459  glm-ocr-base-mmproj-q8_0.gguf

Usage

With llama.cpp. The prompt is GLM-OCR's own, Text Recognition:, with no system prompt and no extra instruction — the fine-tune stays directly comparable with the base model.

llama-mtmd-cli -m glm-ocr-ocr-ita-v8-native-3ep-clean-q8_0.gguf \
  --mmproj glm-ocr-base-mmproj-q8_0.gguf \
  --image page.png -p "Text Recognition:"

Pages are rendered onto a fixed 960×1248 canvas at 150 DPI, as in training.

With the application: put both files in the models folder next to the executable, or point OCR_ITA_MODELS at the folder holding them. Instructions in docs/en/MODEL.md. The app starts llama-server on loopback and applies a cascade of retries when a decode ends up incomplete, so command-line output can differ from the app's on the same page.

Training

LoRA rank 8 on the base model, 3 epochs, on Italian handwritten pages with human reference transcriptions. The split is writer-disjoint: no writer present in training appears in evaluation.

The work is documented in the technical report Teaching a Vision Model When to Stop, which describes the termination collapse observed after fine-tuning, the two stop tokens that caused it and the inference cascade that compensates for it.

The dataset remains private and is not published in any form.

Results

Base model against the shipped system. Every row states the set it was measured on: a percentage without its set is meaningless. The measurement noise floor is 0.3 points, so the margins below are well clear of it.

Set Pages CER WER
Development holdout, writer-disjoint 164 30.88 → 25.71 (−16.7%) 54.68 → 45.54 (−16.7%)
Sealed benchmark, readable part 67 29.28 → 18.78 (−35.9%) 56.74 → 38.99 (−31.3%)
Cohort 136 −8.8% −20.4%
Subset labelled easy a priori 61 12.06 → 11.28 (−6.5%) 36.97 → 30.67 (−17.0%)

The holdout uses capped CER, the only fair statistic where pages are lost; on the other sets neither model loses a page, so raw and capped coincide. The benchmark row excludes one writer at the edge of legibility whom the base model already reads at 52% capped CER — the aggregate over all 103 pages says more about that hand than about either model. On the easy subset the shipped system wins on 43 pages, loses on 17 and ties on 1.

Termination and the cascade

Fine-tuning induced a termination collapse: the model stops emitting the stop token reliably and runs to the context ceiling. Inference-time mitigations, measured on a 48-page panel of affected pages:

Configuration Pages lost / 48 Recovers Loops
Greedy, no mitigation 24 n/a 0
Frequency 0.35 + presence 0.20 5 20 / 24 1
Frequency 0.35 5 21 / 24 5
Presence 0.20 alone 22 4 / 24 0
DRY sampling 22 2 / 24 0

The shipped cascade uses the first configuration as its retry stage. It is cheaper than not having it: on the holdout, a full cascaded run takes less wall clock than a single plain greedy pass, because runaway pages are cut at around 550 tokens instead of reaching the 4,096 ceiling. Cascade behaviour: 140 pages resolved greedily, 19 on retry 1, 3 on retry 2, 2 left incomplete on the holdout; 71 / 21 / 3 / 8 on the benchmark.

Both official stop tokens must be honoured — this is what the --override-kv argument in the application sets up.

Quantisation

Paired per-page difference in character error against the bf16 reference, 126-page cohort. Positive is worse. Each configuration is the language tower plus the vision projector.

Configuration Mean diff. vs bf16 Pages identical to bf16
Q4_K_M + bf16 −0.02 2 / 126
Q8_0 + bf16 +0.08 49 / 126
Q8_0 + Q8_0 (shipped) +0.23 22 / 126
bf16 + Q8_0 +0.55 24 / 126
Q5_K_M + bf16 +0.65 10 / 126

The mean alone is misleading — the count of identical pages is the discriminator. Q8_0 is the only quantisation genuinely indistinguishable from bf16: median difference of zero, better and worse pages close to a coin toss, and 49 pages out of 126 character-identical. Q4_K_M looks fine on the mean yet leaves almost no page untouched (2 / 126) with roughly 60% of pages worse: a small but directional degradation. Against bf16, Q8_0 uses about 30% less video memory and decodes about 35% faster, which is why it ships for both towers.

CPU

Quality is unchanged on the CPU; latency is not. On the test machine a page goes from roughly 1.6 s on the GPU to roughly 14 s on the CPU, and 11–12 s of that is image encoding rather than generation: decoding nearly doubles from bf16 to Q4_K_M (63.8 → 124.2 tok/s) while the page still costs about fourteen seconds. Absolute timings depend on the machine — processor, GPU, drivers, page complexity — and should be read as a ratio. On the CPU, quantisation buys memory rather than latency: 3,410 MB resident at bf16 against 2,402 MB at Q8_0. The lever for latency would be the canvas resolution, not weight precision.

Inference cost is unchanged by fine-tuning itself: about 1.61 s per page for both base and fine-tuned model on the easy subset, ~297 generated tokens per page against a 1,542-token prompt that is almost entirely the image.

Limits

  • Transcription can omit, repeat or invent text, especially with difficult handwriting, irregular layout, formulas and tables. Always compare it with the original.
  • The model is adapted to Italian handwriting: outside that domain there is no claimed advantage over the base model.
  • Internal measurements are not an independent benchmark. Method and limits in BENCHMARKS.en.md.
  • Small variations in GPU decoding can produce different results.

Licence

MIT, like the GLM-OCR base weights. The application code is MIT; third-party components keep their own licences, listed in THIRD-PARTY.md.

The base model is GLM-OCR by Z.ai. "ITA-OCR" is an independent name and implies no affiliation with Z.ai or the GLM project.

Note on the use of AI

The application code, part of the documentation and the graphic assets were produced with the help of artificial-intelligence tools. That is not a guarantee of correctness or accuracy: transcriptions must be verified.


ITA-OCR — pesi GGUF (italiano)

Fine-tuning di GLM-OCR per la scrittura a mano italiana, in formato GGUF per llama.cpp. Sono i pesi usati dall'applicazione desktop ITA-OCR, che esegue il riconoscimento interamente sul computer dell'utente: nessun servizio OCR cloud, nessuna chiave API, nessuna telemetria.

File

File Ruolo Dimensione
glm-ocr-ocr-ita-v8-native-3ep-clean-q8_0.gguf modello linguistico adattato 686 MiB
glm-ocr-base-mmproj-q8_0.gguf proiettore visivo (multimodale) 462 MiB

Servono entrambi: il proiettore da solo non trascrive, il modello da solo non vede la pagina. Quantizzazione Q8_0. I checksum SHA-256 sono quelli riportati sopra.

Uso

Il prompt è quello di GLM-OCR, Text Recognition:, senza prompt di sistema e senza istruzioni aggiuntive. Le pagine vanno rese su tela fissa 960×1248 a 150 DPI, come in addestramento.

llama-mtmd-cli -m glm-ocr-ocr-ita-v8-native-3ep-clean-q8_0.gguf \
  --mmproj glm-ocr-base-mmproj-q8_0.gguf \
  --image pagina.png -p "Text Recognition:"

Con l'applicazione: metti i due file nella cartella models accanto all'eseguibile, oppure indica la cartella con OCR_ITA_MODELS. Istruzioni in docs/MODELLO.md.

Addestramento

LoRA rank 8 sul modello base, 3 epoche, su pagine manoscritte italiane con trascrizione umana di riferimento. Lo split è per scrivente: le persone presenti nell'addestramento non compaiono nella valutazione. Il lavoro è documentato nel report tecnico. Il dataset resta privato e non è pubblicato in nessuna forma.

Limiti

La trascrizione può omettere, ripetere o inventare testo, soprattutto con scrittura difficile, impaginazione irregolare, formule e tabelle: va sempre confrontata con l'originale. Le misure interne non sono un benchmark indipendente — metodo e limiti in BENCHMARKS.md.

Licenza

MIT, come i pesi base di GLM-OCR. I componenti di terze parti mantengono le proprie licenze, elencate in TERZE-PARTI.md. «ITA-OCR» è un nome indipendente e non indica affiliazione con Z.ai o con il progetto GLM.

Nota sull'uso dell'AI

Il codice dell'applicazione, parte della documentazione e gli asset grafici sono stati realizzati con il supporto di strumenti di intelligenza artificiale. Non costituisce garanzia di correttezza o accuratezza: le trascrizioni vanno verificate.

Downloads last month
44
GGUF
Model size
0.7B params
Architecture
glm4
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ueuegio/ITA-OCR

Base model

zai-org/GLM-OCR
Finetuned
(28)
this model