Spaces:
Running
Running
Download benchmarks.html from moebiusT7/book-ocr-studio: direct link, hf CLI and curl.
- Browser
- Download file 7.15 kB
-
https://huggingface.co/spaces/moebiusT7/book-ocr-studio/resolve/main/benchmarks.html
- Command line
-
hf download hf://spaces/moebiusT7/book-ocr-studio/benchmarks.html
-
curl -L -o benchmarks.html https://huggingface.co/spaces/moebiusT7/book-ocr-studio/resolve/main/benchmarks.html
7.15 kB
| <html lang="en"><head><meta charset="utf-8"><meta name="viewport" content="width=device-width,initial-scale=1"><title>12B / 26B C1-OCR comparison and limitations</title><style>body{margin:0;background:#f5f4ef;color:#18342f;font:17px/1.65 system-ui,sans-serif}main{max-width:900px;margin:auto;padding:35px 24px}a{color:#215947}h1{line-height:1.2}pre{background:#e8ebe3;padding:18px;overflow-x:auto;white-space:pre-wrap;overflow-wrap:anywhere}code{font-size:.88em}table{border-collapse:collapse;display:block;overflow:auto}td,th{padding:9px;border:1px solid #becbbf;text-align:left}nav{margin-bottom:32px}</style></head><body><main><nav><a href="index.html">Book OCR Studio</a> · <a href="install.html">Install</a> · <a href="benchmarks.html">Comparison</a> · <a href="release-notes.html">Release notes</a></nav><h1>12B / 26B C1-OCR comparison and limitations</h1> | |
| <h2>Latest test: current pipeline, 100 synthetic cases</h2> | |
| <p>On 2026-09-23, the shipped <code>c1_ocr_v3</code> review pipeline was tested on the same | |
| 100 short synthetic images (50 Japanese, 50 English) with three seeds per | |
| model, sequentially on one RTX 5070 Ti 16 GB. Both models used native Ollama, | |
| temperature 0 and an 8,192-token context. Each image had one corrupted OCR | |
| field and one initially correct field. Scores include image re-verification.</p> | |
| <table> | |
| <thead> | |
| <tr> | |
| <th>Metric</th> | |
| <th style="text-align:right;">12B + C1</th> | |
| <th style="text-align:right;">26B + C1</th> | |
| </tr> | |
| </thead> | |
| <tbody> | |
| <tr> | |
| <td>Corrected fields, exact match / 300</td> | |
| <td style="text-align:right;">165 (55.0%)</td> | |
| <td style="text-align:right;">74 (24.7%)</td> | |
| </tr> | |
| <tr> | |
| <td>Errors left unchanged / 300</td> | |
| <td style="text-align:right;">129</td> | |
| <td style="text-align:right;">226</td> | |
| </tr> | |
| <tr> | |
| <td>Changed but still not exact / 300</td> | |
| <td style="text-align:right;">6</td> | |
| <td style="text-align:right;">0</td> | |
| </tr> | |
| <tr> | |
| <td>Originally correct fields changed / 300</td> | |
| <td style="text-align:right;">3</td> | |
| <td style="text-align:right;">0</td> | |
| </tr> | |
| <tr> | |
| <td>Mean review time per case, excluding loading</td> | |
| <td style="text-align:right;">3.021 s</td> | |
| <td style="text-align:right;">2.519 s</td> | |
| </tr> | |
| </tbody> | |
| </table> | |
| <p>The 300 observations per model repeat 100 cases across three seeds; they are | |
| not 300 independent images. Exact matching includes whitespace and punctuation. | |
| The three changes to initially correct fields in 12B were whitespace removal. | |
| This larger test did <strong>not</strong> support a general quality advantage for 26B suggested | |
| by the older six-image results below. It supports retaining <strong>12B + C1 as the | |
| default</strong>, with 26B optional, rather than promising higher accuracy from 26B.</p> | |
| <p>A separate exploratory follow-up added the same explicit JSON-schema instruction | |
| to both models' prompts: exact corrections were 47/100 for 12B and 41/100 for 26B. | |
| This prompt change is <strong>not shipped</strong>; those results are not pooled with the table. | |
| The paired net-correction difference's 95% interval included zero in that follow-up.</p> | |
| <p>These are short horizontal synthetic cases with shared templates, not 100 books | |
| or a guarantee for real books, vertical text, ruby, complex layouts or languages | |
| beyond those tested. They measure the complete review stage, not capture, OCR or | |
| export time, and do not isolate raw model capability from review-gate behavior. | |
| The newer experiment's raw logs and fixtures are not bundled; this public summary | |
| is not an independently reproducible benchmark package.</p> | |
| <h2>Historical six-image test: earlier C1 prompt</h2> | |
| <p><strong>Exploratory measurements, not a claim about the current release's overall | |
| speed or accuracy.</strong> Measured on 2026-09-23 using an earlier OCR-specific C1 | |
| prompt. The current application uses <code>c1_ocr_v3</code>, with additional review | |
| steps. This is not a new benchmark of that pipeline.</p> | |
| <p>Six authored images each contained three erroneous and three correct OCR | |
| fields. Each arm used seeds 101, 202 and 303 at temperature 0: 18 requests, | |
| 54 repeated erroneous fields and 54 repeated correct fields per arm. | |
| These are repetitions, not 54 independent samples. The 26B+C1 arm was added | |
| later using the same images and scoring; the earlier arms were not rerun.</p> | |
| <table> | |
| <thead> | |
| <tr> | |
| <th>Metric</th> | |
| <th style="text-align:right;">12B + C1</th> | |
| <th style="text-align:right;">26B + C1</th> | |
| </tr> | |
| </thead> | |
| <tbody> | |
| <tr> | |
| <td>Sum of 18 review-request wall times</td> | |
| <td style="text-align:right;">52.291 s</td> | |
| <td style="text-align:right;">40.022 s</td> | |
| </tr> | |
| <tr> | |
| <td>Corrected fields, exact match / 54</td> | |
| <td style="text-align:right;">51</td> | |
| <td style="text-align:right;">45</td> | |
| </tr> | |
| <tr> | |
| <td>Errors left unchanged / 54</td> | |
| <td style="text-align:right;">3</td> | |
| <td style="text-align:right;">0</td> | |
| </tr> | |
| <tr> | |
| <td>Changed but still not exact / 54</td> | |
| <td style="text-align:right;">0</td> | |
| <td style="text-align:right;">9</td> | |
| </tr> | |
| <tr> | |
| <td>Originally correct fields changed / 54</td> | |
| <td style="text-align:right;">0</td> | |
| <td style="text-align:right;">0</td> | |
| </tr> | |
| <tr> | |
| <td>Corrected fields after colon/adjacent-space normalization / 54</td> | |
| <td style="text-align:right;">51</td> | |
| <td style="text-align:right;">54</td> | |
| </tr> | |
| </tbody> | |
| </table> | |
| <p>All nine exact-match failures in 26B+C1 were in three technical fields | |
| across three repeats: the erroneous letters/numbers were corrected, but | |
| full-width colons became half-width colons with different spacing. The last | |
| row is a supplementary, post-hoc sensitivity analysis, not the primary | |
| predefined score. Normalization does not establish perfect source fidelity. | |
| 12B left one authored name error in all three repeats.</p> | |
| <p>In this small test, 26B's summed request time was approximately <strong>23.5% lower</strong>. | |
| This does not include the complete capture/OCR/export workflow or establish | |
| a cold-start/loading comparison. No confidence interval, significance test, | |
| held-out generalization or training-contamination clearance is claimed. | |
| 26B may make a whole job slower if limited VRAM forces OCR and review to | |
| alternate. Both observed model artifacts were Q4_0; exact tags, digests, | |
| sizes and individual request measurements are in | |
| <a href="benchmarks/historical-c1.json">the measurement data</a>.</p> | |
| <p>The release therefore defaults to <strong>12B + C1</strong>. <strong>26B + C1 is optional</strong> for | |
| users able to provide more VRAM. These observations suggest a possible | |
| benefit on some errors, accompanied by a punctuation-fidelity tradeoff; | |
| they do not guarantee a speed or quality improvement on a user's books.</p> | |
| <p>The JSON contains only numeric measurements and synthetic case identifiers. | |
| Source-image fixtures and the old benchmark harness are not included, so | |
| this package supports checking the reported arithmetic, not rerunning the | |
| entire historical experiment. Source-file SHA-256 values identify the local | |
| records used; hashes alone are not independent verification of the experiment.</p> | |
| <hr><p><a href="BENCHMARKS.md" download>Download original text</a></p></main></body></html> | |