File size: 11,545 Bytes
7e3c772
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
584ee03
7e3c772
 
584ee03
7e3c772
 
 
 
 
584ee03
 
7e3c772
 
 
 
 
 
 
 
 
 
 
 
558bd18
7e3c772
 
 
558bd18
 
 
7e3c772
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
---
language:
- en
license: apache-2.0
tags:
- ocr
- captcha
- captcha-recognition
- image-to-text
- onnx
- pytorch
- safetensors
- crnn
- ctc
- mobilenet
- edge-ai
- computer-vision
- fast-inference
- lightweight
- quantization
- fp32
- fp16
- int8
- fp8
- int4
pipeline_tag: image-to-text
inference: false
model-index:
- name: captCHAD
  results:
  - task:
      type: image-to-text
      name: Optical Character Recognition
    dataset:
      name: Multi-Archetype CAPTCHA Robustness Benchmark
      type: captcha-benchmark
    metrics:
    - type: accuracy
      value: 80.0
      name: Exact Match (Unseen Multi-Style)
      verified: false
    - type: accuracy
      value: 95.0
      name: Character Accuracy (Unseen Multi-Style)
      verified: false
    - type: latency
      value: 0.56ms
      name: CPU Inference Latency (ONNX)
      verified: false
---

# captCHAD: Sub-100k Parameter Neural CAPTCHA OCR

<p align="center">
  <img src="sample.png" alt="Sample CAPTCHA" width="280"/>
</p>

<p align="center">
  <a href="https://huggingface.co/spaces/AndresDev/captCHAD-demo"><img src="https://img.shields.io/badge/πŸ€—%20Interactive%20Demo-WebAssembly-yellow" alt="Interactive Demo"></a>
  <a href="https://huggingface.co/AndresDev/captCHAD"><img src="https://img.shields.io/badge/Release-v1.5-blueviolet" alt="Version"></a>
  <a href="https://huggingface.co/AndresDev/captCHAD"><img src="https://img.shields.io/badge/Parameters-97k-brightgreen" alt="Parameters"></a>
  <a href="https://huggingface.co/AndresDev/captCHAD"><img src="https://img.shields.io/badge/CPU_Latency-0.56ms-blue" alt="Latency"></a>
  <a href="https://huggingface.co/AndresDev/captCHAD"><img src="https://img.shields.io/badge/ONNX_Size-400KB-orange" alt="Size"></a>
  <a href="https://huggingface.co/AndresDev/captCHAD"><img src="https://img.shields.io/badge/Formats-PyTorch_%7C_ONNX_%7C_Safetensors-red" alt="Formats"></a>
  <a href="https://huggingface.co/AndresDev/captCHAD"><img src="https://img.shields.io/badge/Quantizations-FP32_%7C_FP16_%7C_INT8_%7C_FP8_%7C_INT4-purple" alt="Quantizations"></a>
  <a href="https://opensource.org/licenses/Apache-2.0"><img src="https://img.shields.io/badge/License-Apache_2.0-green.svg" alt="License"></a>
</p>

> πŸš€ **v1.5 Update:** Weights upgraded to epoch 83 convergence. Real-world challenge benchmark increased to **4/5 (80%)** in the gallery and **5/6 (83.3%)** overall, now solving `PXAZGC` and `JQA3Lg` with 100% exact match. ONNX (FP32/INT8) and Safetensors (FP32/FP16) packages updated.

**captCHAD** is a compact, 97,057-parameter optical character recognition (OCR) model engineered for low-power CPU, edge, and browser environments.

> ⚑ **Try the In-Browser Demo:** You can test captCHAD directly in your browser with zero server latency using WebAssembly at [AndresDev/captCHAD-demo](https://huggingface.co/spaces/AndresDev/captCHAD-demo). Pick from preset challenge datasets or drop/paste your own CAPTCHAs.

Rather than relying on heavy Vision Transformers (TrOCR, 61M–330M parameters) or standard CRNN pipelines that degrade when encountering crossing strike-bars, distorted grid patterns, and high-contrast color shifts, captCHAD employs a lightweight sequence architecture combining **Inverted Residual MobileNet blocks**, a **Dual Contrast Stem**, **Bidirectional GRUs**, and **Connectionist Temporal Classification (CTC)** decoding.

---

## Direct Challenge Evaluation Gallery

Evaluation against open-source models on real-world production CAPTCHAs collected from active websites:

| Image Challenge | Ground Truth | captCHAD v1.5 (97k) | Graf-J CRNN (3.6M) | Conv-Trans (12.3M) | TrOCR (61.6M) |
| :---: | :---: | :---: | :---: | :---: | :---: |
| <img src="assets/K6DXAY.png" width="160"/> | `K6DXAY` | **`K6DXAY` (100%)** | `K6DXAY` (100%) | `K6DxAY` | `x674w7` (0%) |
| <img src="assets/DmS3X.png" width="160"/> | `DmS3X` | **`DmS3X` (100%)** | `PxcC` (0%) | `DxP9` (0%) | `p4c4m` (0%) |
| <img src="assets/PXAZGC.png" width="160"/> | `PXAZGC` | **`PXAZGC` (100%)** | `PxAzGC` (100% case) | `PXAZCc` (83.3%) | `xw6c` (0%) |
| <img src="assets/JQA3Lg.png" width="160"/> | `JQA3Lg` | **`JQA3Lg` (100%)** | `lIJ` (0%) | `J8g` (0%) | `dn86g` (0%) |
| <img src="assets/7yCntT.png" width="160"/> | `7yCntT` | `2rGm5s` | `ajyGcr` | `7GEr` | `cmcn` |

---

## Comprehensive Benchmark: captCHAD vs. Hugging Face Models

Benchmarked on identical hardware (CPU, 4 threads) across 50 unseen multi-archetype test captchas:

| Model | Architecture | # Parameters | Disk Footprint | Exact Match | Mean Char Acc | CPU Latency |
| :--- | :--- | :---: | :---: | :---: | :---: | :---: |
| **captCHAD (ONNX)** | **Inverted Residuals + BiGRU + CTC** | **97,057** | **582 KB** | **80.0%** (40/50) | **95.00%** | **0.56 ms** |
| **captCHAD (PyTorch)** | **Inverted Residuals + BiGRU + CTC** | **97,057** | **436 KB** | **80.0%** (40/50) | **95.00%** | **3.65 ms** |
| **[Graf-J/captcha-crnn-finetuned](https://huggingface.co/Graf-J/captcha-crnn-finetuned)** | CNN + Bi-LSTM (CRNN) | 3,570,943 (3.57M) | 14.3 MB | 12.0% (6/50) | 58.33% | 16.18 ms |
| **[Graf-J/captcha-conv-transformer](https://huggingface.co/Graf-J/captcha-conv-transformer-finetuned)** | CNN + Transformer Encoder | 12,279,551 (12.3M) | 51.7 MB | 10.0% (5/50) | 58.90% | 18.22 ms |
| **[tomofi/trocr-captcha](https://huggingface.co/tomofi/trocr-captcha)** | TrOCR (Vision Transformer) | 61,596,672 (61.6M) | 246.5 MB | 0.0% (0/50) | 15.67% | 564.51 ms |

### Architectural Observations

- **TrOCR (61.6M params):** Autoregressive Vision Transformers trained primarily on clean documents lack inductive edge bias. When facing crossing strike-bars, non-linear waves, or bulge grids, the attention mechanism loses positional alignment and hallucinates, resulting in 0% exact match and 564 ms latency.
- **CRNN and Conv-Transformer (3.6M–12.3M params):** Standard convolutional backbones without edge-separation stages fuse background grid lines and noise specks into character activations, resulting in 10%–12% exact match on multi-archetype noise.
- **captCHAD (97k params):** Uses a pre-convolutional Contrast Stem (normalized luminance and directional Sobel gradients) and Squeeze-and-Excitation channel gating to isolate glyph contours from background clutter, maintaining 80% exact match with significantly lower compute requirements.

---

## Quantization Formats & Benchmark

To support different deployment environments (WASM browsers, microcontrollers, embedded Linux, server inference), captCHAD is provided in multiple precision formats:

| Format / File | Precision | Disk Size | Exact Match | Mean Char Acc | Single-Sample Latency | Profile / Recommendation |
| :--- | :--- | :---: | :---: | :---: | :---: | :--- |
| **`captchad.onnx`** | FP32 | **582 KB** | **80.0%** | **94.00%** | **1.03 ms** (0.56 ms raw) | **Fastest CPU Execution** (Production standard) |
| **`model_fp16.safetensors`** | FP16 | **206 KB** | **80.0%** | **94.00%** | 5.61 ms | **Optimal Balance** (50% size reduction, 0% accuracy loss) |
| **`captchad_fp16.onnx`** | FP16 | **395 KB** | **80.0%** | **94.00%** | 5.75 ms | **Balanced ONNX** (Half-precision graph) |
| **`captchad_int8.pt`** | INT8 | **320 KB** | **80.0%** | **94.00%** | 6.13 ms | **Quantized PyTorch** (Integer dynamic weights) |
| **`captchad_int8.onnx`** | INT8 | **412 KB** | 72.0% | 90.67% | 5.26 ms | **Edge Hardware** (Integer quantized graph) |
| **`model_fp8.safetensors`** | FP8 (`e4m3fn`) | **110 KB** | 68.0% | 91.33% | 5.68 ms | **Ultra-Compact** (72% size reduction, solid accuracy) |
| **`model_int4.safetensors`** | INT4 Packed | **72 KB** | 4.0% | 41.30% | 5.63 ms | **Extreme Compression** (Experimental 4-bit packaging) |

---

## Architecture Details

```
Input Image (3 x 64 x 192)
     β”‚
     β–Ό
[Contrast Stem] ──> Extracts Normalized Luminance + Sobel Spatial Gradients (Sobel-X & Sobel-Y)
     β”‚
     β–Ό
[MobileNet Inverted Residual Blocks] ──> Depthwise-Separable Convolutions + Squeeze-and-Excitation
     β”‚
     β–Ό
[Height Pooling] ──> Vertical feature compression: (B, 52, 48)
     β”‚
     β–Ό
[Bidirectional GRU] ──> Horizontal temporal sequence modeling
     β”‚
     β–Ό
[CTC Sequence Decoder] ──> Arbitrary-length prediction without segmentation
```

- **Input Dimensions:** `(B, 3, 64, 192)`
- **Output Dimensions:** `(48, B, 63)` (Logits for CTC Loss, sequence length $T=48$)
- **Output Character Length:** 1 to 8 characters per image (optimized for 4–7 character CAPTCHAs)
- **Character Vocabulary:** `0123456789abcdefghijklmnopqrstuvwxyzABCDEFGHIJKLMNOPQRSTUVWXYZ` (62 alphanumeric classes + 1 CTC blank token at index 0)
- **Trainable Parameters:** `97,057`

---

## Quick Start

### 1. ONNX Runtime Inference

Requires `onnxruntime`, `numpy`, and `pillow`:

```python
import numpy as np
from PIL import Image
import onnxruntime as ort

# Supports captchad.onnx (FP32), captchad_int8.onnx (INT8), captchad_fp16.onnx (FP16)
session = ort.InferenceSession("captchad.onnx")
input_name = session.get_inputs()[0].name

CHARSET = "0123456789abcdefghijklmnopqrstuvwxyzABCDEFGHIJKLMNOPQRSTUVWXYZ"
IDX2CHAR = {i + 1: ch for i, ch in enumerate(CHARSET)}

# Preprocess image to [1, 3, 64, 192] normalized in [-1.0, 1.0]
img = Image.open("sample.png").convert("RGB").resize((192, 64), Image.BILINEAR)
arr = (np.array(img, dtype=np.float32).transpose(2, 0, 1) - 127.5) / 127.5
inp = arr[np.newaxis, :, :, :]

# Run inference
logits = session.run(None, {input_name: inp})[0]
preds = np.argmax(logits[:, 0, :], axis=-1)

# CTC Collapse
decoded, prev = [], None
for t in preds:
    if t != prev and t != 0:
        decoded.append(IDX2CHAR[t])
    prev = t

print("Prediction:", "".join(decoded))
```

### 2. Safetensors and PyTorch Inference

```python
import torch
from safetensors.torch import load_file
from model import captCHAD, decode_tokens
from PIL import Image
import numpy as np

device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model = captCHAD()

# Load weights (supports model.safetensors, model_fp16.safetensors, model_fp8.safetensors)
state_dict = load_file("model_fp16.safetensors")
state_dict = {k: v.to(torch.float32) if v.is_floating_point() else v for k, v in state_dict.items()}
model.load_state_dict(state_dict)
model.to(device)
model.eval()

img = Image.open("sample.png").convert("RGB").resize((192, 64), Image.BILINEAR)
arr = (np.array(img, dtype=np.float32).transpose(2, 0, 1) - 127.5) / 127.5
tensor = torch.from_numpy(arr).unsqueeze(0).to(device)

with torch.no_grad():
    logits = model(tensor)
    preds = logits.argmax(dim=-1)[:, 0].tolist()
    text = decode_tokens(preds)

print("Prediction:", text)
```

### 3. Command-Line Inference

Test different engines and quantization formats directly via `inference.py`:

```bash
# Fastest CPU production baseline (ONNX FP32)
python inference.py sample.png --engine onnx --quant fp32

# Balanced half-precision (ONNX FP16, 395 KB)
python inference.py sample.png --engine onnx --quant fp16

# Dynamic integer quantization (ONNX INT8, 412 KB)
python inference.py sample.png --engine onnx --quant int8

# Compact half-precision (Safetensors FP16, 206 KB)
python inference.py sample.png --engine safetensors --quant fp16

# Ultra-compact 8-bit float (Safetensors FP8, 110 KB)
python inference.py sample.png --engine safetensors --quant fp8

# Extreme 4-bit packed weights (Safetensors INT4, 72 KB)
python inference.py sample.png --engine safetensors --quant int4
```

---

## License

This project is licensed under the Apache 2.0 License.