File size: 11,545 Bytes
7e3c772 584ee03 7e3c772 584ee03 7e3c772 584ee03 7e3c772 558bd18 7e3c772 558bd18 7e3c772 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 | ---
language:
- en
license: apache-2.0
tags:
- ocr
- captcha
- captcha-recognition
- image-to-text
- onnx
- pytorch
- safetensors
- crnn
- ctc
- mobilenet
- edge-ai
- computer-vision
- fast-inference
- lightweight
- quantization
- fp32
- fp16
- int8
- fp8
- int4
pipeline_tag: image-to-text
inference: false
model-index:
- name: captCHAD
results:
- task:
type: image-to-text
name: Optical Character Recognition
dataset:
name: Multi-Archetype CAPTCHA Robustness Benchmark
type: captcha-benchmark
metrics:
- type: accuracy
value: 80.0
name: Exact Match (Unseen Multi-Style)
verified: false
- type: accuracy
value: 95.0
name: Character Accuracy (Unseen Multi-Style)
verified: false
- type: latency
value: 0.56ms
name: CPU Inference Latency (ONNX)
verified: false
---
# captCHAD: Sub-100k Parameter Neural CAPTCHA OCR
<p align="center">
<img src="sample.png" alt="Sample CAPTCHA" width="280"/>
</p>
<p align="center">
<a href="https://huggingface.co/spaces/AndresDev/captCHAD-demo"><img src="https://img.shields.io/badge/π€%20Interactive%20Demo-WebAssembly-yellow" alt="Interactive Demo"></a>
<a href="https://huggingface.co/AndresDev/captCHAD"><img src="https://img.shields.io/badge/Release-v1.5-blueviolet" alt="Version"></a>
<a href="https://huggingface.co/AndresDev/captCHAD"><img src="https://img.shields.io/badge/Parameters-97k-brightgreen" alt="Parameters"></a>
<a href="https://huggingface.co/AndresDev/captCHAD"><img src="https://img.shields.io/badge/CPU_Latency-0.56ms-blue" alt="Latency"></a>
<a href="https://huggingface.co/AndresDev/captCHAD"><img src="https://img.shields.io/badge/ONNX_Size-400KB-orange" alt="Size"></a>
<a href="https://huggingface.co/AndresDev/captCHAD"><img src="https://img.shields.io/badge/Formats-PyTorch_%7C_ONNX_%7C_Safetensors-red" alt="Formats"></a>
<a href="https://huggingface.co/AndresDev/captCHAD"><img src="https://img.shields.io/badge/Quantizations-FP32_%7C_FP16_%7C_INT8_%7C_FP8_%7C_INT4-purple" alt="Quantizations"></a>
<a href="https://opensource.org/licenses/Apache-2.0"><img src="https://img.shields.io/badge/License-Apache_2.0-green.svg" alt="License"></a>
</p>
> π **v1.5 Update:** Weights upgraded to epoch 83 convergence. Real-world challenge benchmark increased to **4/5 (80%)** in the gallery and **5/6 (83.3%)** overall, now solving `PXAZGC` and `JQA3Lg` with 100% exact match. ONNX (FP32/INT8) and Safetensors (FP32/FP16) packages updated.
**captCHAD** is a compact, 97,057-parameter optical character recognition (OCR) model engineered for low-power CPU, edge, and browser environments.
> β‘ **Try the In-Browser Demo:** You can test captCHAD directly in your browser with zero server latency using WebAssembly at [AndresDev/captCHAD-demo](https://huggingface.co/spaces/AndresDev/captCHAD-demo). Pick from preset challenge datasets or drop/paste your own CAPTCHAs.
Rather than relying on heavy Vision Transformers (TrOCR, 61Mβ330M parameters) or standard CRNN pipelines that degrade when encountering crossing strike-bars, distorted grid patterns, and high-contrast color shifts, captCHAD employs a lightweight sequence architecture combining **Inverted Residual MobileNet blocks**, a **Dual Contrast Stem**, **Bidirectional GRUs**, and **Connectionist Temporal Classification (CTC)** decoding.
---
## Direct Challenge Evaluation Gallery
Evaluation against open-source models on real-world production CAPTCHAs collected from active websites:
| Image Challenge | Ground Truth | captCHAD v1.5 (97k) | Graf-J CRNN (3.6M) | Conv-Trans (12.3M) | TrOCR (61.6M) |
| :---: | :---: | :---: | :---: | :---: | :---: |
| <img src="assets/K6DXAY.png" width="160"/> | `K6DXAY` | **`K6DXAY` (100%)** | `K6DXAY` (100%) | `K6DxAY` | `x674w7` (0%) |
| <img src="assets/DmS3X.png" width="160"/> | `DmS3X` | **`DmS3X` (100%)** | `PxcC` (0%) | `DxP9` (0%) | `p4c4m` (0%) |
| <img src="assets/PXAZGC.png" width="160"/> | `PXAZGC` | **`PXAZGC` (100%)** | `PxAzGC` (100% case) | `PXAZCc` (83.3%) | `xw6c` (0%) |
| <img src="assets/JQA3Lg.png" width="160"/> | `JQA3Lg` | **`JQA3Lg` (100%)** | `lIJ` (0%) | `J8g` (0%) | `dn86g` (0%) |
| <img src="assets/7yCntT.png" width="160"/> | `7yCntT` | `2rGm5s` | `ajyGcr` | `7GEr` | `cmcn` |
---
## Comprehensive Benchmark: captCHAD vs. Hugging Face Models
Benchmarked on identical hardware (CPU, 4 threads) across 50 unseen multi-archetype test captchas:
| Model | Architecture | # Parameters | Disk Footprint | Exact Match | Mean Char Acc | CPU Latency |
| :--- | :--- | :---: | :---: | :---: | :---: | :---: |
| **captCHAD (ONNX)** | **Inverted Residuals + BiGRU + CTC** | **97,057** | **582 KB** | **80.0%** (40/50) | **95.00%** | **0.56 ms** |
| **captCHAD (PyTorch)** | **Inverted Residuals + BiGRU + CTC** | **97,057** | **436 KB** | **80.0%** (40/50) | **95.00%** | **3.65 ms** |
| **[Graf-J/captcha-crnn-finetuned](https://huggingface.co/Graf-J/captcha-crnn-finetuned)** | CNN + Bi-LSTM (CRNN) | 3,570,943 (3.57M) | 14.3 MB | 12.0% (6/50) | 58.33% | 16.18 ms |
| **[Graf-J/captcha-conv-transformer](https://huggingface.co/Graf-J/captcha-conv-transformer-finetuned)** | CNN + Transformer Encoder | 12,279,551 (12.3M) | 51.7 MB | 10.0% (5/50) | 58.90% | 18.22 ms |
| **[tomofi/trocr-captcha](https://huggingface.co/tomofi/trocr-captcha)** | TrOCR (Vision Transformer) | 61,596,672 (61.6M) | 246.5 MB | 0.0% (0/50) | 15.67% | 564.51 ms |
### Architectural Observations
- **TrOCR (61.6M params):** Autoregressive Vision Transformers trained primarily on clean documents lack inductive edge bias. When facing crossing strike-bars, non-linear waves, or bulge grids, the attention mechanism loses positional alignment and hallucinates, resulting in 0% exact match and 564 ms latency.
- **CRNN and Conv-Transformer (3.6Mβ12.3M params):** Standard convolutional backbones without edge-separation stages fuse background grid lines and noise specks into character activations, resulting in 10%β12% exact match on multi-archetype noise.
- **captCHAD (97k params):** Uses a pre-convolutional Contrast Stem (normalized luminance and directional Sobel gradients) and Squeeze-and-Excitation channel gating to isolate glyph contours from background clutter, maintaining 80% exact match with significantly lower compute requirements.
---
## Quantization Formats & Benchmark
To support different deployment environments (WASM browsers, microcontrollers, embedded Linux, server inference), captCHAD is provided in multiple precision formats:
| Format / File | Precision | Disk Size | Exact Match | Mean Char Acc | Single-Sample Latency | Profile / Recommendation |
| :--- | :--- | :---: | :---: | :---: | :---: | :--- |
| **`captchad.onnx`** | FP32 | **582 KB** | **80.0%** | **94.00%** | **1.03 ms** (0.56 ms raw) | **Fastest CPU Execution** (Production standard) |
| **`model_fp16.safetensors`** | FP16 | **206 KB** | **80.0%** | **94.00%** | 5.61 ms | **Optimal Balance** (50% size reduction, 0% accuracy loss) |
| **`captchad_fp16.onnx`** | FP16 | **395 KB** | **80.0%** | **94.00%** | 5.75 ms | **Balanced ONNX** (Half-precision graph) |
| **`captchad_int8.pt`** | INT8 | **320 KB** | **80.0%** | **94.00%** | 6.13 ms | **Quantized PyTorch** (Integer dynamic weights) |
| **`captchad_int8.onnx`** | INT8 | **412 KB** | 72.0% | 90.67% | 5.26 ms | **Edge Hardware** (Integer quantized graph) |
| **`model_fp8.safetensors`** | FP8 (`e4m3fn`) | **110 KB** | 68.0% | 91.33% | 5.68 ms | **Ultra-Compact** (72% size reduction, solid accuracy) |
| **`model_int4.safetensors`** | INT4 Packed | **72 KB** | 4.0% | 41.30% | 5.63 ms | **Extreme Compression** (Experimental 4-bit packaging) |
---
## Architecture Details
```
Input Image (3 x 64 x 192)
β
βΌ
[Contrast Stem] ββ> Extracts Normalized Luminance + Sobel Spatial Gradients (Sobel-X & Sobel-Y)
β
βΌ
[MobileNet Inverted Residual Blocks] ββ> Depthwise-Separable Convolutions + Squeeze-and-Excitation
β
βΌ
[Height Pooling] ββ> Vertical feature compression: (B, 52, 48)
β
βΌ
[Bidirectional GRU] ββ> Horizontal temporal sequence modeling
β
βΌ
[CTC Sequence Decoder] ββ> Arbitrary-length prediction without segmentation
```
- **Input Dimensions:** `(B, 3, 64, 192)`
- **Output Dimensions:** `(48, B, 63)` (Logits for CTC Loss, sequence length $T=48$)
- **Output Character Length:** 1 to 8 characters per image (optimized for 4β7 character CAPTCHAs)
- **Character Vocabulary:** `0123456789abcdefghijklmnopqrstuvwxyzABCDEFGHIJKLMNOPQRSTUVWXYZ` (62 alphanumeric classes + 1 CTC blank token at index 0)
- **Trainable Parameters:** `97,057`
---
## Quick Start
### 1. ONNX Runtime Inference
Requires `onnxruntime`, `numpy`, and `pillow`:
```python
import numpy as np
from PIL import Image
import onnxruntime as ort
# Supports captchad.onnx (FP32), captchad_int8.onnx (INT8), captchad_fp16.onnx (FP16)
session = ort.InferenceSession("captchad.onnx")
input_name = session.get_inputs()[0].name
CHARSET = "0123456789abcdefghijklmnopqrstuvwxyzABCDEFGHIJKLMNOPQRSTUVWXYZ"
IDX2CHAR = {i + 1: ch for i, ch in enumerate(CHARSET)}
# Preprocess image to [1, 3, 64, 192] normalized in [-1.0, 1.0]
img = Image.open("sample.png").convert("RGB").resize((192, 64), Image.BILINEAR)
arr = (np.array(img, dtype=np.float32).transpose(2, 0, 1) - 127.5) / 127.5
inp = arr[np.newaxis, :, :, :]
# Run inference
logits = session.run(None, {input_name: inp})[0]
preds = np.argmax(logits[:, 0, :], axis=-1)
# CTC Collapse
decoded, prev = [], None
for t in preds:
if t != prev and t != 0:
decoded.append(IDX2CHAR[t])
prev = t
print("Prediction:", "".join(decoded))
```
### 2. Safetensors and PyTorch Inference
```python
import torch
from safetensors.torch import load_file
from model import captCHAD, decode_tokens
from PIL import Image
import numpy as np
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model = captCHAD()
# Load weights (supports model.safetensors, model_fp16.safetensors, model_fp8.safetensors)
state_dict = load_file("model_fp16.safetensors")
state_dict = {k: v.to(torch.float32) if v.is_floating_point() else v for k, v in state_dict.items()}
model.load_state_dict(state_dict)
model.to(device)
model.eval()
img = Image.open("sample.png").convert("RGB").resize((192, 64), Image.BILINEAR)
arr = (np.array(img, dtype=np.float32).transpose(2, 0, 1) - 127.5) / 127.5
tensor = torch.from_numpy(arr).unsqueeze(0).to(device)
with torch.no_grad():
logits = model(tensor)
preds = logits.argmax(dim=-1)[:, 0].tolist()
text = decode_tokens(preds)
print("Prediction:", text)
```
### 3. Command-Line Inference
Test different engines and quantization formats directly via `inference.py`:
```bash
# Fastest CPU production baseline (ONNX FP32)
python inference.py sample.png --engine onnx --quant fp32
# Balanced half-precision (ONNX FP16, 395 KB)
python inference.py sample.png --engine onnx --quant fp16
# Dynamic integer quantization (ONNX INT8, 412 KB)
python inference.py sample.png --engine onnx --quant int8
# Compact half-precision (Safetensors FP16, 206 KB)
python inference.py sample.png --engine safetensors --quant fp16
# Ultra-compact 8-bit float (Safetensors FP8, 110 KB)
python inference.py sample.png --engine safetensors --quant fp8
# Extreme 4-bit packed weights (Safetensors INT4, 72 KB)
python inference.py sample.png --engine safetensors --quant int4
```
---
## License
This project is licensed under the Apache 2.0 License.
|