AX-Unlimited-OCR-3B-MoE-CUDA-AWQ-W4A16
CUDA PREVIEW — not a full multi-architecture GA release. NVIDIA Thor
aarch64passed formal native and immutable OCI acceptance.x86_64remains experimental until the exact Production A100 OCI gate passes.
This is one shared AWQ W4A16 checkpoint intended for CUDA/vLLM. The same model files may be downloaded on Thor and x86 Linux, but this preview makes a formal support claim only for NVIDIA Thor.
Architecture status
| Target | Status | Release claim |
|---|---|---|
NVIDIA Thor (aarch64) |
Formally validated | Supported by this preview |
A100-SXM4-80GB (x86_64, SM80) |
Experimental | Native evidence exists; formal Production OCI pending |
RTX A6000 (x86_64, SM86) |
Experimental | No formal compatibility or performance claim |
The available x86 development hosts currently report Thunder
mode=prototyping, Docker fastvfs, and /etc/thunder/libthunder.so
injection. Increasing their disks to 200 GB solved storage capacity but did not
turn them into standard Production NVIDIA runtimes. This is an evidence-host
limitation, not a claim that the checkpoint itself failed on x86.
native_x86_64_acceptance remains recorded as false; it was not removed or
changed to pass. The existing full CUDA publisher still requires both
architectures and all 31 GA gates. No immutable SemVer GA tag is created for
this preview.
Thor evidence
- Formal OCI child:
127.0.0.1:5000/axocr-cuda@sha256:73257ded4e5a0b763ce7c117dca846c6796b786b8933c236ab479e4f58845575 - Image config:
sha256:d6418adfcb11c371a13595d5df1af3d76f78542fc5f4050ac2534ae8285a0565 - OCI receipt:
7ef7bb054547ad3ec479ace48a2ec22d18ded21fd6738017fadee055d461f6d4 - Frozen runtime source:
9b71932f621ee2939b40507a8ab942559851e757c4e54fe59528b498072fce10 - Upstream source:
baidu/Unlimited-OCR@ee63731b6461c8afcdcc7b15352e7d2ffecc2ead - Evaluation set:
ax-ocr-accuracy-v3-20260723(208 samples)
| Thor metric | BF16 | AWQ |
|---|---|---|
| Aggregate throughput (completion tok/s) | 30.9081 | 31.0139 |
| p95 latency (s) | 23.2570 | 23.0460 |
| Mean CER | 0.089845 | 0.088245 |
| Digit CER | 0.076077 | 0.076705 |
The exact public-safe evidence summary is in preview_manifest.json.
Serving
Use AX-OCR 0.5.0 source matching the frozen runtime identity and public vLLM 0.25.1 with CUDA 13.0. On Thor the AX-OCR vLLM plugin selects the validated eager compatibility path.
ax-ocr vllm-serve \
--model-path AutomatosX/AX-Unlimited-OCR-3B-MoE-CUDA-AWQ-W4A16 \
--host 127.0.0.1 \
--port 8000
Requests must preserve the literal <image> prompt, set
skip_special_tokens=false, and use the Unlimited-OCR no-repeat/R-SWA request
settings recorded by AX-OCR.
Limitations
- This preview is not a Production support statement for x86 CUDA.
- A6000 and other NVIDIA GPUs may require different kernels or fallbacks.
- The model is not MLX or GGUF.
- OCR can hallucinate or omit content; verify high-stakes documents.
- Downloads last month
- 64
Model tree for AutomatosX/AX-Unlimited-OCR-3B-MoE-CUDA-AWQ-W4A16
Base model
baidu/Unlimited-OCR