AX-Unlimited-OCR-3B-MoE-CUDA-AWQ-W4A16

CUDA PREVIEW — not a full multi-architecture GA release. NVIDIA Thor aarch64 passed formal native and immutable OCI acceptance. x86_64 remains experimental until the exact Production A100 OCI gate passes.

This is one shared AWQ W4A16 checkpoint intended for CUDA/vLLM. The same model files may be downloaded on Thor and x86 Linux, but this preview makes a formal support claim only for NVIDIA Thor.

Architecture status

Target Status Release claim
NVIDIA Thor (aarch64) Formally validated Supported by this preview
A100-SXM4-80GB (x86_64, SM80) Experimental Native evidence exists; formal Production OCI pending
RTX A6000 (x86_64, SM86) Experimental No formal compatibility or performance claim

The available x86 development hosts currently report Thunder mode=prototyping, Docker fastvfs, and /etc/thunder/libthunder.so injection. Increasing their disks to 200 GB solved storage capacity but did not turn them into standard Production NVIDIA runtimes. This is an evidence-host limitation, not a claim that the checkpoint itself failed on x86.

native_x86_64_acceptance remains recorded as false; it was not removed or changed to pass. The existing full CUDA publisher still requires both architectures and all 31 GA gates. No immutable SemVer GA tag is created for this preview.

Thor evidence

  • Formal OCI child: 127.0.0.1:5000/axocr-cuda@sha256:73257ded4e5a0b763ce7c117dca846c6796b786b8933c236ab479e4f58845575
  • Image config: sha256:d6418adfcb11c371a13595d5df1af3d76f78542fc5f4050ac2534ae8285a0565
  • OCI receipt: 7ef7bb054547ad3ec479ace48a2ec22d18ded21fd6738017fadee055d461f6d4
  • Frozen runtime source: 9b71932f621ee2939b40507a8ab942559851e757c4e54fe59528b498072fce10
  • Upstream source: baidu/Unlimited-OCR@ee63731b6461c8afcdcc7b15352e7d2ffecc2ead
  • Evaluation set: ax-ocr-accuracy-v3-20260723 (208 samples)
Thor metric BF16 AWQ
Aggregate throughput (completion tok/s) 30.9081 31.0139
p95 latency (s) 23.2570 23.0460
Mean CER 0.089845 0.088245
Digit CER 0.076077 0.076705

The exact public-safe evidence summary is in preview_manifest.json.

Serving

Use AX-OCR 0.5.0 source matching the frozen runtime identity and public vLLM 0.25.1 with CUDA 13.0. On Thor the AX-OCR vLLM plugin selects the validated eager compatibility path.

ax-ocr vllm-serve \
  --model-path AutomatosX/AX-Unlimited-OCR-3B-MoE-CUDA-AWQ-W4A16 \
  --host 127.0.0.1 \
  --port 8000

Requests must preserve the literal <image> prompt, set skip_special_tokens=false, and use the Unlimited-OCR no-repeat/R-SWA request settings recorded by AX-OCR.

Limitations

  • This preview is not a Production support statement for x86 CUDA.
  • A6000 and other NVIDIA GPUs may require different kernels or fallbacks.
  • The model is not MLX or GGUF.
  • OCR can hallucinate or omit content; verify high-stakes documents.
Downloads last month
64
Safetensors
Model size
3B params
Tensor type
I64
·
I32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AutomatosX/AX-Unlimited-OCR-3B-MoE-CUDA-AWQ-W4A16

Quantized
(20)
this model