Validation
Release requires FP16 and BF16 parity against PyTorch SDPA at key lengths 41, 277, 1024, 1025, and 2048; poisoned padded scratch; CUDA Graph replay; source and installed-artifact runs on SM110; and native-vs-package timing.
The additive attention_mha_fp16_masked and
attention_mha_bf16_masked entries must remain bitwise-equal to
forward_static. The BF16 entry additionally checks qkv_token_stride
against the actual Q tensor stride so fused-QKV views fail early instead of
reading the wrong token row. Release gating includes poisoned padded columns,
bitwise graph replay, and an explicit/static latency ratio no greater than
1.05x.
On 2026-08-06, source full passed 10/10 on NVIDIA Thor (SM110), PyTorch
2.13.0+cu130, and CUDA 13.0. The additive forward_seqused_static rows cover
PI0.5 (Sq,H,D)=(10,8,256), Sk_max=456/968, and device-resident
valid_k=456/712/968. Worst p99 absolute error was 0.000122, worst cosine
was 0.99999982, and graph replay was bitwise deterministic.
On 2026-08-08, the source and clean installed-artifact gates both passed
10/10 on the same Thor class. The explicit FP16/BF16 entries measured
1.005x/1.000x of forward_static; BF16 fused-stride execution also passed
torch.compile(fullgraph=True) exactly. For the GROOT row
Sq=Sk=41,H=32,D=48, worst p99 absolute error was 0.003906 and cosine was
0.99999213; sequence lengths 1024, 1025, and 2048 also passed.