masked-mha-runtime / VALIDATION.md
liangsu9988's picture
Promote latest kernel artifacts to main
5974ffe verified
|
Raw
History Blame Contribute Delete
1.44 kB

Validation

Release requires FP16 and BF16 parity against PyTorch SDPA at key lengths 41, 277, 1024, 1025, and 2048; poisoned padded scratch; CUDA Graph replay; source and installed-artifact runs on SM110; and native-vs-package timing.

The additive attention_mha_fp16_masked and attention_mha_bf16_masked entries must remain bitwise-equal to forward_static. The BF16 entry additionally checks qkv_token_stride against the actual Q tensor stride so fused-QKV views fail early instead of reading the wrong token row. Release gating includes poisoned padded columns, bitwise graph replay, and an explicit/static latency ratio no greater than 1.05x.

On 2026-08-06, source full passed 10/10 on NVIDIA Thor (SM110), PyTorch 2.13.0+cu130, and CUDA 13.0. The additive forward_seqused_static rows cover PI0.5 (Sq,H,D)=(10,8,256), Sk_max=456/968, and device-resident valid_k=456/712/968. Worst p99 absolute error was 0.000122, worst cosine was 0.99999982, and graph replay was bitwise deterministic.

On 2026-08-08, the source and clean installed-artifact gates both passed 10/10 on the same Thor class. The explicit FP16/BF16 entries measured 1.005x/1.000x of forward_static; BF16 fused-stride execution also passed torch.compile(fullgraph=True) exactly. For the GROOT row Sq=Sk=41,H=32,D=48, worst p99 absolute error was 0.003906 and cosine was 0.99999213; sequence lengths 1024, 1025, and 2048 also passed.