|
Download VALIDATION.md from flashrt/fp8-gemm: direct link, hf CLI and curl.
- Browser
- Download file 5.51 kB
-
https://huggingface.co/kernels/flashrt/fp8-gemm/resolve/v1/VALIDATION.md
- Command line
-
hf download hf://flashrt/fp8-gemm@v1/VALIDATION.md
-
curl -L -o VALIDATION.md https://huggingface.co/kernels/flashrt/fp8-gemm/resolve/v1/VALIDATION.md
5.51 kB
| # Validation | |
| Date: June 20, 2026 | |
| Local environment: | |
| - GPU: NVIDIA GeForce RTX 5090 | |
| - PyTorch: 2.9.1+cu128 | |
| - CUDA runtime reported by PyTorch: 12.8 | |
| - Source build target: `sm_120a` | |
| ## Source Correctness | |
| Command: | |
| ```bash | |
| python fp8-gemm/tests/test_fp8_gemm.py --backend source --mode full | |
| ``` | |
| Result: 14/14 checks passed, plus the blockwise custom op passed | |
| `torch.compile(fullgraph=True)` with bitwise-equal output to the eager wrapper. | |
| Covered public v1 rows: | |
| - M=1 decode GEMV: `K in {512,4096}`, `N in {512,2048,8192}` | |
| - small-M GEMM: `M in {8,16,32,64}` with representative | |
| transformer/diffuser-adjacent `K,N` rows | |
| - M=1 residual-add GEMV | |
| - block-128 scaled FP8 GEMM at: | |
| - `(M,K,N)=(1,1024,1024)` | |
| - `(51,1536,1536)` | |
| - `(277,2048,2048)` | |
| - `(1024,1152,1152)` | |
| - `(2520,3072,3072)` | |
| - `(128,4096,12288)` | |
| Metrics: | |
| - `max_abs` | |
| - `mean_abs` | |
| - `p99_abs` | |
| - cosine similarity | |
| - output dtype | |
| - tolerance | |
| The blockwise rows use the stricter gate: | |
| - `max_abs <= 0.0625` | |
| - `mean_abs <= 0.003` | |
| - `p99_abs <= 0.015625` | |
| - cosine similarity `>= 0.9999` | |
| The release benchmark also compares the Tensor wrapper against an independent | |
| binding of the original FlashRT pointer API. Matching source code alone is not | |
| treated as proof of zero wrapper overhead. | |
| ## Source Benchmark | |
| Command: | |
| ```bash | |
| python fp8-gemm/benchmarks/benchmark.py \ | |
| --backend source --mode headline --warmup 20 --iterations 100 --compile-ref | |
| ``` | |
| Result: all public rows passed. Headline rows are recorded in | |
| `benchmarks/RESULTS.md`. | |
| ## Architecture Scope Boundary | |
| On SM120, the public per-tensor path supports `M=1` and `2 <= M <= 64`. | |
| The blockwise path retains its independent unrestricted-M contract. | |
| On SM110, the public per-tensor path uses the production CUTLASS Sq/T1/Wide | |
| family and supports the validated model-shape matrix. The current full sweep | |
| covers the large-M boundary `65`, exact PI0.5 prefill rows `712/768/970`, and | |
| representative `K,N` rows from PI0.5, GROOT N1.6/N1.7, Cosmos Edge, and | |
| LingBot VLA. It also gates BF16 bias, in-place bias+residual, and tanh-GELU | |
| bias epilogues on SigLIP dimensions `1152/3456/4304`. SM110 blockwise scaling | |
| is not claimed. | |
| The August 8 Thor source gate passed `39/39` with zero failures. Plain large-M | |
| GEMM was bitwise equal to the reference, and every PI0.5 prefill auto tile was | |
| within 2% of the fastest validated package/native tile. | |
| ## Thor SM110 Increment | |
| Validated August 2, 2026: | |
| - GPU: NVIDIA Thor, compute capability 11.0; | |
| - PyTorch: 2.11.0+cu130; | |
| - CUDA: 13.0; | |
| - CUTLASS: 4.5.2, matching the current `kernel-builder` `cutlass_4_5` | |
| dependency; | |
| - pinned builder: `e9152aa24e0d99eca255ca9f1beb996de32f9ca4`; | |
| - source correctness: 23/23; | |
| - locally installed aarch64 artifact correctness: 23/23; | |
| - `torch.compile(fullgraph=True)`: exact output parity; | |
| - CUDA Graph capture/replay: exact output parity; | |
| - original SM120 source regression on RTX 5090: 14/14. | |
| The 23 Thor rows include 20 production auto-dispatch checks and three forced | |
| Sq/T1/Wide diagnostics. Ordinary GEMMs were bitwise equal to the FP32 | |
| accumulation reference after BF16 output conversion. The residual row passed | |
| with `max_abs=0.0625`, `p99_abs=0.0625`, and cosine `0.9999958` under the | |
| documented BF16 residual contract. | |
| Source-to-installed-artifact performance parity passed over 17 public | |
| auto-dispatch shapes: median artifact/source `0.9986`, p95 `1.0195`, and max | |
| `1.0244`. | |
| Comparisons against the original FlashRT pointer entry are reported separately | |
| in `benchmarks/RESULTS.md`. | |
| The final clean local artifact was built from | |
| `d31c69b1cb97ecd703aba01e29f423097f11c86a`. All 17 production rows passed the | |
| dispatcher gate; the worst auto/fastest-valid-tile paired ratio was `1.0028`. | |
| Sixteen rows matched the original CUTLASS 4.4.2 native entry within about 1.3% | |
| in the paired graph comparison. The PI0.5 gate/up row is a documented CUTLASS | |
| 4.5.2 version outlier at `1.128x`; it is not described as native-performance | |
| parity. | |
| Before the SM110 update was published, the existing Thor pipeline dependency | |
| set was cold-loaded from Hub using both `kernels==0.16.0` and | |
| `kernels==0.12.3`: 20/20 package imports passed for each client. The Thor host | |
| required `HF_ENDPOINT=https://hf-mirror.com`; direct access to | |
| `huggingface.co` timed out, so official-endpoint cold loading remains a | |
| post-publication check on a host with direct Hub access. | |
| ## HF Jobs Publish Status | |
| `flashrt/fp8-gemm` v1 was built and uploaded through the repository HF Jobs | |
| workflow. | |
| - Hub revision checked on June 20, 2026: `166f09be` | |
| - Uploaded variants: | |
| - `torch211-cxx11-cu128-x86_64-linux` | |
| - `torch211-cxx11-cu130-x86_64-linux` | |
| - `torch212-cxx11-cu130-x86_64-linux` | |
| - `torch212-cxx11-cu132-x86_64-linux` | |
| The existing SM120 Hub variants remain published. The new | |
| `torch211-cxx11-cu130-aarch64-linux` SM110 artifact is not included in the | |
| older Hub revision above; it must be published only after the clean-commit | |
| artifact rebuild and cold-cache checks. | |
| ## SM89 Increment | |
| The source now also exposes block-128 scaled FP8 GEMM/GEMV and | |
| `fp8_blockwise_swiglu_quantize_fp8` on SM89. The fused producer performs the | |
| gate/up GEMMs, SiLU product, and block-128 FP8 requantization in one launch. | |
| It requires `1<=M<=256`, `N%128==0`, and `K%128==0` and uses the upstream | |
| measured `32x128-w4-s1` tile. SM120 source regression remains 14/14. SM89 | |
| installed correctness, tile parity, and performance claims remain gated on an | |
| SM89 release artifact run; source presence alone is not recorded as runtime | |
| validation. | |