|
Download benchmarks/RESULTS.md from flashrt/fp4-gemm: direct link, hf CLI and curl.
- Browser
- Download file 2.42 kB
-
https://huggingface.co/flashrt/fp4-gemm/resolve/main/benchmarks/RESULTS.md
- Command line
-
hf download hf://flashrt/fp4-gemm/benchmarks/RESULTS.md
-
curl -L -o RESULTS.md https://huggingface.co/flashrt/fp4-gemm/resolve/main/benchmarks/RESULTS.md
2.42 kB
| # fp4-gemm Benchmark Results | |
| Installed kernel-builder artifact benchmark on NVIDIA GeForce RTX 5090, | |
| PyTorch `2.11.0+cu128`. | |
| Command: | |
| ```bash | |
| python fp4-gemm/benchmarks/benchmark.py \ | |
| --backend installed \ | |
| --artifact fp4-gemm/build/torch211-cxx11-cu128-x86_64-linux \ | |
| --mode headline \ | |
| --warmup 100 \ | |
| --iterations 500 \ | |
| --json-out internal-tests/fp4-gemm-installed-benchmark.json | |
| ``` | |
| Reference is PyTorch GEMM over the same dequantized FP4/SFA and FP4/SFB inputs | |
| that the FlashRT kernel consumes. | |
| | Shape | Variant | FlashRT us | Eager us | Compile us | vs eager | vs compile | Max abs | | |
| | --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | | |
| | M=16, N=128, K=128 | 0 | 6.156 | 15.156 | 27.748 | 2.46x | 4.51x | 0.0 | | |
| | M=16, N=128, K=128 | 1 | 6.152 | 15.156 | 27.748 | 2.46x | 4.51x | 0.0 | | |
| | M=16, N=128, K=128 | 2 | 6.145 | 15.156 | 27.748 | 2.47x | 4.52x | 0.0 | | |
| | M=32, N=256, K=256 | 0 | 6.153 | 16.685 | 35.690 | 2.71x | 5.80x | 0.0 | | |
| | M=32, N=256, K=256 | 1 | 8.201 | 16.685 | 35.690 | 2.03x | 4.35x | 0.0 | | |
| | M=32, N=256, K=256 | 2 | 6.147 | 16.685 | 35.690 | 2.71x | 5.81x | 0.0 | | |
| | M=64, N=512, K=512 | 0 | 6.152 | 16.480 | 36.205 | 2.68x | 5.89x | 0.0 | | |
| | M=64, N=512, K=512 | 1 | 10.246 | 16.480 | 36.205 | 1.61x | 3.53x | 0.0 | | |
| | M=64, N=512, K=512 | 2 | 6.152 | 16.480 | 36.205 | 2.68x | 5.89x | 0.0 | | |
| Variant notes: | |
| - `variant=0` is the stable default. | |
| - `variant=1` is the widen schedule intended for very large `N`; it is not the | |
| best choice for these small validation shapes. | |
| - `variant=2` is competitive on small shapes and remains exposed for explicit | |
| A/B testing. | |
| The PyTorch references consume the same already-dequantized FP4 tensors and do | |
| not include quantization. The compiled reference is warmed before timing. | |
| ## BF16 Direct Producer | |
| Source benchmark on RTX 5090 with 100 warmup and 1000 measured iterations: | |
| | Shape | Direct BF16 us | Cast + FP16 producer us | Speedup | Native BF16 us | Wrapper/native | | |
| | --- | ---: | ---: | ---: | ---: | ---: | | |
| | M=1, K=5120 | 4.098 | 6.404 | 1.563x | 6.150 | 0.666x | | |
| | M=1, K=6144 | 4.098 | 6.403 | 1.562x | 8.190 | 0.500x | | |
| | M=1, K=17408 | 4.096 | 6.413 | 1.566x | 18.442 | 0.222x | | |
| The direct entry is byte-exact against the package's established | |
| BF16-to-FP16 plus FP16-producer contract. The native timing is reported as a | |
| performance reference only because that producer uses a distinct quantization | |
| strategy. | |