Instructions to use SuperexponentialAI/relu with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Kernels
How to use SuperexponentialAI/relu with Kernels:
# !pip install kernels from kernels import get_kernel kernel = get_kernel("SuperexponentialAI/relu") - Notebooks
- Google Colab
- Kaggle
relu kernel β measured performance to date
Consolidated view of every backend/platform measured so far. Detailed
methodology per platform: OPTIMIZATION.md (4090), H100_BENCHMARK.md
(H100), DTYPE_THROUGHPUT.md (dtype scaling), INVESTIGATION_L2.md
(timing pitfalls), README.md (CPU/XPU). All GPU numbers are nsys
GPU-kernel durations, interleaved 1:1, correctness-gated (bit-identical
to F.relu; NaN β 0 matches upstream).
CUDA β RTX 4090 (Ada, sm_89, tuned here) β corrected in v3
v3 retraction + fix (see L2_BAND_CLIFF.md): the previously published
fp16/bf16 "1.85x/1.81x" and fp32 "1.20x DRAM" vs torch were artifacts of a
fixed-order shared-buffer interleave, and v2 additionally had a real
launch-config cliff (0.34β0.65x vs torch) for working sets in
(32 MiB, 72 MiB L2]. v3 numbers, fair buffer handling:
| comparison | 1024Β² (L2) | 32 MiBβL2 band | DRAM-bound |
|---|---|---|---|
| vs torch.relu fp32 | ~parity (0.95β1.04x) | 0.99β1.06x (was 0.34β0.65x) | ~parityβ1.09x |
| vs torch fp16 / bf16 | 1.17β1.27x | ~parity | ~parity |
| vs torch int8 | 1.35β1.37x | 1.14x | ~parity |
| vs upstream Hub (fp32) | 0.96β1.22x (run-sensitive) | β | parity (both at wall) |
DRAM wall ~944 GB/s (93.7% of 1008 peak), dtype-invariant β every config hits it; upstream can't do fp16/bf16/int8 at all, which remains the main 4090 advantage over the Hub kernel.
CUDA β H100 80GB HBM3 (sm_90, incl. same-day retune) [2026-07-30]
L2-resident vs-torch numbers below are from separate-buffer symmetric
pairs (the shared-buffer 3-way suite's 1.79x-class readings were
position-biased β see L2_BAND_CLIFF.md); DRAM-bound numbers are
position-insensitive and unchanged.
| comparison | 1024Β² (L2) | 4096Β² | 8192Β² |
|---|---|---|---|
| vs upstream Hub kernel (fp32) | 1.2β1.5x (run-sensitive) | 1.44x | 1.24x |
| vs torch fp32 | 1.52x | 1.10x | 1.01x |
| vs torch fp16 | 1.38x | 1.20x | 0.99x |
| vs torch bf16 | 1.38x | 1.16x | 1.01x |
| vs torch int8 | 1.23x | 1.61x | 1.16x |
sm_90 retune: +9.3% (pairwise A/B) at sub-16 MB working sets via a
Hopper-only unroll-2/stream-store/block-256 branch; larger sizes keep the
4090 config β the sweep's claimed 33β52 MB win did not survive in-harness
verification (see H100_BENCHMARK.md Round 2 for the methodology finding
on standalone-sweep and interleave-position bias).
Ours sits at the HBM3 wall (~3.0 TB/s, ~90% of 3,350 peak); upstream Hub never reaches it (2.2β2.4 TB/s) β the DRAM-bound margin grows on H100. Robustness: on flat 1-D tensors upstream launches grid=(1,1,1) and runs 140x slower than ours (its grid heuristic is shape-dependent; ours is a shape-agnostic grid-stride launch).
CPU β i9-13900K (AVX2 + OpenMP + non-temporal stores)
fp32/fp16/bf16/int8 all sit at the memory-bandwidth ceiling (thread/ILP sweep confirmed bandwidth-, not compute-, limited).
XPU β Intel UHD 770 iGPU (SYCL, 16 B/work-item), vs torch 2.12+xpu
| dtype | fp32 | fp16 | bf16 | int8 |
|---|---|---|---|---|
| ours vs torch | 1.01x | 0.96x | 1.16x | 1.06x |
Both bound by the same shared-DDR5 wall; vectorization gave 2.6x over the
naive SYCL version. Consumer iGPU caveats in ../AITESTING.md.
Dtype scaling (memory-bound law, both GPUs)
Element throughput scales 1/sizeof(dtype): bf16/fp16 β 2x fp32, int8/fp8
β 4x fp32 intrinsically; up to ~8β15x at L2-boundary working sets
(DTYPE_THROUGHPUT.md). Upstream Hub kernel is fp32-only; this build
ships fp32/fp16/bf16/int8 on all three backends.