|
Download VALIDATION.md from flashrt/flashrt-adaptive-norms: direct link, hf CLI and curl.
- Browser
- Download file 1.4 kB
-
https://huggingface.co/flashrt/flashrt-adaptive-norms/resolve/main/VALIDATION.md
- Command line
-
hf download hf://flashrt/flashrt-adaptive-norms/VALIDATION.md
-
curl -L -o VALIDATION.md https://huggingface.co/flashrt/flashrt-adaptive-norms/resolve/main/VALIDATION.md
1.4 kB
Validation: flashrt-adaptive-norms
Required before publishing this package:
Source-extension correctness:
python flashrt-adaptive-norms/tests/test_adaptive_norms.py \ --backend source \ --mode fullSource-extension benchmark:
python flashrt-adaptive-norms/benchmarks/benchmark.py \ --backend source \ --shapes allKernel-builder artifact build:
kernel-builder build-and-copy flashrt-adaptive-normsBuilt-artifact correctness:
PYTHONPATH=<artifact-path> \ python flashrt-adaptive-norms/tests/test_adaptive_norms.py \ --backend installed \ --mode fullBuilt-artifact benchmark:
python flashrt-adaptive-norms/benchmarks/benchmark.py \ --backend installed \ --artifact <artifact-path> \ --shapes allMulti-hardware matrix:
Add hardware claims only after the same correctness and benchmark commands pass on that machine.
Style Broadcast Gate
Both APIs accept style as either (rows, 3*dim) or (1, 3*dim). The full
source and installed-artifact runs must execute both layouts at rows
64/2520/4096. The broadcast path additionally requires bit-identical gate
output, exact residual update, FP8 p99 absolute error zero, CUDA Graph replay,
and installed-wrapper torch.compile(fullgraph=True) parity.