Uploaded using `kernel-builder`.
Browse files
README.md
ADDED
|
@@ -0,0 +1,46 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# atlas-gdn
|
| 2 |
+
|
| 3 |
+
Hand-tuned Gated DeltaNet kernels for the linear-attention path of
|
| 4 |
+
Qwen3.6 hybrid models on NVIDIA GB10 (DGX Spark, SM121).
|
| 5 |
+
|
| 6 |
+
## What's inside
|
| 7 |
+
|
| 8 |
+
| Op | Use |
|
| 9 |
+
|-----------------------------|--------------------------------------------------------|
|
| 10 |
+
| `gdn_decode` | Single-token recurrent decode (FP32 Q/K/V, BF16 out) |
|
| 11 |
+
| `gdn_prefill` | Multi-token prefill (BF16 throughout) |
|
| 12 |
+
| `gdn_chunk2` / `gdn_chunk3` | MTP K=2/3 chunkwise verify (Qwen3.6 NVFP4 specialized) |
|
| 13 |
+
| `gdn_wy2` / `wy3` / `wy4` | 2-pass WY-chunkwise verify (general K=2/3/4) |
|
| 14 |
+
| `causal_conv1d_fwd` | Depthwise causal Conv1d (SSM input projection) |
|
| 15 |
+
| `causal_conv1d_update` | Single-step Conv1d update (decode) |
|
| 16 |
+
|
| 17 |
+
## Hardware
|
| 18 |
+
|
| 19 |
+
These kernels target **only** NVIDIA GB10 (compute capability 12.1,
|
| 20 |
+
`sm_121f`). They will not load on any other GPU. GB10 has:
|
| 21 |
+
|
| 22 |
+
- Unified LPDDR5X memory (~273 GB/s) — bandwidth-bound, not occupancy-bound
|
| 23 |
+
- No multi-CTA clusters (ClusterShape forced to 1×1×1)
|
| 24 |
+
- No `cvt.rn.satfinite.e2m1x2.f32` PTX (software E2M1 conversion path)
|
| 25 |
+
- Cooperative-only scheduling (no Pingpong)
|
| 26 |
+
|
| 27 |
+
`build.toml` pins `cuda-capabilities = ["12.1"]` so the build matrix
|
| 28 |
+
yields a single SM121 binary; no fallback binaries are produced.
|
| 29 |
+
|
| 30 |
+
## Models tested
|
| 31 |
+
|
| 32 |
+
| Model | Layers using these kernels |
|
| 33 |
+
|-----------------------------------------|----------------------------|
|
| 34 |
+
| Qwen/Qwen3.6-27B (dense, hybrid) | 48 GDN layers |
|
| 35 |
+
| Qwen/Qwen3.6-35B-A3B (sparse MoE, hybrid) | 30 GDN layers |
|
| 36 |
+
|
| 37 |
+
## Provenance
|
| 38 |
+
|
| 39 |
+
Sources are extracted from the Atlas inference engine
|
| 40 |
+
(<https://github.com/Avarok-Cybersecurity/atlas>, AGPL-3.0). The GDN
|
| 41 |
+
NVFP4 variant ships with `__launch_bounds__` annotations specific to
|
| 42 |
+
Qwen3.6 hidden dimensions (k_dim=128, v_dim=128, 16/32 K/V heads).
|
| 43 |
+
|
| 44 |
+
## License
|
| 45 |
+
|
| 46 |
+
AGPL-3.0-only.
|