nologik commited on
Commit
3f6ddf9
·
verified ·
1 Parent(s): 057d8be

Uploaded using `kernel-builder`.

Browse files
Files changed (1) hide show
  1. README.md +46 -0
README.md ADDED
@@ -0,0 +1,46 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # atlas-gdn
2
+
3
+ Hand-tuned Gated DeltaNet kernels for the linear-attention path of
4
+ Qwen3.6 hybrid models on NVIDIA GB10 (DGX Spark, SM121).
5
+
6
+ ## What's inside
7
+
8
+ | Op | Use |
9
+ |-----------------------------|--------------------------------------------------------|
10
+ | `gdn_decode` | Single-token recurrent decode (FP32 Q/K/V, BF16 out) |
11
+ | `gdn_prefill` | Multi-token prefill (BF16 throughout) |
12
+ | `gdn_chunk2` / `gdn_chunk3` | MTP K=2/3 chunkwise verify (Qwen3.6 NVFP4 specialized) |
13
+ | `gdn_wy2` / `wy3` / `wy4` | 2-pass WY-chunkwise verify (general K=2/3/4) |
14
+ | `causal_conv1d_fwd` | Depthwise causal Conv1d (SSM input projection) |
15
+ | `causal_conv1d_update` | Single-step Conv1d update (decode) |
16
+
17
+ ## Hardware
18
+
19
+ These kernels target **only** NVIDIA GB10 (compute capability 12.1,
20
+ `sm_121f`). They will not load on any other GPU. GB10 has:
21
+
22
+ - Unified LPDDR5X memory (~273 GB/s) — bandwidth-bound, not occupancy-bound
23
+ - No multi-CTA clusters (ClusterShape forced to 1×1×1)
24
+ - No `cvt.rn.satfinite.e2m1x2.f32` PTX (software E2M1 conversion path)
25
+ - Cooperative-only scheduling (no Pingpong)
26
+
27
+ `build.toml` pins `cuda-capabilities = ["12.1"]` so the build matrix
28
+ yields a single SM121 binary; no fallback binaries are produced.
29
+
30
+ ## Models tested
31
+
32
+ | Model | Layers using these kernels |
33
+ |-----------------------------------------|----------------------------|
34
+ | Qwen/Qwen3.6-27B (dense, hybrid) | 48 GDN layers |
35
+ | Qwen/Qwen3.6-35B-A3B (sparse MoE, hybrid) | 30 GDN layers |
36
+
37
+ ## Provenance
38
+
39
+ Sources are extracted from the Atlas inference engine
40
+ (<https://github.com/Avarok-Cybersecurity/atlas>, AGPL-3.0). The GDN
41
+ NVFP4 variant ships with `__launch_bounds__` annotations specific to
42
+ Qwen3.6 hidden dimensions (k_dim=128, v_dim=128, 16/32 K/V heads).
43
+
44
+ ## License
45
+
46
+ AGPL-3.0-only.