AzeezIsh commited on
Commit
a3ce773
·
verified ·
1 Parent(s): e19a44c

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +159 -113
README.md CHANGED
@@ -1,161 +1,207 @@
1
- ---
2
- title: Atlas Inference
3
- emoji: 🚀
4
- colorFrom: red
5
- colorTo: yellow
6
- sdk: static
7
- pinned: false
8
- license: agpl-3.0
9
- short_description: Pure Rust LLM Inference.
10
- ---
11
-
12
- <p align="center">
13
- <video src="https://huggingface.co/spaces/Atlas-Inference/README/resolve/main/atlas-demo.mov" controls muted playsinline width="820"></video>
14
- </p>
15
-
16
  <p align="center">
17
- <a href="https://x.com/AIshaqui81766/status/2052121270506930276"><strong>📣 Read the launch announcement on X →</strong></a>
18
  </p>
19
-
20
  <p align="center">
21
- <h1 align="center">Atlas Inference</h1>
 
 
 
 
22
  <p align="center">
23
- <strong>Pure Rust LLM Inference.</strong>
 
 
24
  </p>
25
  <p align="center">
26
- <a href="https://atlasinference.io"><img alt="Website" src="https://img.shields.io/badge/web-atlasinference.io-orange?style=flat-square"></a>
27
- <a href="https://github.com/Avarok-Cybersecurity/atlas"><img alt="GitHub" src="https://img.shields.io/badge/source-github-181717?style=flat-square&logo=github&logoColor=white"></a>
28
- <a href="https://hub.docker.com/r/avarok/atlas-gb10"><img alt="Docker Hub" src="https://img.shields.io/badge/Docker%20Hub-avarok%2Fatlas--gb10-2496ED?style=flat-square&logo=docker&logoColor=white"></a>
29
- <a href="https://discord.gg/DwF3brBMpw"><img alt="Discord" src="https://img.shields.io/badge/community-discord-5865F2?style=flat-square&logo=discord&logoColor=white"></a>
30
- <a href="https://github.com/Avarok-Cybersecurity/atlas/blob/master/LICENSE"><img alt="License: AGPLv3" src="https://img.shields.io/badge/license-AGPLv3-yellow?style=flat-square"></a>
31
  </p>
32
  </p>
33
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
34
  ---
35
 
36
- ## What is Atlas?
 
 
 
 
 
 
 
 
 
 
 
 
 
 
37
 
38
- Atlas is a from-scratch Rust + CUDA inference engine built for the next decade of LLM deployment. No Python interpreter. No PyTorch. No 20 GB Docker image. One ~2.5 GB binary that boots in under two minutes and pins the bandwidth ceiling on every supported (Hardware × Model × Quantization) target.
39
 
40
- We started on NVIDIA's DGX Spark (GB10 / SM121) with twelve hand-tuned model targets and a plug-and-play architecture designed so AMD, Intel, and Apple Silicon can land as community contributions, and so the next round of model families slot in the same way the Qwens did this quarter.
41
 
42
- ## Why Atlas
43
 
44
- | | Atlas | vLLM (same hardware) |
45
- | ---------------------------- | -------------------- | -------------------- |
46
- | Image size | **~2.5 GB** | 20+ GB |
47
- | Cold start | **<2 min** | ~10 min |
48
- | Runtime | **Rust + CUDA** | Python + PyTorch |
49
- | Dependencies | **None** | 200+ packages |
50
- | Peak Qwen3.5-35B (NVFP4) | **130 tok/s** | ~38 tok/s |
51
- | Average across workloads | **111 tok/s (3.0×)** | 37 tok/s |
 
 
 
 
 
 
 
 
 
 
 
 
 
52
 
53
- Same hardware. Same model weights. Bring your own benchmark — `scripts/sweep_all_models.sh` is in the repo and we publish the vLLM baseline command alongside ours so you can verify both. If you reproduce a faster vLLM number, file an issue. We would rather be measured than congratulated.
54
 
55
- ## Quick Start
 
 
 
 
 
 
 
 
 
 
56
 
57
  ```bash
58
- docker pull avarok/atlas-gb10:latest
 
59
 
60
- docker run --gpus all --ipc=host -p 8888:8888 \
61
- -v ~/.cache/huggingface:/root/.cache/huggingface \
62
- avarok/atlas-gb10:latest \
63
- serve Sehyo/Qwen3.5-35B-A3B-NVFP4 --speculative --mtp-quantization nvfp4
64
- ```
65
 
66
- Anything OpenAI- or Anthropic-compatible — `curl`, the OpenAI SDK, opencode, Claude Code, Cline, Open WebUI — points at port 8888:
 
 
67
 
 
68
  ```bash
69
- curl http://localhost:8888/v1/chat/completions \
70
- -H "Content-Type: application/json" \
71
- -d '{"model":"atlas","messages":[{"role":"user","content":"Hello!"}],"max_tokens":256}'
72
  ```
73
 
74
- Per-model recipes (vision, MoE, multi-node EP=2, single-GPU 122B with the tighter budget) live in [`QUICKSTART.md`](https://github.com/Avarok-Cybersecurity/atlas/blob/master/QUICKSTART.md).
75
 
76
- ## What Ships Today
 
 
 
 
 
 
 
77
 
78
- Thirteen hand-tuned (Hardware × Model × Quantization) targets across Qwen3 / Qwen3.5 / Qwen3.6 / Qwen3-Next / Qwen3-VL / Gemma-4 / Mistral / MiniMax / Nemotron-H families. Every supported model runs off one multi-model binary; the right kernel set is selected at startup from the model's `config.json`.
79
 
80
- | Model | Params / active | Quant | Architecture | Throughput |
81
- | --------------------------- | ------------------- | ----------- | ----------------------------- | --------------- |
82
- | Qwen3.5-35B-A3B (MTP K=2) | 35B / 3B | NVFP4 | GDN + Attention + MoE | **~130 tok/s** |
83
- | Qwen3-VL-30B-A3B | 30B / 3B | NVFP4 | Vision + Attention + MoE | ~97 tok/s |
84
- | Nemotron-3-Nano-30B-A3B | 30B / 3.5B | NVFP4 / FP8 | Mamba-2 + Attention + MoE | ~88 tok/s |
85
- | Qwen3-Next-80B-A3B | 80B / 3B | NVFP4 | SSM + Attention + MoE | ~74–87 tok/s |
86
- | Qwen3.6-35B-A3B | 35B / 3B | FP8 | GDN + Attention + MoE + ViT | ~71 tok/s |
87
- | Gemma-4-26B-A4B | 26B / 4B | NVFP4 | Attention + MoE (GeGLU) | ~67 tok/s |
88
- | Qwen3.5-122B-A10B (EP=2) | 122B / 10B | NVFP4 | GDN + Attention + MoE | ~46 tok/s |
89
- | Mistral-Small-4-119B | 119B / 6.5B | NVFP4 | MLA + MoE | ~33 tok/s |
90
- | Nemotron-3-Super-120B-A12B | 120B / 12B | NVFP4 / FP8 | Mamba-2 + Attention + MoE | ~24 tok/s |
91
- | MiniMax-M2.7 (EP=2) | 229B / ~10B | NVFP4 | Attention + 256-expert MoE | ~15 tok/s |
92
- | Qwen3.5-27B (dense hybrid) | 27B | NVFP4 | Hybrid SSM + Attention | ~13 tok/s |
93
- | Gemma-4-31B | 31B | NVFP4 | Attention (sliding + full) | ~9–11 tok/s |
94
 
95
- Full HuggingFace IDs, methodology, and the kernel-by-kernel comparison against PyTorch eager live in the [GitHub README](https://github.com/Avarok-Cybersecurity/atlas#readme).
96
 
97
- ## What Works Today
 
 
 
 
 
 
98
 
99
- | Component | Status |
100
- |---|---|
101
- | OpenAI- and Anthropic-compatible HTTP API (streaming + non-streaming) | ✅ |
102
- | Tool calling (Hermes, Qwen3-Coder, Mistral formats) with grammar-constrained decoding | ✅ |
103
- | Reasoning / thinking tokens with budget cap | ✅ |
104
- | Concurrent batched decode + per-batch CUDA graphs | ✅ |
105
- | MTP speculative decoding (K=2, pipelined verify) | ✅ |
106
- | Prefix caching via radix tree (RadixAttention) + SSM snapshot cache (Marconi) — 10× warm-cache TTFT | ✅ |
107
- | KV cache dtypes — BF16, FP8, NVFP4, turbo3, turbo4 | ✅ |
108
- | MoE routing up to 512 experts | ✅ |
109
- | Vision encoder (Qwen3-VL, Qwen3.6 ViT) | ✅ |
110
- | Multi-GPU expert parallelism (EP=2 over RoCEv2) | ✅ |
111
- | SLO-aware scheduling, chunked prefill, active context compaction | ✅ |
112
- | High-speed NVMe KV swap (sliding-window aware) | ✅ |
113
- | Auto OOM pre-flight + UVM fallback on host OOM | ✅ |
114
 
115
- ## Plug & Play Architecture
116
 
117
- Atlas is built around a small set of Rust traits and a kernel registry — each marked with 🔌 below is the abstraction boundary where a new integration plugs in without touching anything above or below it:
118
 
119
- | Plug Point | What It Abstracts | To Add Support |
120
- |---|---|---|
121
- | 🔌 `trait ModelWeightLoader` | HuggingFace → layer translation | Implement one struct + add a match arm in `factory.rs` |
122
- | 🔌 `trait TransformerLayer` | Per-layer compute (attn, SSM, MoE, FFN) | Compose existing primitives or implement a new layer type |
123
- | 🔌 `trait GpuBackend` | All GPU memory and kernel ops | Swap CUDA for another accelerator backend |
124
- | 🔌 `kernels/<hw>/<model>/<quant>/` | Hardware-tuned CUDA kernels | Drop a directory with `MODEL.toml` + `.cu` files; `build.rs` auto-discovers it |
125
- | 🔌 `trait CommBackend` | Multi-GPU collectives | Implement for MPI, GDR, custom interconnects |
126
- | 🔌 `trait StorageBackend` | NVMe KV-cache offload I/O | Implement for CXL, RDMA, other storage tiers |
127
 
128
- A `MockGpuBackend` in `spark-runtime` lets you write and test the entire scaffold without owning the hardware — every layer above the GPU trait is hardware-agnostic.
 
129
 
130
- ## What the Community is Saying
 
 
131
 
132
- > *"103 tok/s sustained on the 35B, startup in 15 seconds. Night and day compared to vLLM's 10-minute torch.compile cycle. Then tried the 122B, 43.8 tok/s with MTP, a 41% speedup over our vLLM hybrid, same hardware, 2-minute startup."*
133
- > — **ronald_15496**, [Discord #general](https://discord.gg/DwF3brBMpw)
 
 
134
 
135
- > *"Testing atlas-qwen3.5-35b for over an hour on a PNY DGX Spark in an agentic workflow. Super impressed. Spark is actually awesome with Atlas."*
136
- > — **PersonWhoThinks**, [r/LocalLLaMA](https://www.reddit.com/r/LocalLLaMA/comments/1rmvxo3/)
137
 
138
- > *"I've grown tired of vLLM and have been hoping for something. I was really surprised and impressed. I'm so glad I bought Spark because I came across this."*
139
- > — **tetsuro59**, [Discord #general](https://discord.gg/DwF3brBMpw)
140
 
141
- ## Citations
142
 
143
- We did not invent the kernels we ship. We picked the right ideas from the right papers, fused them together, and tuned them for one chip until they pinned the bandwidth ceiling. Direct intellectual debts: **FlashAttention-2** (Dao, 2024), **FlashAttention-4** (Shah et al., 2025), **FlashInfer** (Ye et al., MLSys 2025), **SageAttention 3** (Zhang et al., NeurIPS 2025), **LeanAttention** (Roy et al., 2024). Full references in the [GitHub README](https://github.com/Avarok-Cybersecurity/atlas#citations).
144
 
145
- ## License & Enterprise Edition
 
 
 
 
 
 
 
 
146
 
147
- Atlas operates under a **dual-license** model. Both are real and intentional.
148
 
149
- 1. **Community Edition — AGPLv3.** Free, open, copyleft. Use it on your own hardware for research, hobby, side-projects, hosted demos.
150
- 2. **Enterprise Edition — commercial license.** Ship Atlas inside a closed-source product, run it as a SaaS backend without inheriting the AGPLv3 source-disclosure obligation, get a support relationship with the people who wrote the kernels, and prioritized model and hardware ports. Reach us via the [website](https://atlasinference.io) or Discord.
151
 
152
- A permissive license keeps us building Atlas full-time; the AGPL community license keeps the project honest. What is in this repository is what we run.
 
 
 
153
 
154
  ---
155
 
156
- <p align="center">
157
- <a href="https://atlasinference.io">atlasinference.io</a> ·
158
- <a href="https://github.com/Avarok-Cybersecurity/atlas">GitHub</a> ·
159
- <a href="https://hub.docker.com/r/avarok/atlas-gb10">Docker Hub</a> ·
160
- <a href="https://discord.gg/DwF3brBMpw">Discord</a>
161
- </p>
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  <p align="center">
2
+ <img src="assets/logo.svg" alt="Atlas Inference Engine" width="640" />
3
  </p>
 
4
  <p align="center">
5
+ <h1 align="center">Atlas Inference Engine</h1>
6
+ <p align="center">
7
+ <strong>Pure Rust & CUDA LLM Inference</strong><br>
8
+ <em>Universal Inference At Unimaginable Speeds</em>
9
+ </p>
10
  <p align="center">
11
+ <img alt="NVIDIA" src="https://img.shields.io/badge/NVIDIA-76B900?style=flat-square&logo=nvidia&logoColor=white">
12
+ <img alt="AMD" src="https://img.shields.io/badge/AMD-ED1C24?style=flat-square&logo=amd&logoColor=white">
13
+ <img alt="Intel" src="https://img.shields.io/badge/Intel-0071C5?style=flat-square&logo=intel&logoColor=white">
14
  </p>
15
  <p align="center">
16
+ <a href="LICENSE"><img alt="License: AGPLv3" src="https://img.shields.io/badge/license-AGPLv3-yellow?style=flat-square"></a>
17
+ <a href="#quick-start"><img alt="Pure Rust" src="https://img.shields.io/badge/runtime-pure%20Rust-orange?style=flat-square"></a>
18
+ <a href="https://hub.docker.com/r/azeezish/atlas-gb10:latest"><img alt="Docker Hub" src="https://img.shields.io/badge/Docker%20Hub-azeezish%2Fatlas--gb10-2496ED?style=flat-square&logo=docker&logoColor=white"></a>
19
+ <a href="https://discord.com/invite/6vDbKaKrKD"><img alt="Discord" src="https://img.shields.io/badge/dynamic/json?url=https%3A%2F%2Fdiscord.com%2Fapi%2Fv10%2Finvites%2F6vDbKaKrKD%3Fwith_counts%3Dtrue&query=%24.approximate_member_count&label=discord&suffix=%20members&style=flat-square&logo=discord&logoColor=white&color=5865F2"></a>
20
+ <a href="https://x.com/AtlasInference"><img alt="X / Twitter" src="https://img.shields.io/badge/X-%40AtlasInference-000000?style=flat-square&logo=x&logoColor=white"></a>
21
  </p>
22
  </p>
23
 
24
+ <p align="center">
25
+ <a href="assets/atlas-demo.mp4"><img alt="Atlas demo, click for full-quality MP4" src="assets/atlas-demo.gif" width="820" /></a>
26
+ </p>
27
+
28
+ <p align="center">
29
+ <a href="#quick-start"><img alt="Quick Start — under 2 minutes" src="https://img.shields.io/badge/%E2%9A%A1%20Quick%20Start%20%E2%80%94%20%3C%202%20min-2EA44F?style=for-the-badge&logo=docker&logoColor=white"></a>
30
+ <a href="https://atlasinference.dev"><img alt="atlasinference.dev" src="https://img.shields.io/badge/%F0%9F%8C%90%20atlasinference.dev-F48C06?style=for-the-badge"></a>
31
+ <a href="https://mlcommons.org/2026/07/mlperf-inference-v61-edge-agentic/"><img alt="MLPerf v6.1 Edge Agentic" src="https://img.shields.io/badge/MLPerf%20v6.1-Edge%20Agentic-blue?style=for-the-badge"></a>
32
+ <a href="docs/GB10_DEPLOYMENT_GUIDE.md"><img alt="Deployment Guide" src="https://img.shields.io/badge/%F0%9F%93%96%20GB10%20Deployment%20Guide-4A154B?style=for-the-badge"></a>
33
+ </p>
34
+
35
+ ---
36
+
37
+ ## ⚡ What is Atlas?
38
+
39
+ Atlas is a high-performance, pure Rust & CUDA LLM inference engine purpose-built for prosumer workstations (NVIDIA DGX Spark / GB10 SM121 and AMD Strix Halo). No Python, no PyTorch, no bloated dependency trees—just one compact binary with hand-tuned micro-kernels.
40
+
41
+ - **Sub-90s First Token**: Boots in seconds with cached weights; zero JIT compile or Python startup lag.
42
+ - **Default Flagship Qwen 3.8 27B**: Dense hybrid GDN + Attention running at 23.59 tok/s single-stream with MTP speculative decoding on a single GB10.
43
+ - **Nemotron 3.5 Lightning + DSpark**: Full bring-up of hybrid Mamba-2 SSM + MoE paired with DSpark speculative decoding drafters for sub-10ms token decode latencies.
44
+ - **Qwen 3.8 Flash-Next Support**: Stream massive ~180B hybrid MoE models inside ~90 GB resident VRAM using direct parallel `pread` NVMe offloading.
45
+ - **Turnkey Sparkrun Integration**: Launch any verified model recipe instantly with `sparkrun run @atlas/<recipe>`.
46
+ - **OpenAI & Anthropic Compatible**: Drop-in API endpoint supporting streaming, tool calling, and reasoning traces.
47
+ - **MLPerf Proven**: Official contributor to the MLPerf Inference v6.1 Edge Agentic benchmark.
48
+
49
  ---
50
 
51
+ <a id="philosophy"></a>
52
+
53
+ ## 🧭 Philosophy
54
+
55
+ Atlas began as a solution to a widely known problem with Python inference engines. The code was steeped in an ever-shifting ecosystem of dependencies, patches, and cross-dependencies. One day your workaround for running a model works. The next, you are updating to a nightly branch of several dependencies and injecting a new workaround. That is how you build a proof of concept, not a software ecosystem. Data scientists proved LLMs can revolutionize our world. Now software engineers turn the proof of concept into something designed to withstand the test of time. The main objective mirrors what llama.cpp proved for cheap GPUs. As hardware advances, nobody should have to pay premium Cloud API prices for inference. Atlas maximizes speed for each hardware and model combination so a meaningfully powerful local model is truly useful.
56
+
57
+ - **Free and open source, always.** Great software comes from opening the source. The more eyes the better.
58
+ - **Community-first.** We build for you, and with source in hand you build for others in ways that triumph over existing solutions. We are the Pirates of the inference space.
59
+ - **Monorepo.** Everything in one place, so a data scientist, engineer, or agent can land a meaningful PR anywhere in the stack.
60
+ - **Hardware and model specific kernels.** No compromises or generalizations. Each combination gets fine-tuned custom kernels, which is where the 2-3x speedups come from.
61
+ - **AI-friendly codebase.** Built with enough railguards, structure, and abstraction for an AI to absorb the monorepo and contribute meaningfully. Fork it, point your agent at a model, and get a working port in hours instead of weeks. **AI-authored PRs are the default, and the target.** If you write code by hand, say which parts and why the human beat the AI. Every such case marks a gap in tooling we would rather close. The contribution loop lives in [`CONTRIBUTING.md`](CONTRIBUTING.md#pull-request-process).
62
+ - **Theory-friendly.** Relevant arXiv results are welcome as PoC PRs. Explain what you did and why.
63
+ - **Plug-and-play design.** Modular traits keep the business logic identical across hardware and model combinations. Only the concrete implementations differ. The extension points are `ModelWeightLoader`, `TransformerLayer`, `GpuBackend`, `CommBackend`, `StorageBackend`, and the `kernels/<hw>/<model>/<quant>/` directory convention.
64
+
65
+ <a id="models"></a>
66
 
67
+ ## 📦 What We Ship Today
68
 
69
+ One multi-model binary. The right kernel set is selected at startup from the model's `config.json`, so there is no swapping images and no rebuilding. Point Atlas at a HuggingFace ID, or launch a verified recipe with `sparkrun run @atlas/<recipe>`. The recipe SSOT is [sparkrun-recipes](https://github.com/Atlas-Inf/sparkrun-recipes) and the interactive browser lives at [atlasinference.dev/#models](https://atlasinference.dev/#models).
70
 
71
+ Numbers are single-GB10 decode tok/s measured end-to-end through the HTTP API on a short prompt (`max_tokens ≤ 30`, `temperature = 0.1`), reproducible via `scripts/sweep_all_models.sh`.
72
 
73
+ | Model | HuggingFace ID | Params / active | Architecture | Recipe | tok/s |
74
+ |---|---|---|---:|---|---:|
75
+ | **Qwen3.8-27B flagship** | `nvidia/Qwen3.8-27B-NVFP4` | 27B dense | GDN + attention hybrid, MTP | `@atlas/qwen3.8-27b-nvfp4` | 23.59 |
76
+ | **Qwen3.8-Flash-Next** | `RadixArk/Qwen3.8-Flash-Next-NVFP4` | ~180B MoE | GDN + attention + MoE, PLE NVMe streaming | `@atlas/qwen3.8-flash-next-nvfp4` | 36.7 |
77
+ | Nemotron-3.5-Lightning + DSpark | `nvidia` checkpoint via recipe | 30B / 3B | Mamba-2 + attention + MoE + DSpark drafter | `@atlas/nemotron-3.5-lightning-30b-a3b-nvfp4-dspark` | sub-10ms/token |
78
+ | Qwen3.5-27B | `Kbenkhaled/Qwen3.5-27B-NVFP4` | 27B dense | Hybrid SSM + attention, dense FFN, MRoPE | | 13 |
79
+ | Qwen3.5-35B-A3B | `Sehyo/Qwen3.5-35B-A3B-NVFP4` | 35B / 3B | GDN + attention + MoE, MTP | | **131** (MTP K=2) |
80
+ | Qwen3.5-122B-A10B | `Sehyo/Qwen3.5-122B-A10B-NVFP4` | 122B / 10B | GDN + attention + MoE, MTP | | 46 (EP=2) |
81
+ | Qwen3.6-35B-A3B | `Qwen/Qwen3.6-35B-A3B-FP8` | 35B / 3B | GDN + attention + MoE, MRoPE, vision | `@atlas/qwen3.6-35b-a3b-fp8-mtp` | |
82
+ | Holo-3.1-35B-A3B | `Hcompany/Holo-3.1-35B-A3B-NVFP4` | 35B / 3B | GDN + attention + 256-expert MoE, Qwen3-VL vision | | |
83
+ | Holo-3.1-0.8B | `Hcompany/Holo-3.1-0.8B` | 0.8B dense | GDN + attention + dense FFN, Qwen3-VL vision | | |
84
+ | Ornith-1.0-9B | `deepreinforce-ai/Ornith-1.0-9B` | 9B dense | GDN + attention + dense FFN, Qwen3-VL vision, MRoPE | | |
85
+ | Qwen3-Next-80B-A3B | `nvidia/Qwen3-Next-80B-A3B-Instruct-NVFP4` | 80B / 3B | SSM + attention + MoE | | 74 |
86
+ | Qwen3-VL-30B-A3B | `ig1/Qwen3-VL-30B-A3B-Instruct-NVFP4` | 30B / 3B | Vision + attention + MoE | | 97 |
87
+ | Gemma-4-26B-A4B | `bg-digitalservices/Gemma-4-26B-A4B-it-NVFP4A16` | 26B / 4B | Attention + MoE, GeGLU | `@atlas/gemma-4-26b-a4b-nvfp4` | 67 |
88
+ | Gemma-4-31B | `nvidia/Gemma-4-31B-IT-NVFP4` | 31B dense | Attention (sliding + full), GeGLU | | 9 |
89
+ | Mistral-Small-4-119B | `mistralai/Mistral-Small-4-119B-2603-NVFP4` | 119B / 6.5B | Attention + MoE | | 33 |
90
+ | MiniMax-M2.7 | `lukealonso/MiniMax-M2.7-NVFP4` | 229B / ~10B | Attention + 256-expert MoE + MTP | | |
91
+ | Nemotron-3-Nano-30B-A3B | `nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4` | 30B / 3B | Mamba-2 + attention + MoE | `@atlas/nemotron-3-nano-30b-a3b-nvfp4` | 88 |
92
+ | Nemotron-3-Super-120B-A12B | `nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4` | 120B / 12B | Mamba-2 + attention + MoE | `@atlas/nemotron-3-super-120b-a12b-nvfp4` | 24 |
93
+ | DeepSeek-V4-Flash | via recipe | | Attention + MoE | `@atlas/deepseek-v4-flash-nvfp4-ep2` | EP=2, 2 Sparks |
94
 
95
+ Flagship and Flash-Next also ship `-latency` and `-throughput` recipe variants for single-stream latency and 1 to 128 stream concurrency. On Qwen3.5-35B-A3B with MTP speculative decoding, Atlas decodes faster than NVIDIA's own vLLM build on the same hardware, on numbers we can hand you the script for. If you reproduce a faster vLLM number, file an issue. We would rather be measured than congratulated. The kernel-by-kernel comparison against PyTorch eager lives in the [benchmarks chapter](book/src/operations/benchmarks.md).
96
 
97
+ And this is only the beginning. The plug-and-play design has already carried AMD into the garden, Apple Silicon and Intel are next, and next quarter's models will slot in the same way the Qwens did this quarter.
98
+
99
+ > **New to Atlas on a Spark?** The [**GB10 Deployment & Compatibility Guide**](docs/GB10_DEPLOYMENT_GUIDE.md) is the one page to read first. It covers which model and quant fit your box, what to do when it OOMs, the known gotchas, and what verified means, then hands you the exact recipe.
100
+
101
+ <a id="performance"></a>
102
+
103
+ <a id="quick-start"></a>
104
+
105
+ ## 🚀 Quick Start
106
+
107
+ ### 1. NVIDIA GB10, flagship recipe `@atlas/qwen3.8-27b-nvfp4`
108
 
109
  ```bash
110
+ # Step 1, install sparkrun
111
+ pip install sparkrun # or uvx sparkrun setup install
112
 
113
+ # Step 2, download weights (or let sparkrun fetch automatically)
114
+ huggingface-cli download unsloth/Qwen3.8-27B-NVFP4 \
115
+ --local-dir ~/.cache/huggingface/hub/models--unsloth--Qwen3.8-27B-NVFP4
 
 
116
 
117
+ # Step 3, launch the flagship (port 8888)
118
+ sparkrun run @atlas/qwen3.8-27b-nvfp4 --hosts localhost
119
+ ```
120
 
121
+ Prefer the one-line quickstart script?
122
  ```bash
123
+ curl -fsSL https://atlasinference.dev/quickstart.sh | sh
 
 
124
  ```
125
 
126
+ The other flagships work the same way. The recipe is the command.
127
 
128
+ ```bash
129
+ # Nemotron 3.5 Lightning 30B + DSpark drafter, sub-10ms decode
130
+ export ATLAS_DFLASH_OPTION_B=1 ATLAS_NO_TOOL_INJECT=1
131
+ sparkrun run @atlas/nemotron-3.5-lightning-30b-a3b-nvfp4-dspark --hosts localhost
132
+
133
+ # Qwen 3.8 Flash-Next ~180B, streams the 47.7 GB PLE table off NVMe (~90 GB resident)
134
+ sparkrun run @atlas/qwen3.8-flash-next-nvfp4 --hosts localhost
135
+ ```
136
 
137
+ ### 2. AMD Strix Halo (gfx1151) on Linux and Windows
138
 
139
+ Strix Halo is a unified-memory APU, so the validated path is the native binary built against ROCm with no container. Both legs merged into main via [PR #8](https://github.com/Atlas-Inf/atlas/pull/8) (Linux) and [PR #9](https://github.com/Atlas-Inf/atlas/pull/9) (Windows).
 
 
 
 
 
 
 
 
 
 
 
 
 
140
 
141
+ **Linux** (Ubuntu 24.04, ROCm 6.2+)
142
 
143
+ ```bash
144
+ git clone https://github.com/Atlas-Inf/atlas.git
145
+ cd atlas
146
+ ./build-amd.sh # strix-hip backend, all targets, needs ROCm + cargo
147
+ ./serve-amd.sh # nvidia/Qwen3.8-27B-NVFP4, validated config
148
+ ./serve-amd.sh unsloth/Qwen3.8-27B-NVFP4 # or the preservation checkpoint
149
+ ```
150
 
151
+ Defaults are the validated config. K=4 MTP speculative, the W4A8 DP4A decode arm, BF16 KV. `NUM_DRAFTS=0` disables speculation, `LM_HEAD=bf16` switches for the unsloth checkpoint's per-row-FP8 lm_head, and `ATLAS_W4A16_DP4A=0` opts out of DP4A. The `ATLAS_FP8_DEQUANT_*` and `ATLAS_GDN_BF16_WEIGHTS` exports are baked in and required for unsloth.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
152
 
153
+ Measured on a Ryzen AI Max+ 395 with Radeon 8060S under ROCm 7.13 at ~60 GB GTT. **28.3 to 28.6 tok/s** K=4 decode, 13.3 to 13.6 tok/s at 30k context, and **83.02 / 80.41** on the 995-row bfcl-subset golden draw, matching the NVIDIA shipped reference (83.22 / 79.02). Details in [`BENCH.toml`](kernels/strix-hip/qwen3.8-27b/BENCH.toml) and [`amd-strix-halo-scale.md`](docs/porting/amd-strix-halo-scale.md).
154
 
155
+ **Windows** (native HIP on Windows 11, no WSL, no container. `spark.exe` is MSVC-built and reaches the GPU through a CUDA to HIP shim over ROCm 10 / TheRock)
156
 
157
+ ```powershell
158
+ git clone https://github.com/Atlas-Inf/atlas.git
159
+ cd atlas
 
 
 
 
 
160
 
161
+ hf download nvidia/Qwen3.8-27B-NVFP4 `
162
+ --local-dir "$env:USERPROFILE\models\nvidia-Qwen3.8-27B-NVFP4"
163
 
164
+ # Build from PowerShell, not Git Bash (bash puts coreutils link.exe ahead
165
+ # of MSVC's). Needs MSVC Desktop C++, the ROCm SDK (HIP_PATH), and cargo.
166
+ .\build-amd.ps1
167
 
168
+ # Serve, the fingerprinted fp8d recipe. Native FP8 GDN, K=4 MTP, BF16 KV
169
+ # and lm_head. Preflight hard-fails on non-gfx1151 GPUs.
170
+ .\serve-amd.ps1
171
+ ```
172
 
173
+ Prebuilt instead? Unzip `spark-windows-x86_64-amd-hip` keeping every DLL beside `spark.exe`, set `$env:ATLAS_BIN` to it, and run `.\serve-amd.ps1`. That skips the toolchain check and the build. The script runs a detached smoke probe (`first_run_smoke.log`), and the 64K-context record config is [`win_serve_qwen38_nvfp4.ps1`](scripts/strix-windows/win_serve_qwen38_nvfp4.ps1).
 
174
 
175
+ Measured on a Framework Desktop (Ryzen AI Max+ 395, Radeon 8060S, Windows 11, ROCm 10.0.0 TheRock, Adrenalin 32.0.31041.1004). **83.32 / 78.70** on the bfcl-subset golden draw with zero faults in 3.6 hours, and **17.25 tok/s** decode with MTP engaged at p1 0.834. Provenance in [`BENCH.toml`](kernels/strix-hip/qwen3.8-27b/BENCH.toml) and [`STRIX_WINDOWS_HIP.md`](docs/porting/STRIX_WINDOWS_HIP.md).
 
176
 
177
+ ### Hitting the Endpoint
178
 
179
+ Atlas speaks OpenAI, Anthropic, and Responses APIs on the same port. `curl`, the OpenAI SDK, Open WebUI, opencode, Cline, and Claude Code all work when pointed at the served port. `sparkrun` recipes default to 8888 and the AMD scripts default to 8081.
180
 
181
+ ```bash
182
+ curl http://localhost:8888/v1/chat/completions \
183
+ -H "Content-Type: application/json" \
184
+ -d '{
185
+ "model": "atlas",
186
+ "messages": [{"role": "user", "content": "Explain quantum computing in three sentences."}],
187
+ "max_tokens": 256
188
+ }'
189
+ ```
190
 
191
+ ---
192
 
193
+ ## 🤝 Community & Support
 
194
 
195
+ - **Website**: [atlasinference.dev](https://atlasinference.dev)
196
+ - **Discord**: [Join our Discord](https://discord.com/invite/6vDbKaKrKD) — Active daily development, live kernel tuning, and model requests.
197
+ - **Recipes Repository**: [Atlas-Inf/sparkrun-recipes](https://github.com/Atlas-Inf/sparkrun-recipes)
198
+ - **Deployment Guide**: [GB10 Deployment Guide](docs/GB10_DEPLOYMENT_GUIDE.md)
199
 
200
  ---
201
 
202
+ ## ⚖️ Dual License
203
+
204
+ - **Community Edition**: Licensed under **AGPLv3**. Free and open for personal use, research, and non-commercial local deployments.
205
+ - **Enterprise Edition**: Commercial licensing for proprietary applications, SaaS hosting without AGPLv3 copyleft obligations, dedicated support, and custom hardware/kernel porting. Contact `debaterishaqui@gmail.com`.
206
+
207
+ <sub><b>Continuity notice.</b> Atlas is continuing. This repository, the <a href="https://github.com/Atlas-Inf">Atlas-Inf</a> GitHub organization, and <a href="https://atlasinference.dev">atlasinference.dev</a> are the replacement official Atlas channels. The existing website and GitHub repository remain disputed Atlas assets that have not been relinquished.</sub>