Spaces:
Configuration error
Configuration error
Update README.md
Browse files
README.md
CHANGED
|
@@ -1,161 +1,207 @@
|
|
| 1 |
-
---
|
| 2 |
-
title: Atlas Inference
|
| 3 |
-
emoji: 🚀
|
| 4 |
-
colorFrom: red
|
| 5 |
-
colorTo: yellow
|
| 6 |
-
sdk: static
|
| 7 |
-
pinned: false
|
| 8 |
-
license: agpl-3.0
|
| 9 |
-
short_description: Pure Rust LLM Inference.
|
| 10 |
-
---
|
| 11 |
-
|
| 12 |
-
<p align="center">
|
| 13 |
-
<video src="https://huggingface.co/spaces/Atlas-Inference/README/resolve/main/atlas-demo.mov" controls muted playsinline width="820"></video>
|
| 14 |
-
</p>
|
| 15 |
-
|
| 16 |
<p align="center">
|
| 17 |
-
<
|
| 18 |
</p>
|
| 19 |
-
|
| 20 |
<p align="center">
|
| 21 |
-
<h1 align="center">Atlas Inference</h1>
|
|
|
|
|
|
|
|
|
|
|
|
|
| 22 |
<p align="center">
|
| 23 |
-
<
|
|
|
|
|
|
|
| 24 |
</p>
|
| 25 |
<p align="center">
|
| 26 |
-
<a href="
|
| 27 |
-
<a href="
|
| 28 |
-
<a href="https://hub.docker.com/r/
|
| 29 |
-
<a href="https://discord.
|
| 30 |
-
<a href="https://
|
| 31 |
</p>
|
| 32 |
</p>
|
| 33 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 34 |
---
|
| 35 |
|
| 36 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 37 |
|
| 38 |
-
|
| 39 |
|
| 40 |
-
|
| 41 |
|
| 42 |
-
|
| 43 |
|
| 44 |
-
|
|
| 45 |
-
|
|
| 46 |
-
|
|
| 47 |
-
|
|
| 48 |
-
|
|
| 49 |
-
|
|
| 50 |
-
|
|
| 51 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 52 |
|
| 53 |
-
|
| 54 |
|
| 55 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 56 |
|
| 57 |
```bash
|
| 58 |
-
|
|
|
|
| 59 |
|
| 60 |
-
|
| 61 |
-
|
| 62 |
-
|
| 63 |
-
serve Sehyo/Qwen3.5-35B-A3B-NVFP4 --speculative --mtp-quantization nvfp4
|
| 64 |
-
```
|
| 65 |
|
| 66 |
-
|
|
|
|
|
|
|
| 67 |
|
|
|
|
| 68 |
```bash
|
| 69 |
-
curl
|
| 70 |
-
-H "Content-Type: application/json" \
|
| 71 |
-
-d '{"model":"atlas","messages":[{"role":"user","content":"Hello!"}],"max_tokens":256}'
|
| 72 |
```
|
| 73 |
|
| 74 |
-
|
| 75 |
|
| 76 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 77 |
|
| 78 |
-
|
| 79 |
|
| 80 |
-
|
| 81 |
-
| --------------------------- | ------------------- | ----------- | ----------------------------- | --------------- |
|
| 82 |
-
| Qwen3.5-35B-A3B (MTP K=2) | 35B / 3B | NVFP4 | GDN + Attention + MoE | **~130 tok/s** |
|
| 83 |
-
| Qwen3-VL-30B-A3B | 30B / 3B | NVFP4 | Vision + Attention + MoE | ~97 tok/s |
|
| 84 |
-
| Nemotron-3-Nano-30B-A3B | 30B / 3.5B | NVFP4 / FP8 | Mamba-2 + Attention + MoE | ~88 tok/s |
|
| 85 |
-
| Qwen3-Next-80B-A3B | 80B / 3B | NVFP4 | SSM + Attention + MoE | ~74–87 tok/s |
|
| 86 |
-
| Qwen3.6-35B-A3B | 35B / 3B | FP8 | GDN + Attention + MoE + ViT | ~71 tok/s |
|
| 87 |
-
| Gemma-4-26B-A4B | 26B / 4B | NVFP4 | Attention + MoE (GeGLU) | ~67 tok/s |
|
| 88 |
-
| Qwen3.5-122B-A10B (EP=2) | 122B / 10B | NVFP4 | GDN + Attention + MoE | ~46 tok/s |
|
| 89 |
-
| Mistral-Small-4-119B | 119B / 6.5B | NVFP4 | MLA + MoE | ~33 tok/s |
|
| 90 |
-
| Nemotron-3-Super-120B-A12B | 120B / 12B | NVFP4 / FP8 | Mamba-2 + Attention + MoE | ~24 tok/s |
|
| 91 |
-
| MiniMax-M2.7 (EP=2) | 229B / ~10B | NVFP4 | Attention + 256-expert MoE | ~15 tok/s |
|
| 92 |
-
| Qwen3.5-27B (dense hybrid) | 27B | NVFP4 | Hybrid SSM + Attention | ~13 tok/s |
|
| 93 |
-
| Gemma-4-31B | 31B | NVFP4 | Attention (sliding + full) | ~9–11 tok/s |
|
| 94 |
|
| 95 |
-
|
| 96 |
|
| 97 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 98 |
|
| 99 |
-
|
| 100 |
-
|---|---|
|
| 101 |
-
| OpenAI- and Anthropic-compatible HTTP API (streaming + non-streaming) | ✅ |
|
| 102 |
-
| Tool calling (Hermes, Qwen3-Coder, Mistral formats) with grammar-constrained decoding | ✅ |
|
| 103 |
-
| Reasoning / thinking tokens with budget cap | ✅ |
|
| 104 |
-
| Concurrent batched decode + per-batch CUDA graphs | ✅ |
|
| 105 |
-
| MTP speculative decoding (K=2, pipelined verify) | ✅ |
|
| 106 |
-
| Prefix caching via radix tree (RadixAttention) + SSM snapshot cache (Marconi) — 10× warm-cache TTFT | ✅ |
|
| 107 |
-
| KV cache dtypes — BF16, FP8, NVFP4, turbo3, turbo4 | ✅ |
|
| 108 |
-
| MoE routing up to 512 experts | ✅ |
|
| 109 |
-
| Vision encoder (Qwen3-VL, Qwen3.6 ViT) | ✅ |
|
| 110 |
-
| Multi-GPU expert parallelism (EP=2 over RoCEv2) | ✅ |
|
| 111 |
-
| SLO-aware scheduling, chunked prefill, active context compaction | ✅ |
|
| 112 |
-
| High-speed NVMe KV swap (sliding-window aware) | ✅ |
|
| 113 |
-
| Auto OOM pre-flight + UVM fallback on host OOM | ✅ |
|
| 114 |
|
| 115 |
-
|
| 116 |
|
| 117 |
-
|
| 118 |
|
| 119 |
-
|
| 120 |
-
|
| 121 |
-
|
| 122 |
-
| 🔌 `trait TransformerLayer` | Per-layer compute (attn, SSM, MoE, FFN) | Compose existing primitives or implement a new layer type |
|
| 123 |
-
| 🔌 `trait GpuBackend` | All GPU memory and kernel ops | Swap CUDA for another accelerator backend |
|
| 124 |
-
| 🔌 `kernels/<hw>/<model>/<quant>/` | Hardware-tuned CUDA kernels | Drop a directory with `MODEL.toml` + `.cu` files; `build.rs` auto-discovers it |
|
| 125 |
-
| 🔌 `trait CommBackend` | Multi-GPU collectives | Implement for MPI, GDR, custom interconnects |
|
| 126 |
-
| 🔌 `trait StorageBackend` | NVMe KV-cache offload I/O | Implement for CXL, RDMA, other storage tiers |
|
| 127 |
|
| 128 |
-
|
|
|
|
| 129 |
|
| 130 |
-
#
|
|
|
|
|
|
|
| 131 |
|
| 132 |
-
|
| 133 |
-
|
|
|
|
|
|
|
| 134 |
|
| 135 |
-
|
| 136 |
-
> — **PersonWhoThinks**, [r/LocalLLaMA](https://www.reddit.com/r/LocalLLaMA/comments/1rmvxo3/)
|
| 137 |
|
| 138 |
-
|
| 139 |
-
> — **tetsuro59**, [Discord #general](https://discord.gg/DwF3brBMpw)
|
| 140 |
|
| 141 |
-
##
|
| 142 |
|
| 143 |
-
|
| 144 |
|
| 145 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 146 |
|
| 147 |
-
|
| 148 |
|
| 149 |
-
|
| 150 |
-
2. **Enterprise Edition — commercial license.** Ship Atlas inside a closed-source product, run it as a SaaS backend without inheriting the AGPLv3 source-disclosure obligation, get a support relationship with the people who wrote the kernels, and prioritized model and hardware ports. Reach us via the [website](https://atlasinference.io) or Discord.
|
| 151 |
|
| 152 |
-
|
|
|
|
|
|
|
|
|
|
| 153 |
|
| 154 |
---
|
| 155 |
|
| 156 |
-
|
| 157 |
-
|
| 158 |
-
|
| 159 |
-
|
| 160 |
-
|
| 161 |
-
</
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
<p align="center">
|
| 2 |
+
<img src="assets/logo.svg" alt="Atlas Inference Engine" width="640" />
|
| 3 |
</p>
|
|
|
|
| 4 |
<p align="center">
|
| 5 |
+
<h1 align="center">Atlas Inference Engine</h1>
|
| 6 |
+
<p align="center">
|
| 7 |
+
<strong>Pure Rust & CUDA LLM Inference</strong><br>
|
| 8 |
+
<em>Universal Inference At Unimaginable Speeds</em>
|
| 9 |
+
</p>
|
| 10 |
<p align="center">
|
| 11 |
+
<img alt="NVIDIA" src="https://img.shields.io/badge/NVIDIA-76B900?style=flat-square&logo=nvidia&logoColor=white">
|
| 12 |
+
<img alt="AMD" src="https://img.shields.io/badge/AMD-ED1C24?style=flat-square&logo=amd&logoColor=white">
|
| 13 |
+
<img alt="Intel" src="https://img.shields.io/badge/Intel-0071C5?style=flat-square&logo=intel&logoColor=white">
|
| 14 |
</p>
|
| 15 |
<p align="center">
|
| 16 |
+
<a href="LICENSE"><img alt="License: AGPLv3" src="https://img.shields.io/badge/license-AGPLv3-yellow?style=flat-square"></a>
|
| 17 |
+
<a href="#quick-start"><img alt="Pure Rust" src="https://img.shields.io/badge/runtime-pure%20Rust-orange?style=flat-square"></a>
|
| 18 |
+
<a href="https://hub.docker.com/r/azeezish/atlas-gb10:latest"><img alt="Docker Hub" src="https://img.shields.io/badge/Docker%20Hub-azeezish%2Fatlas--gb10-2496ED?style=flat-square&logo=docker&logoColor=white"></a>
|
| 19 |
+
<a href="https://discord.com/invite/6vDbKaKrKD"><img alt="Discord" src="https://img.shields.io/badge/dynamic/json?url=https%3A%2F%2Fdiscord.com%2Fapi%2Fv10%2Finvites%2F6vDbKaKrKD%3Fwith_counts%3Dtrue&query=%24.approximate_member_count&label=discord&suffix=%20members&style=flat-square&logo=discord&logoColor=white&color=5865F2"></a>
|
| 20 |
+
<a href="https://x.com/AtlasInference"><img alt="X / Twitter" src="https://img.shields.io/badge/X-%40AtlasInference-000000?style=flat-square&logo=x&logoColor=white"></a>
|
| 21 |
</p>
|
| 22 |
</p>
|
| 23 |
|
| 24 |
+
<p align="center">
|
| 25 |
+
<a href="assets/atlas-demo.mp4"><img alt="Atlas demo, click for full-quality MP4" src="assets/atlas-demo.gif" width="820" /></a>
|
| 26 |
+
</p>
|
| 27 |
+
|
| 28 |
+
<p align="center">
|
| 29 |
+
<a href="#quick-start"><img alt="Quick Start — under 2 minutes" src="https://img.shields.io/badge/%E2%9A%A1%20Quick%20Start%20%E2%80%94%20%3C%202%20min-2EA44F?style=for-the-badge&logo=docker&logoColor=white"></a>
|
| 30 |
+
<a href="https://atlasinference.dev"><img alt="atlasinference.dev" src="https://img.shields.io/badge/%F0%9F%8C%90%20atlasinference.dev-F48C06?style=for-the-badge"></a>
|
| 31 |
+
<a href="https://mlcommons.org/2026/07/mlperf-inference-v61-edge-agentic/"><img alt="MLPerf v6.1 Edge Agentic" src="https://img.shields.io/badge/MLPerf%20v6.1-Edge%20Agentic-blue?style=for-the-badge"></a>
|
| 32 |
+
<a href="docs/GB10_DEPLOYMENT_GUIDE.md"><img alt="Deployment Guide" src="https://img.shields.io/badge/%F0%9F%93%96%20GB10%20Deployment%20Guide-4A154B?style=for-the-badge"></a>
|
| 33 |
+
</p>
|
| 34 |
+
|
| 35 |
+
---
|
| 36 |
+
|
| 37 |
+
## ⚡ What is Atlas?
|
| 38 |
+
|
| 39 |
+
Atlas is a high-performance, pure Rust & CUDA LLM inference engine purpose-built for prosumer workstations (NVIDIA DGX Spark / GB10 SM121 and AMD Strix Halo). No Python, no PyTorch, no bloated dependency trees—just one compact binary with hand-tuned micro-kernels.
|
| 40 |
+
|
| 41 |
+
- **Sub-90s First Token**: Boots in seconds with cached weights; zero JIT compile or Python startup lag.
|
| 42 |
+
- **Default Flagship Qwen 3.8 27B**: Dense hybrid GDN + Attention running at 23.59 tok/s single-stream with MTP speculative decoding on a single GB10.
|
| 43 |
+
- **Nemotron 3.5 Lightning + DSpark**: Full bring-up of hybrid Mamba-2 SSM + MoE paired with DSpark speculative decoding drafters for sub-10ms token decode latencies.
|
| 44 |
+
- **Qwen 3.8 Flash-Next Support**: Stream massive ~180B hybrid MoE models inside ~90 GB resident VRAM using direct parallel `pread` NVMe offloading.
|
| 45 |
+
- **Turnkey Sparkrun Integration**: Launch any verified model recipe instantly with `sparkrun run @atlas/<recipe>`.
|
| 46 |
+
- **OpenAI & Anthropic Compatible**: Drop-in API endpoint supporting streaming, tool calling, and reasoning traces.
|
| 47 |
+
- **MLPerf Proven**: Official contributor to the MLPerf Inference v6.1 Edge Agentic benchmark.
|
| 48 |
+
|
| 49 |
---
|
| 50 |
|
| 51 |
+
<a id="philosophy"></a>
|
| 52 |
+
|
| 53 |
+
## 🧭 Philosophy
|
| 54 |
+
|
| 55 |
+
Atlas began as a solution to a widely known problem with Python inference engines. The code was steeped in an ever-shifting ecosystem of dependencies, patches, and cross-dependencies. One day your workaround for running a model works. The next, you are updating to a nightly branch of several dependencies and injecting a new workaround. That is how you build a proof of concept, not a software ecosystem. Data scientists proved LLMs can revolutionize our world. Now software engineers turn the proof of concept into something designed to withstand the test of time. The main objective mirrors what llama.cpp proved for cheap GPUs. As hardware advances, nobody should have to pay premium Cloud API prices for inference. Atlas maximizes speed for each hardware and model combination so a meaningfully powerful local model is truly useful.
|
| 56 |
+
|
| 57 |
+
- **Free and open source, always.** Great software comes from opening the source. The more eyes the better.
|
| 58 |
+
- **Community-first.** We build for you, and with source in hand you build for others in ways that triumph over existing solutions. We are the Pirates of the inference space.
|
| 59 |
+
- **Monorepo.** Everything in one place, so a data scientist, engineer, or agent can land a meaningful PR anywhere in the stack.
|
| 60 |
+
- **Hardware and model specific kernels.** No compromises or generalizations. Each combination gets fine-tuned custom kernels, which is where the 2-3x speedups come from.
|
| 61 |
+
- **AI-friendly codebase.** Built with enough railguards, structure, and abstraction for an AI to absorb the monorepo and contribute meaningfully. Fork it, point your agent at a model, and get a working port in hours instead of weeks. **AI-authored PRs are the default, and the target.** If you write code by hand, say which parts and why the human beat the AI. Every such case marks a gap in tooling we would rather close. The contribution loop lives in [`CONTRIBUTING.md`](CONTRIBUTING.md#pull-request-process).
|
| 62 |
+
- **Theory-friendly.** Relevant arXiv results are welcome as PoC PRs. Explain what you did and why.
|
| 63 |
+
- **Plug-and-play design.** Modular traits keep the business logic identical across hardware and model combinations. Only the concrete implementations differ. The extension points are `ModelWeightLoader`, `TransformerLayer`, `GpuBackend`, `CommBackend`, `StorageBackend`, and the `kernels/<hw>/<model>/<quant>/` directory convention.
|
| 64 |
+
|
| 65 |
+
<a id="models"></a>
|
| 66 |
|
| 67 |
+
## 📦 What We Ship Today
|
| 68 |
|
| 69 |
+
One multi-model binary. The right kernel set is selected at startup from the model's `config.json`, so there is no swapping images and no rebuilding. Point Atlas at a HuggingFace ID, or launch a verified recipe with `sparkrun run @atlas/<recipe>`. The recipe SSOT is [sparkrun-recipes](https://github.com/Atlas-Inf/sparkrun-recipes) and the interactive browser lives at [atlasinference.dev/#models](https://atlasinference.dev/#models).
|
| 70 |
|
| 71 |
+
Numbers are single-GB10 decode tok/s measured end-to-end through the HTTP API on a short prompt (`max_tokens ≤ 30`, `temperature = 0.1`), reproducible via `scripts/sweep_all_models.sh`.
|
| 72 |
|
| 73 |
+
| Model | HuggingFace ID | Params / active | Architecture | Recipe | tok/s |
|
| 74 |
+
|---|---|---|---:|---|---:|
|
| 75 |
+
| **Qwen3.8-27B flagship** | `nvidia/Qwen3.8-27B-NVFP4` | 27B dense | GDN + attention hybrid, MTP | `@atlas/qwen3.8-27b-nvfp4` | 23.59 |
|
| 76 |
+
| **Qwen3.8-Flash-Next** | `RadixArk/Qwen3.8-Flash-Next-NVFP4` | ~180B MoE | GDN + attention + MoE, PLE NVMe streaming | `@atlas/qwen3.8-flash-next-nvfp4` | 36.7 |
|
| 77 |
+
| Nemotron-3.5-Lightning + DSpark | `nvidia` checkpoint via recipe | 30B / 3B | Mamba-2 + attention + MoE + DSpark drafter | `@atlas/nemotron-3.5-lightning-30b-a3b-nvfp4-dspark` | sub-10ms/token |
|
| 78 |
+
| Qwen3.5-27B | `Kbenkhaled/Qwen3.5-27B-NVFP4` | 27B dense | Hybrid SSM + attention, dense FFN, MRoPE | | 13 |
|
| 79 |
+
| Qwen3.5-35B-A3B | `Sehyo/Qwen3.5-35B-A3B-NVFP4` | 35B / 3B | GDN + attention + MoE, MTP | | **131** (MTP K=2) |
|
| 80 |
+
| Qwen3.5-122B-A10B | `Sehyo/Qwen3.5-122B-A10B-NVFP4` | 122B / 10B | GDN + attention + MoE, MTP | | 46 (EP=2) |
|
| 81 |
+
| Qwen3.6-35B-A3B | `Qwen/Qwen3.6-35B-A3B-FP8` | 35B / 3B | GDN + attention + MoE, MRoPE, vision | `@atlas/qwen3.6-35b-a3b-fp8-mtp` | |
|
| 82 |
+
| Holo-3.1-35B-A3B | `Hcompany/Holo-3.1-35B-A3B-NVFP4` | 35B / 3B | GDN + attention + 256-expert MoE, Qwen3-VL vision | | |
|
| 83 |
+
| Holo-3.1-0.8B | `Hcompany/Holo-3.1-0.8B` | 0.8B dense | GDN + attention + dense FFN, Qwen3-VL vision | | |
|
| 84 |
+
| Ornith-1.0-9B | `deepreinforce-ai/Ornith-1.0-9B` | 9B dense | GDN + attention + dense FFN, Qwen3-VL vision, MRoPE | | |
|
| 85 |
+
| Qwen3-Next-80B-A3B | `nvidia/Qwen3-Next-80B-A3B-Instruct-NVFP4` | 80B / 3B | SSM + attention + MoE | | 74 |
|
| 86 |
+
| Qwen3-VL-30B-A3B | `ig1/Qwen3-VL-30B-A3B-Instruct-NVFP4` | 30B / 3B | Vision + attention + MoE | | 97 |
|
| 87 |
+
| Gemma-4-26B-A4B | `bg-digitalservices/Gemma-4-26B-A4B-it-NVFP4A16` | 26B / 4B | Attention + MoE, GeGLU | `@atlas/gemma-4-26b-a4b-nvfp4` | 67 |
|
| 88 |
+
| Gemma-4-31B | `nvidia/Gemma-4-31B-IT-NVFP4` | 31B dense | Attention (sliding + full), GeGLU | | 9 |
|
| 89 |
+
| Mistral-Small-4-119B | `mistralai/Mistral-Small-4-119B-2603-NVFP4` | 119B / 6.5B | Attention + MoE | | 33 |
|
| 90 |
+
| MiniMax-M2.7 | `lukealonso/MiniMax-M2.7-NVFP4` | 229B / ~10B | Attention + 256-expert MoE + MTP | | |
|
| 91 |
+
| Nemotron-3-Nano-30B-A3B | `nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4` | 30B / 3B | Mamba-2 + attention + MoE | `@atlas/nemotron-3-nano-30b-a3b-nvfp4` | 88 |
|
| 92 |
+
| Nemotron-3-Super-120B-A12B | `nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4` | 120B / 12B | Mamba-2 + attention + MoE | `@atlas/nemotron-3-super-120b-a12b-nvfp4` | 24 |
|
| 93 |
+
| DeepSeek-V4-Flash | via recipe | | Attention + MoE | `@atlas/deepseek-v4-flash-nvfp4-ep2` | EP=2, 2 Sparks |
|
| 94 |
|
| 95 |
+
Flagship and Flash-Next also ship `-latency` and `-throughput` recipe variants for single-stream latency and 1 to 128 stream concurrency. On Qwen3.5-35B-A3B with MTP speculative decoding, Atlas decodes faster than NVIDIA's own vLLM build on the same hardware, on numbers we can hand you the script for. If you reproduce a faster vLLM number, file an issue. We would rather be measured than congratulated. The kernel-by-kernel comparison against PyTorch eager lives in the [benchmarks chapter](book/src/operations/benchmarks.md).
|
| 96 |
|
| 97 |
+
And this is only the beginning. The plug-and-play design has already carried AMD into the garden, Apple Silicon and Intel are next, and next quarter's models will slot in the same way the Qwens did this quarter.
|
| 98 |
+
|
| 99 |
+
> **New to Atlas on a Spark?** The [**GB10 Deployment & Compatibility Guide**](docs/GB10_DEPLOYMENT_GUIDE.md) is the one page to read first. It covers which model and quant fit your box, what to do when it OOMs, the known gotchas, and what verified means, then hands you the exact recipe.
|
| 100 |
+
|
| 101 |
+
<a id="performance"></a>
|
| 102 |
+
|
| 103 |
+
<a id="quick-start"></a>
|
| 104 |
+
|
| 105 |
+
## 🚀 Quick Start
|
| 106 |
+
|
| 107 |
+
### 1. NVIDIA GB10, flagship recipe `@atlas/qwen3.8-27b-nvfp4`
|
| 108 |
|
| 109 |
```bash
|
| 110 |
+
# Step 1, install sparkrun
|
| 111 |
+
pip install sparkrun # or uvx sparkrun setup install
|
| 112 |
|
| 113 |
+
# Step 2, download weights (or let sparkrun fetch automatically)
|
| 114 |
+
huggingface-cli download unsloth/Qwen3.8-27B-NVFP4 \
|
| 115 |
+
--local-dir ~/.cache/huggingface/hub/models--unsloth--Qwen3.8-27B-NVFP4
|
|
|
|
|
|
|
| 116 |
|
| 117 |
+
# Step 3, launch the flagship (port 8888)
|
| 118 |
+
sparkrun run @atlas/qwen3.8-27b-nvfp4 --hosts localhost
|
| 119 |
+
```
|
| 120 |
|
| 121 |
+
Prefer the one-line quickstart script?
|
| 122 |
```bash
|
| 123 |
+
curl -fsSL https://atlasinference.dev/quickstart.sh | sh
|
|
|
|
|
|
|
| 124 |
```
|
| 125 |
|
| 126 |
+
The other flagships work the same way. The recipe is the command.
|
| 127 |
|
| 128 |
+
```bash
|
| 129 |
+
# Nemotron 3.5 Lightning 30B + DSpark drafter, sub-10ms decode
|
| 130 |
+
export ATLAS_DFLASH_OPTION_B=1 ATLAS_NO_TOOL_INJECT=1
|
| 131 |
+
sparkrun run @atlas/nemotron-3.5-lightning-30b-a3b-nvfp4-dspark --hosts localhost
|
| 132 |
+
|
| 133 |
+
# Qwen 3.8 Flash-Next ~180B, streams the 47.7 GB PLE table off NVMe (~90 GB resident)
|
| 134 |
+
sparkrun run @atlas/qwen3.8-flash-next-nvfp4 --hosts localhost
|
| 135 |
+
```
|
| 136 |
|
| 137 |
+
### 2. AMD Strix Halo (gfx1151) on Linux and Windows
|
| 138 |
|
| 139 |
+
Strix Halo is a unified-memory APU, so the validated path is the native binary built against ROCm with no container. Both legs merged into main via [PR #8](https://github.com/Atlas-Inf/atlas/pull/8) (Linux) and [PR #9](https://github.com/Atlas-Inf/atlas/pull/9) (Windows).
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 140 |
|
| 141 |
+
**Linux** (Ubuntu 24.04, ROCm 6.2+)
|
| 142 |
|
| 143 |
+
```bash
|
| 144 |
+
git clone https://github.com/Atlas-Inf/atlas.git
|
| 145 |
+
cd atlas
|
| 146 |
+
./build-amd.sh # strix-hip backend, all targets, needs ROCm + cargo
|
| 147 |
+
./serve-amd.sh # nvidia/Qwen3.8-27B-NVFP4, validated config
|
| 148 |
+
./serve-amd.sh unsloth/Qwen3.8-27B-NVFP4 # or the preservation checkpoint
|
| 149 |
+
```
|
| 150 |
|
| 151 |
+
Defaults are the validated config. K=4 MTP speculative, the W4A8 DP4A decode arm, BF16 KV. `NUM_DRAFTS=0` disables speculation, `LM_HEAD=bf16` switches for the unsloth checkpoint's per-row-FP8 lm_head, and `ATLAS_W4A16_DP4A=0` opts out of DP4A. The `ATLAS_FP8_DEQUANT_*` and `ATLAS_GDN_BF16_WEIGHTS` exports are baked in and required for unsloth.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 152 |
|
| 153 |
+
Measured on a Ryzen AI Max+ 395 with Radeon 8060S under ROCm 7.13 at ~60 GB GTT. **28.3 to 28.6 tok/s** K=4 decode, 13.3 to 13.6 tok/s at 30k context, and **83.02 / 80.41** on the 995-row bfcl-subset golden draw, matching the NVIDIA shipped reference (83.22 / 79.02). Details in [`BENCH.toml`](kernels/strix-hip/qwen3.8-27b/BENCH.toml) and [`amd-strix-halo-scale.md`](docs/porting/amd-strix-halo-scale.md).
|
| 154 |
|
| 155 |
+
**Windows** (native HIP on Windows 11, no WSL, no container. `spark.exe` is MSVC-built and reaches the GPU through a CUDA to HIP shim over ROCm 10 / TheRock)
|
| 156 |
|
| 157 |
+
```powershell
|
| 158 |
+
git clone https://github.com/Atlas-Inf/atlas.git
|
| 159 |
+
cd atlas
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 160 |
|
| 161 |
+
hf download nvidia/Qwen3.8-27B-NVFP4 `
|
| 162 |
+
--local-dir "$env:USERPROFILE\models\nvidia-Qwen3.8-27B-NVFP4"
|
| 163 |
|
| 164 |
+
# Build from PowerShell, not Git Bash (bash puts coreutils link.exe ahead
|
| 165 |
+
# of MSVC's). Needs MSVC Desktop C++, the ROCm SDK (HIP_PATH), and cargo.
|
| 166 |
+
.\build-amd.ps1
|
| 167 |
|
| 168 |
+
# Serve, the fingerprinted fp8d recipe. Native FP8 GDN, K=4 MTP, BF16 KV
|
| 169 |
+
# and lm_head. Preflight hard-fails on non-gfx1151 GPUs.
|
| 170 |
+
.\serve-amd.ps1
|
| 171 |
+
```
|
| 172 |
|
| 173 |
+
Prebuilt instead? Unzip `spark-windows-x86_64-amd-hip` keeping every DLL beside `spark.exe`, set `$env:ATLAS_BIN` to it, and run `.\serve-amd.ps1`. That skips the toolchain check and the build. The script runs a detached smoke probe (`first_run_smoke.log`), and the 64K-context record config is [`win_serve_qwen38_nvfp4.ps1`](scripts/strix-windows/win_serve_qwen38_nvfp4.ps1).
|
|
|
|
| 174 |
|
| 175 |
+
Measured on a Framework Desktop (Ryzen AI Max+ 395, Radeon 8060S, Windows 11, ROCm 10.0.0 TheRock, Adrenalin 32.0.31041.1004). **83.32 / 78.70** on the bfcl-subset golden draw with zero faults in 3.6 hours, and **17.25 tok/s** decode with MTP engaged at p1 0.834. Provenance in [`BENCH.toml`](kernels/strix-hip/qwen3.8-27b/BENCH.toml) and [`STRIX_WINDOWS_HIP.md`](docs/porting/STRIX_WINDOWS_HIP.md).
|
|
|
|
| 176 |
|
| 177 |
+
### Hitting the Endpoint
|
| 178 |
|
| 179 |
+
Atlas speaks OpenAI, Anthropic, and Responses APIs on the same port. `curl`, the OpenAI SDK, Open WebUI, opencode, Cline, and Claude Code all work when pointed at the served port. `sparkrun` recipes default to 8888 and the AMD scripts default to 8081.
|
| 180 |
|
| 181 |
+
```bash
|
| 182 |
+
curl http://localhost:8888/v1/chat/completions \
|
| 183 |
+
-H "Content-Type: application/json" \
|
| 184 |
+
-d '{
|
| 185 |
+
"model": "atlas",
|
| 186 |
+
"messages": [{"role": "user", "content": "Explain quantum computing in three sentences."}],
|
| 187 |
+
"max_tokens": 256
|
| 188 |
+
}'
|
| 189 |
+
```
|
| 190 |
|
| 191 |
+
---
|
| 192 |
|
| 193 |
+
## 🤝 Community & Support
|
|
|
|
| 194 |
|
| 195 |
+
- **Website**: [atlasinference.dev](https://atlasinference.dev)
|
| 196 |
+
- **Discord**: [Join our Discord](https://discord.com/invite/6vDbKaKrKD) — Active daily development, live kernel tuning, and model requests.
|
| 197 |
+
- **Recipes Repository**: [Atlas-Inf/sparkrun-recipes](https://github.com/Atlas-Inf/sparkrun-recipes)
|
| 198 |
+
- **Deployment Guide**: [GB10 Deployment Guide](docs/GB10_DEPLOYMENT_GUIDE.md)
|
| 199 |
|
| 200 |
---
|
| 201 |
|
| 202 |
+
## ⚖️ Dual License
|
| 203 |
+
|
| 204 |
+
- **Community Edition**: Licensed under **AGPLv3**. Free and open for personal use, research, and non-commercial local deployments.
|
| 205 |
+
- **Enterprise Edition**: Commercial licensing for proprietary applications, SaaS hosting without AGPLv3 copyleft obligations, dedicated support, and custom hardware/kernel porting. Contact `debaterishaqui@gmail.com`.
|
| 206 |
+
|
| 207 |
+
<sub><b>Continuity notice.</b> Atlas is continuing. This repository, the <a href="https://github.com/Atlas-Inf">Atlas-Inf</a> GitHub organization, and <a href="https://atlasinference.dev">atlasinference.dev</a> are the replacement official Atlas channels. The existing website and GitHub repository remain disputed Atlas assets that have not been relinquished.</sub>
|