Instructions to use AtomicChat/Laguna-S-2.1-NVFP4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use AtomicChat/Laguna-S-2.1-NVFP4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="AtomicChat/Laguna-S-2.1-NVFP4", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("AtomicChat/Laguna-S-2.1-NVFP4", trust_remote_code=True) model = AutoModelForCausalLM.from_pretrained("AtomicChat/Laguna-S-2.1-NVFP4", trust_remote_code=True, device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use AtomicChat/Laguna-S-2.1-NVFP4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "AtomicChat/Laguna-S-2.1-NVFP4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AtomicChat/Laguna-S-2.1-NVFP4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/AtomicChat/Laguna-S-2.1-NVFP4
- SGLang
How to use AtomicChat/Laguna-S-2.1-NVFP4 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "AtomicChat/Laguna-S-2.1-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AtomicChat/Laguna-S-2.1-NVFP4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "AtomicChat/Laguna-S-2.1-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AtomicChat/Laguna-S-2.1-NVFP4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use AtomicChat/Laguna-S-2.1-NVFP4 with Docker Model Runner:
docker model run hf.co/AtomicChat/Laguna-S-2.1-NVFP4
Laguna S 2.1, self-quantized to NVFP4 by Atomic Chat. Built straight from Poolside's original weights with a per-tensor importance matrix, so this is not a repack of somebody else's files. Runs fully offline.
Highlights
- 117.6B parameters: the weights this repo quantizes.
- Context length: 1,048,576 tokens (1M), as published by Poolside.
- 48 layers: Mixture-of-Experts, hybrid sliding-window (512) and global attention.
- Full imatrix ladder: every quant is calibrated with an importance matrix.
- Mixed SWA and global attention layout: 48 layers in a 1:3 global-to-SWA ratio (12 global attention layers, 36 sliding-window layers, window 512), with softplus attention gating and per-layer-type rotary scales.
- Native reasoning support: interleaved thinking between tool calls, with per-request control via enable_thinking.
- Speculative decoding: a trained DFlash draft model is available for lower-latency serving.
These NVFP4s are self-quantized from the original weights, not a repack. The importance matrix keeps low-bit quants closer to the full-precision model.
Model Overview
| Property | Value |
|---|---|
| Base model | poolside/Laguna-S-2.1 |
| Parameters | 117.6B |
| Layers | 48 |
| Experts | 256 routed (top-10) |
| Sliding window | 512 tokens |
| Context length | 1,048,576 tokens (1M) |
| Vocabulary | 100,352 |
| Modalities | Text |
| Architecture | Mixture-of-Experts, 256 experts (top-10), hybrid sliding-window (512) and global attention, 48 attention heads over 8 KV heads, LagunaForCausalLM |
| This repo | NVFP4 weights |
Benchmarks
| Benchmark | Score |
|---|---|
| Laguna S 2.1 | 70.2% |
| Tencent Hy3 | 71.7% |
| Inkling | 63.8% |
| Nemotron 3 Ultra | 56.4% |
| DeepSeek-V4-Pro Max | 64.0% |
| Kimi K3 | 88.3% |
| Qwen 3.7 Max | 74.5% |
| Muse Spark 1.1 | 80% |
| Claude Fable 5 | 88% |
Scores are Poolside's published results for the base poolside/Laguna-S-2.1, not our own measurements. Quantization preserves the large majority of this; Q4_K_M and up stay close to full precision.
Get started
- Atomic Chat: search
AtomicChat/Laguna-S-2.1-NVFP4and hit Use this model. - vLLM:
vllm serve AtomicChat/Laguna-S-2.1-NVFP4 --max-model-len 8192
Best practices
| Parameter | Value |
|---|---|
| temperature | 1.0 |
| top_p | 1.0 |
| top_k | 20 |
| min_p | 0.0 |
Poolside's recommended sampling configuration for poolside/Laguna-S-2.1.
How these were made
- Download
poolside/Laguna-S-2.1(original weights). - Quantize to NVFP4 with
llm-compressorover a calibration corpus.
License
Original model by Poolside, released under the OpenMDW-1.1 license. Full terms: OpenMDW-1.1. Quantized by Atomic Chat.
- Downloads last month
- 196
Model tree for AtomicChat/Laguna-S-2.1-NVFP4
Base model
poolside/Laguna-S-2.1

