Instructions to use PoSTMEDIA/Rosetta-7B-Think-NVFP4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use PoSTMEDIA/Rosetta-7B-Think-NVFP4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="PoSTMEDIA/Rosetta-7B-Think-NVFP4", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("PoSTMEDIA/Rosetta-7B-Think-NVFP4", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use PoSTMEDIA/Rosetta-7B-Think-NVFP4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "PoSTMEDIA/Rosetta-7B-Think-NVFP4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "PoSTMEDIA/Rosetta-7B-Think-NVFP4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/PoSTMEDIA/Rosetta-7B-Think-NVFP4
- SGLang
How to use PoSTMEDIA/Rosetta-7B-Think-NVFP4 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "PoSTMEDIA/Rosetta-7B-Think-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "PoSTMEDIA/Rosetta-7B-Think-NVFP4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "PoSTMEDIA/Rosetta-7B-Think-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "PoSTMEDIA/Rosetta-7B-Think-NVFP4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use PoSTMEDIA/Rosetta-7B-Think-NVFP4 with Docker Model Runner:
docker model run hf.co/PoSTMEDIA/Rosetta-7B-Think-NVFP4
Introduction
Rosetta-7B-Think-NVFP4 is the official NVFP4 (4-bit floating point) quantization of Rosetta-7B-Think, PoSTMEDIA's bilingual (Korean-English) reasoning model that wraps an explicit reasoning trace in <think> ... </think> before its final answer. It was produced with NVIDIA TensorRT Model Optimizer using the official NVFP4 post-training-quantization recipe, and targets NVIDIA Blackwell GPUs — DGX Spark, GeForce RTX 50 series, and B200/GB200-class datacenter parts — where vLLM selects native NVFP4 GEMM kernels automatically.
Reasoning models generate long traces, so decode throughput matters doubly: the checkpoint shrinks from 14.5 GB (BF16) to 6.3 GB (~2.3× smaller), and on bandwidth-bound devices such as DGX Spark this translates directly into faster token generation over long <think> sequences. The same checkpoint also loads on pre-Blackwell GPUs (Hopper, Ada, Ampere) through vLLM's weight-only Marlin fallback.
| Model | Download | Note |
|---|---|---|
| Rosetta-7B-Base | HuggingFace | Foundation model (completion-style) |
| Rosetta-7B-Instruct | HuggingFace | Instruction following / chat (BF16) |
| Rosetta-7B-Instruct-NVFP4 | HuggingFace | NVFP4 quantization of Instruct |
| Rosetta-7B-Think | HuggingFace | Explicit reasoning (<think>, BF16) |
| Rosetta-7B-Think-NVFP4 | HuggingFace | NVFP4 quantization (this model) |
Quantization Details
| Method | NVIDIA TensorRT Model Optimizer (nvidia-modelopt 0.46.1), official NVFP4 PTQ recipe |
| Weight format | NVFP4 — FP4 (E2M1) elements, 16-element blocks, FP8 (E4M3) per-block scales + FP32 per-tensor scale |
| Activations | NVFP4 (W4A4) with calibrated static input scales |
| KV cache | FP8 (E4M3) with calibrated static scales (activated when serving with FP8 KV cache) |
| Kept in BF16 | embeddings, lm_head, normalization layers |
| Calibration | 1,024 bilingual samples with a 4,096-token calibration window — Korean long-form reasoning traces from PoSTMEDIA's in-house synthetic data assets (8 domains, chat-template rendered with <think> traces) + English mathematical-reasoning and news text |
| Checkpoint size | 6.3 GB (vs. 14.5 GB BF16) |
| Format | ModelOpt unified HuggingFace checkpoint (quantization_config + hf_quant_config.json) |
Calibration was tailored to the reasoning workload: a 4,096-token window (8× the standard recipe) so that full <think> spans are observed during activation calibration, with reasoning-trace data in both languages. In internal side-by-side evaluations against the BF16 model under an identical protocol, reasoning behavior — including reliable <think> termination on Korean inputs — is preserved. As with any 4-bit quantization, minor differences can surface on the hardest reasoning chains; for maximum-accuracy use cases, prefer the BF16 model.
Quickstart
vLLM
Use the PoSTMEDIA vLLM distribution — native Rosetta support and a built-in reasoning parser, no trust_remote_code required:
VLLM_USE_PRECOMPILED=1 pip install git+https://github.com/PoSTMEDIA-AI/vllm@rosetta-v0.26.0
vllm serve PoSTMEDIA/Rosetta-7B-Think-NVFP4 \
--reasoning-parser rosetta
vLLM detects the ModelOpt NVFP4 checkpoint automatically:
- Blackwell (SM 100/120/121 — B200, RTX 50, DGX Spark) — native NVFP4 GEMM kernels (CUTLASS / FlashInfer / Marlin, auto-selected)
- Hopper / Ada / Ampere (SM ≥ 80) — weight-only Marlin fallback: same 6.3 GB footprint, BF16 arithmetic
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
resp = client.chat.completions.create(
model="PoSTMEDIA/Rosetta-7B-Think-NVFP4",
messages=[{"role": "user", "content": "소수가 무한히 많음을 증명해줘."}],
temperature=0.6,
top_p=0.95,
)
print("REASONING:", resp.choices[0].message.reasoning)
print("ANSWER:", resp.choices[0].message.content)
vLLM v0.26 or later is required. Recommended sampling:
temperature 0.6, top_p 0.95. Allow a generousmax_tokens(≥ 4,096; 32,768 for competition math) so reasoning traces can complete. This NVFP4 checkpoint is designed for serving stacks that understand the ModelOpt unified format (vLLM; TensorRT-LLM- and SGLang-compatible layout). For plaintransformersinference, use the BF16 model.
DGX Spark
On DGX Spark (GB10, 128 GB unified memory), the 6.3 GB NVFP4 checkpoint leaves nearly all memory free for the long-context KV cache that reasoning workloads demand, and decode speed over long <think> traces improves markedly over BF16 because decoding on Spark is bound by weight-streaming bandwidth. Install the PoSTMEDIA vLLM distribution in a CUDA 13 environment and serve with the same command as above.
Model Summary
Identical to Rosetta-7B-Think: Rosetta dense decoder-only Transformer (RosettaForCausalLM), 7B parameters, 32 layers, interleaved sliding-window (4,096) + global attention (3:1) with QK-normalization, 65,536-token context, 161,425-token Korean-extended vocabulary, <think> ... </think> reasoning format. See the base model card for training details and full benchmark results.
Limitations
- Inherits the limitations of the BF16 base model (reasoning latency/token budget, factuality inside fluent traces, Korean/English focus, 32K alignment window).
- 4-bit quantization can introduce small deviations from BF16 outputs; verify quality on your workload — especially the hardest reasoning tasks — before production use.
- Native FP4 acceleration requires NVIDIA Blackwell GPUs and a serving stack with NVFP4 kernels (vLLM ≥ 0.26 recommended).
License
Apache License 2.0 — see LICENSE. If you build something with Rosetta, we'd appreciate a "Built with Rosetta" attribution.
Citation
@misc{rosetta2026,
title = {Rosetta-7B: A Bilingual Korean-English Language Model Family},
author = {{PoSTMEDIA AI Lab}},
year = {2026},
url = {https://huggingface.co/collections/PoSTMEDIA/rosetta-6a9db30fd1b4585b0c1845e9}
}
Contact
Questions and feedback — please open a discussion on the model page.
- Downloads last month
- 445