Instructions to use PoSTMEDIA/Rosetta-7B-Instruct-NVFP4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use PoSTMEDIA/Rosetta-7B-Instruct-NVFP4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="PoSTMEDIA/Rosetta-7B-Instruct-NVFP4", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("PoSTMEDIA/Rosetta-7B-Instruct-NVFP4", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use PoSTMEDIA/Rosetta-7B-Instruct-NVFP4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "PoSTMEDIA/Rosetta-7B-Instruct-NVFP4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "PoSTMEDIA/Rosetta-7B-Instruct-NVFP4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/PoSTMEDIA/Rosetta-7B-Instruct-NVFP4
- SGLang
How to use PoSTMEDIA/Rosetta-7B-Instruct-NVFP4 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "PoSTMEDIA/Rosetta-7B-Instruct-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "PoSTMEDIA/Rosetta-7B-Instruct-NVFP4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "PoSTMEDIA/Rosetta-7B-Instruct-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "PoSTMEDIA/Rosetta-7B-Instruct-NVFP4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use PoSTMEDIA/Rosetta-7B-Instruct-NVFP4 with Docker Model Runner:
docker model run hf.co/PoSTMEDIA/Rosetta-7B-Instruct-NVFP4
Introduction
Rosetta-7B-Instruct-NVFP4 is the official NVFP4 (4-bit floating point) quantization of Rosetta-7B-Instruct, PoSTMEDIA's bilingual (Korean-English) instruction-tuned model. It was produced with NVIDIA TensorRT Model Optimizer using the official NVFP4 post-training-quantization recipe, and targets NVIDIA Blackwell GPUs — DGX Spark, GeForce RTX 50 series, and B200/GB200-class datacenter parts — where vLLM selects native NVFP4 GEMM kernels automatically.
The checkpoint shrinks from 14.5 GB (BF16) to 6.3 GB (~2.3× smaller), and because single-stream decoding is dominated by weight memory bandwidth, decode throughput improves substantially on bandwidth-bound devices such as DGX Spark. The same checkpoint also loads on pre-Blackwell GPUs (Hopper, Ada, Ampere) through vLLM's weight-only Marlin fallback — with the identical memory savings.
| Model | Download | Note |
|---|---|---|
| Rosetta-7B-Base | HuggingFace | Foundation model (completion-style) |
| Rosetta-7B-Instruct | HuggingFace | Instruction following / chat (BF16) |
| Rosetta-7B-Instruct-NVFP4 | HuggingFace | NVFP4 quantization (this model) |
| Rosetta-7B-Think | HuggingFace | Explicit reasoning (<think>, BF16) |
| Rosetta-7B-Think-NVFP4 | HuggingFace | NVFP4 quantization of Think |
Quantization Details
| Method | NVIDIA TensorRT Model Optimizer (nvidia-modelopt 0.46.1), official NVFP4 PTQ recipe |
| Weight format | NVFP4 — FP4 (E2M1) elements, 16-element blocks, FP8 (E4M3) per-block scales + FP32 per-tensor scale |
| Activations | NVFP4 (W4A4) with calibrated static input scales |
| KV cache | FP8 (E4M3) with calibrated static scales (activated when serving with FP8 KV cache) |
| Kept in BF16 | embeddings, lm_head, normalization layers |
| Calibration | 1,024 bilingual samples — Korean instruction conversations from PoSTMEDIA's in-house synthetic data assets (8 domains, rendered through the model's chat template) + English news articles |
| Checkpoint size | 6.3 GB (vs. 14.5 GB BF16) |
| Format | ModelOpt unified HuggingFace checkpoint (quantization_config + hf_quant_config.json) |
In internal side-by-side evaluations against the BF16 model under an identical protocol, output quality is broadly preserved across Korean and English instruction following, knowledge, and math. As with any 4-bit quantization, minor differences can surface on knowledge-recall edge cases; for maximum-accuracy use cases, prefer the BF16 model.
Quickstart
vLLM
Use the PoSTMEDIA vLLM distribution with native Rosetta support — no trust_remote_code required:
VLLM_USE_PRECOMPILED=1 pip install git+https://github.com/PoSTMEDIA-AI/vllm@rosetta-v0.26.0
vllm serve PoSTMEDIA/Rosetta-7B-Instruct-NVFP4
vLLM detects the ModelOpt NVFP4 checkpoint automatically:
- Blackwell (SM 100/120/121 — B200, RTX 50, DGX Spark) — native NVFP4 GEMM kernels (CUTLASS / FlashInfer / Marlin, auto-selected)
- Hopper / Ada / Ampere (SM ≥ 80) — weight-only Marlin fallback: same 6.3 GB footprint, BF16 arithmetic
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
resp = client.chat.completions.create(
model="PoSTMEDIA/Rosetta-7B-Instruct-NVFP4",
messages=[{"role": "user", "content": "부산 여행 1박 2일 코스를 짜줘."}],
temperature=0.7,
)
print(resp.choices[0].message.content)
vLLM v0.26 or later is required. Recommended sampling:
temperature 0.7, top_p 0.9(or greedy for deterministic tasks). This NVFP4 checkpoint is designed for serving stacks that understand the ModelOpt unified format (vLLM; TensorRT-LLM- and SGLang-compatible layout). For plaintransformersinference, use the BF16 model.
DGX Spark
On DGX Spark (GB10, 128 GB unified memory), the 6.3 GB NVFP4 checkpoint leaves nearly all memory free for KV cache and other workloads, and decode speed improves markedly over BF16 because decoding on Spark is bound by weight-streaming bandwidth. Install the PoSTMEDIA vLLM distribution in a CUDA 13 environment and serve with the same command as above.
Model Summary
Identical to Rosetta-7B-Instruct: Rosetta dense decoder-only Transformer (RosettaForCausalLM), 7B parameters, 32 layers, interleaved sliding-window (4,096) + global attention (3:1) with QK-normalization, 65,536-token context, 161,425-token Korean-extended vocabulary. See the base model card for training details and full benchmark results.
Limitations
- Inherits the limitations of the BF16 base model (factuality, bias, Korean/English focus, 32K alignment window).
- 4-bit quantization can introduce small deviations from BF16 outputs; verify quality on your workload before production use.
- Native FP4 acceleration requires NVIDIA Blackwell GPUs and a serving stack with NVFP4 kernels (vLLM ≥ 0.26 recommended).
License
Apache License 2.0 — see LICENSE. If you build something with Rosetta, we'd appreciate a "Built with Rosetta" attribution.
Citation
@misc{rosetta2026,
title = {Rosetta-7B: A Bilingual Korean-English Language Model Family},
author = {{PoSTMEDIA AI Lab}},
year = {2026},
url = {https://huggingface.co/collections/PoSTMEDIA/rosetta-6a9db30fd1b4585b0c1845e9}
}
Contact
Questions and feedback — please open a discussion on the model page.
- Downloads last month
- 249
Model tree for PoSTMEDIA/Rosetta-7B-Instruct-NVFP4
Base model
PoSTMEDIA/Rosetta-7B-Base