YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
model_card = """
Mistral-7B-Instruct-INT8 (Quantized with llmcompressor)
This is a quantized version of mistralai/Mistral-7B-Instruct-v0.2, compressed using llmcompressor (a quantization toolkit developed by the vLLM team).
The goal of this quantization was to reduce memory and improve inference efficiency while preserving model accuracy. The result is a Hugging Face-compatible INT8 model with significantly lower memory usage and good performance on commodity GPUs.
Quantization Details or recipe
The quantization was done using the llmcompressor library. The following modifiers and configuration were used:
from llmcompressor.modifiers.quantization import GPTQModifier
from llmcompressor.modifiers.smoothquant import SmoothQuantModifier
from llmcompressor.transformers import oneshot
recipe = [
SmoothQuantModifier(smoothing_strength=0.8),
GPTQModifier(scheme="W8A8", targets="Linear", ignore=["lm_head"]),
]
oneshot(
model="mistralai/Mistral-7B-Instruct-v0.2",
....
....
)
How to use the model
from transformers import AutoTokenizer, AutoModelForCausalLM
model = AutoModelForCausalLM.from_pretrained("Roy2144/Mistral-7B-Instruct-INT8")
tokenizer = AutoTokenizer.from_pretrained("Roy2144/Mistral-7B-Instruct-INT8")
inputs = tokenizer("Why is quantization useful for LLMs?", return_tensors="pt").to("cuda")
outputs = model.generate(**inputs, max_new_tokens=64)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Benchmark Results
- Benchmarked on a single NVIDIA A100 40GB GPU using:
- Transformers backend (FP16 and INT8 variants). Note: The current implementation of
transformersdoes not include optimized kernels for direct INT8 computation. So transformers (INT8) is optional here. - vLLM engine with INT8 inference acceleration for improved performance.
- Transformers backend (FP16 and INT8 variants). Note: The current implementation of
| Metric | Transformers (INT8) | Transformers (FP16) | vLLM (INT8) |
|---|---|---|---|
| Avg Latency | 150,045 ms โ | 2,278 ms โ | 1,100 ms โก |
| Throughput | 0.51 tokens/sec โ | 33.72 tokens/sec โ | 76.47 tokens/sec โก |
| Peak VRAM | 20.5 GB โ | 29.0 GB โ | 20.5 GB โ |
Benchmarking Process Overview
To evaluate the impact of INT8 quantization and compare different inference backends, we conducted controlled benchmarks using the following setup:
- Model:
mistralai/Mistral-7B-Instruct-v0.2(base) and its quantized variant - Device: Single NVIDIA A100 GPU (40GB)
- Backends Evaluated:
- Transformers (PyTorch) for both FP16 and INT8
- vLLM runtime for INT8 acceleration
- Prompt Set: 5 diverse user-like queries related to model usage, quantization, and vLLM
- Inference Config:
max_new_tokens=64,batch_size=1
We measured three key metrics:
- Latency (end-to-end response time)
- Throughput (tokens/sec)
- Peak VRAM Usage
The benchmarking script ran each prompt individually across the backends, capturing time and memory performance using PyTorch and CUDA APIs.
Scope & Future Work
This project focused on compressing and benchmarking the Mistral-7B-Instruct model using both:
- Hugging Face Transformers (FP16 and INT8 inference)
- vLLM runtime with OpenAI-compatible HTTP interface
Future improvements and scope of work include:
- Load testing with GuideLLM for more robust, concurrent benchmark evaluation
- End-to-end latency profiling with multi-turn prompts and larger batch sizes
- Accuracy comparison between original and quantized outputs (task-level metrics)
- Exploring hybrid quantization + sparsity for further efficiency
GuideLLM is a promising benchmarking tool that can simulate realistic OpenAI-style load and measure key inference metrics with concurrency and scale.
Credits
- Base model: mistralai/Mistral-7B-Instruct-v0.2
- Quantized with: llmcompressor
- Downloads last month
- 3
