YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

model_card = """

Mistral-7B-Instruct-INT8 (Quantized with llmcompressor)

This is a quantized version of mistralai/Mistral-7B-Instruct-v0.2, compressed using llmcompressor (a quantization toolkit developed by the vLLM team).

The goal of this quantization was to reduce memory and improve inference efficiency while preserving model accuracy. The result is a Hugging Face-compatible INT8 model with significantly lower memory usage and good performance on commodity GPUs.

Quantization Details or recipe

The quantization was done using the llmcompressor library. The following modifiers and configuration were used:

from llmcompressor.modifiers.quantization import GPTQModifier
from llmcompressor.modifiers.smoothquant import SmoothQuantModifier
from llmcompressor.transformers import oneshot

recipe = [
    SmoothQuantModifier(smoothing_strength=0.8),
    GPTQModifier(scheme="W8A8", targets="Linear", ignore=["lm_head"]),
]

oneshot(
    model="mistralai/Mistral-7B-Instruct-v0.2",
....
....
)

How to use the model

from transformers import AutoTokenizer, AutoModelForCausalLM

model = AutoModelForCausalLM.from_pretrained("Roy2144/Mistral-7B-Instruct-INT8")
tokenizer = AutoTokenizer.from_pretrained("Roy2144/Mistral-7B-Instruct-INT8")

inputs = tokenizer("Why is quantization useful for LLMs?", return_tensors="pt").to("cuda")
outputs = model.generate(**inputs, max_new_tokens=64)

print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Benchmark Results

  • Benchmarked on a single NVIDIA A100 40GB GPU using:
    • Transformers backend (FP16 and INT8 variants). Note: The current implementation of transformers does not include optimized kernels for direct INT8 computation. So transformers (INT8) is optional here.
    • vLLM engine with INT8 inference acceleration for improved performance.
Metric Transformers (INT8) Transformers (FP16) vLLM (INT8)
Avg Latency 150,045 ms โŒ 2,278 ms โœ… 1,100 ms โšก
Throughput 0.51 tokens/sec โŒ 33.72 tokens/sec โœ… 76.47 tokens/sec โšก
Peak VRAM 20.5 GB โœ… 29.0 GB โŒ 20.5 GB โœ…

Benchmarking Process Overview

To evaluate the impact of INT8 quantization and compare different inference backends, we conducted controlled benchmarks using the following setup:

  • Model: mistralai/Mistral-7B-Instruct-v0.2 (base) and its quantized variant
  • Device: Single NVIDIA A100 GPU (40GB)
  • Backends Evaluated:
    • Transformers (PyTorch) for both FP16 and INT8
    • vLLM runtime for INT8 acceleration
  • Prompt Set: 5 diverse user-like queries related to model usage, quantization, and vLLM
  • Inference Config: max_new_tokens=64, batch_size=1

We measured three key metrics:

  • Latency (end-to-end response time)
  • Throughput (tokens/sec)
  • Peak VRAM Usage

The benchmarking script ran each prompt individually across the backends, capturing time and memory performance using PyTorch and CUDA APIs.

benchmark chart


Scope & Future Work

This project focused on compressing and benchmarking the Mistral-7B-Instruct model using both:

  • Hugging Face Transformers (FP16 and INT8 inference)
  • vLLM runtime with OpenAI-compatible HTTP interface

Future improvements and scope of work include:

  • Load testing with GuideLLM for more robust, concurrent benchmark evaluation
  • End-to-end latency profiling with multi-turn prompts and larger batch sizes
  • Accuracy comparison between original and quantized outputs (task-level metrics)
  • Exploring hybrid quantization + sparsity for further efficiency

GuideLLM is a promising benchmarking tool that can simulate realistic OpenAI-style load and measure key inference metrics with concurrency and scale.

Credits

  • Base model: mistralai/Mistral-7B-Instruct-v0.2
  • Quantized with: llmcompressor
Downloads last month
3
Safetensors
Model size
7B params
Tensor type
BF16
ยท
I8
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support