YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Quantized Meta-LLaMA-3.1-8B-Instruct vLLM Inference Project

πŸš€ Project Overview

This project benchmarks the inference performance of:

  • vLLM server with Quantized (W8A8) Meta-LLaMA-3.1-8B-Instruct model using LLM Compressor
  • Hugging Face Transformers running full-precision Meta-LLaMA-3.1-8B-Instruct in float16.

We evaluate:

  • Average latency (seconds)
  • Throughput (tokens/sec)
  • Output quality for identical prompts

πŸ–₯️ System Setup

  • GPU: NVIDIA A100-SXM4-40GB
  • CUDA Version: 12.4
  • Driver Version: 550.54.15
  • Environment: Google Colab (Virtual Environment: llm_comp)

βš™οΈ Tools and Libraries

  • vLLM (Optimized LLM serving engine)
  • llmcompressor (One-shot quantization compression)

πŸ›  Workflow Summary

  1. Quantized Meta-LLaMA-3.1-8B-Instruct model using LLM Compressor, applying SmoothQuant + GPTQ techniques with W8A8 quantization scheme.
  2. Uploaded the quantized model to Hugging Face Hub: Roy2144/Meta-LLama.
  3. Deployed the quantized model using vLLM OpenAI-compatible API server.
  4. Ran Benchmarks:
    • vLLM server performance (Quantized)
    • Hugging Face Transformers local inference (float16)
  5. Measured:
    • Latency per request
    • Throughput (tokens/sec)
    • Response quality comparison
  6. Analyzed differences between quantized and float16 models.

πŸ“ˆ Benchmark Results

Metric vLLM + Quantized Meta-LLaMA W8A8 HF Transformers + Meta-LLaMA float16
Avg Latency 0.62 seconds 2.66 seconds
Avg Throughput 163.99 tokens/sec 30.46 tokens/sec

🧠 Key Findings

  • vLLM + Quantization massively boosts inference speed (~4–5x faster than float16).
  • Quantized Meta-LLaMA-3.1-8B-Instruct maintains excellent response quality despite INT8 compression.
  • HF Transformers serve as a baseline for unoptimized float16 inference.

✨ Sample Prompt for Comparison

Prompt: "Explain quantization in simple words for a beginner."

Both vLLM Quantized and HF float16 models generated detailed, coherent responses.
Quantized model responses were slightly more concise.


πŸ“š Future Enhancements

  • Stress testing with GuideLLM: simulate 100+ concurrent users.
  • Deploy a chatbot frontend with Gradio connected to vLLM server.
  • Experiment with even lighter quantization (e.g., W4A8, W6A8).
  • Multi-GPU tensor parallelism scaling.

🏷️ Repo Details

  • Model Repo: Roy2144/Meta-LLama
  • Quantization Tool: LLM Compressor
  • Quantization Scheme: SmoothQuant + GPTQ (W8A8)
  • Inference Engines: vLLM, Hugging Face Transformers

Downloads last month
3
Safetensors
Model size
8B params
Tensor type
BF16
Β·
I8
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support