YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Quantized Meta-LLaMA-3.1-8B-Instruct vLLM Inference Project
π Project Overview
This project benchmarks the inference performance of:
- vLLM server with Quantized (W8A8) Meta-LLaMA-3.1-8B-Instruct model using LLM Compressor
- Hugging Face Transformers running full-precision Meta-LLaMA-3.1-8B-Instruct in float16.
We evaluate:
- Average latency (seconds)
- Throughput (tokens/sec)
- Output quality for identical prompts
π₯οΈ System Setup
- GPU: NVIDIA A100-SXM4-40GB
- CUDA Version: 12.4
- Driver Version: 550.54.15
- Environment: Google Colab (Virtual Environment:
llm_comp)
βοΈ Tools and Libraries
- vLLM (Optimized LLM serving engine)
- llmcompressor (One-shot quantization compression)
π Workflow Summary
- Quantized Meta-LLaMA-3.1-8B-Instruct model using LLM Compressor, applying SmoothQuant + GPTQ techniques with W8A8 quantization scheme.
- Uploaded the quantized model to Hugging Face Hub:
Roy2144/Meta-LLama. - Deployed the quantized model using vLLM OpenAI-compatible API server.
- Ran Benchmarks:
- vLLM server performance (Quantized)
- Hugging Face Transformers local inference (float16)
- Measured:
- Latency per request
- Throughput (tokens/sec)
- Response quality comparison
- Analyzed differences between quantized and float16 models.
π Benchmark Results
| Metric | vLLM + Quantized Meta-LLaMA W8A8 | HF Transformers + Meta-LLaMA float16 |
|---|---|---|
| Avg Latency | 0.62 seconds | 2.66 seconds |
| Avg Throughput | 163.99 tokens/sec | 30.46 tokens/sec |
π§ Key Findings
- vLLM + Quantization massively boosts inference speed (~4β5x faster than float16).
- Quantized Meta-LLaMA-3.1-8B-Instruct maintains excellent response quality despite INT8 compression.
- HF Transformers serve as a baseline for unoptimized float16 inference.
β¨ Sample Prompt for Comparison
Prompt: "Explain quantization in simple words for a beginner."
Both vLLM Quantized and HF float16 models generated detailed, coherent responses.
Quantized model responses were slightly more concise.
π Future Enhancements
- Stress testing with GuideLLM: simulate 100+ concurrent users.
- Deploy a chatbot frontend with Gradio connected to vLLM server.
- Experiment with even lighter quantization (e.g., W4A8, W6A8).
- Multi-GPU tensor parallelism scaling.
π·οΈ Repo Details
- Model Repo: Roy2144/Meta-LLama
- Quantization Tool: LLM Compressor
- Quantization Scheme: SmoothQuant + GPTQ (W8A8)
- Inference Engines: vLLM, Hugging Face Transformers
- Downloads last month
- 3
Inference Providers NEW
This model isn't deployed by any Inference Provider. π Ask for provider support