Instructions to use CompressedGemma/Qwen3.8-27B-Q2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use CompressedGemma/Qwen3.8-27B-Q2 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf CompressedGemma/Qwen3.8-27B-Q2 # Run inference directly in the terminal: llama cli -hf CompressedGemma/Qwen3.8-27B-Q2
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf CompressedGemma/Qwen3.8-27B-Q2 # Run inference directly in the terminal: llama cli -hf CompressedGemma/Qwen3.8-27B-Q2
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf CompressedGemma/Qwen3.8-27B-Q2 # Run inference directly in the terminal: ./llama-cli -hf CompressedGemma/Qwen3.8-27B-Q2
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf CompressedGemma/Qwen3.8-27B-Q2 # Run inference directly in the terminal: ./build/bin/llama-cli -hf CompressedGemma/Qwen3.8-27B-Q2
Use Docker
docker model run hf.co/CompressedGemma/Qwen3.8-27B-Q2
- LM Studio
- Jan
- Ollama
How to use CompressedGemma/Qwen3.8-27B-Q2 with Ollama:
ollama run hf.co/CompressedGemma/Qwen3.8-27B-Q2
- Unsloth Desktop
- Pi
How to use CompressedGemma/Qwen3.8-27B-Q2 with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf CompressedGemma/Qwen3.8-27B-Q2
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "CompressedGemma/Qwen3.8-27B-Q2" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use CompressedGemma/Qwen3.8-27B-Q2 with Docker Model Runner:
docker model run hf.co/CompressedGemma/Qwen3.8-27B-Q2
- Lemonade
How to use CompressedGemma/Qwen3.8-27B-Q2 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull CompressedGemma/Qwen3.8-27B-Q2
Run and chat with the model
lemonade run user.Qwen3.8-27B-Q2-{{QUANT_TAG}}List all available models
lemonade list
- Hermes Agent
How to use CompressedGemma/Qwen3.8-27B-Q2 with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf CompressedGemma/Qwen3.8-27B-Q2
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default CompressedGemma/Qwen3.8-27B-Q2
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use CompressedGemma/Qwen3.8-27B-Q2 with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf CompressedGemma/Qwen3.8-27B-Q2
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "CompressedGemma/Qwen3.8-27B-Q2" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
THIS DOES INCLUDE MTP LAYER unlike most other quants.
RC2.
Example:
Test 1: Code me a HTML5 page of a rotating cube, use a 4D projection matrix to give it a cool look. .
Test 2: .
Note
Use Pi agent, it is the best harness and allows compaction.
Also this DOES include the MTP layer, invoke it with: --spec-type draft-mtp
Qwen3.8-27B โ Mixed Q2
A selectively quantized Qwen3.8-27B GGUF using a mixed-precision Q2 strategy.
This quantization is designed around the observation that not all tensors contribute equally to model quality. Instead of forcing the entire model into the same low-bit representation, sensitive components are retained at higher precision while less sensitive weights are aggressively compressed.
Overview
| Property | Value |
|---|---|
| Base model | Qwen3.8-27B |
| Parameters | ~27B |
| Format | GGUF |
| Quantization | Mixed Q2 |
| Primary goal | Maximum quality per GB |
| Runtime | llama.cpp / compatible GGUF loaders |
Why Mixed Q2?
Uniform Q2 quantization treats every tensor as though it has the same tolerance for information loss.
It doesn't.
Some weights are substantially more important to preserving:
- reasoning quality
- instruction following
- long-context behavior
- mathematical precision
- attention routing
- residual information
- output quality
This release therefore uses selective precision allocation.
The bulk of the model is compressed aggressively, while particularly sensitive tensors are preserved using higher-precision representations.
The result is intended to retain substantially more of the original model's behavior than a naive all-Q2 conversion at a comparable storage budget.
Quantization Philosophy
The guiding principle is:
Spend bits where they matter. Save bits where they don't.
Rather than optimizing solely for an average bits-per-weight number, the quantizer considers the structural role of individual tensors.
This makes the resulting GGUF a mixed quantization, not simply a conventional Q2 file with a different name.
High-sensitivity components
Certain tensors are deliberately excluded from aggressive quantization when their information content or numerical sensitivity makes them disproportionately important.
Intended Use
This model is intended for users who want to run a 27B-class Qwen model locally at an extremely constrained memory footprint without accepting the full quality loss normally associated with uniform Q2 quantization.
It is particularly interesting for:
- local inference
- laptops and desktops with limited VRAM
- CPU inference
- hybrid CPU/GPU inference
- experimentation with ultra-low-bit LLMs
- reasoning and coding workloads
Performance
Actual performance depends heavily on:
- CPU/GPU
- llama.cpp build
- context length
- batch size
- GPU offload
- KV-cache precision
- operating system
- backend
This release prioritizes quality retention per unit of storage.
Quality
The purpose of this quantization is not merely to make the model smaller.
The goal is to preserve the behaviors that tend to disappear first when a large model is pushed toward extremely low bitrates.
Credits
Base model: Qwen3.8-27B
Quantization: HPC-Quantize
Format: GGUF
Disclaimer
This is an experimental low-bit quantization.
Ultra-low-bit inference necessarily involves information loss relative to the original model. The mixed-precision strategy is intended to reduce that loss by allocating additional precision selectively, but it does not reproduce the full-precision model.
Results may vary substantially depending on workload and inference configuration.
If you find interesting differences between this release and other Q2/Q3 quantizations, please report the workload and inference configuration along with your results.
- Downloads last month
- 1,360
We're not able to determine the quantization variants.


