GGUF
conversational

THIS DOES INCLUDE MTP LAYER unlike most other quants.

RC2.

Example:

Test 1: Code me a HTML5 page of a rotating cube, use a 4D projection matrix to give it a cool look. .

image

Test 2: .

image

image

Note

Use Pi agent, it is the best harness and allows compaction.

Also this DOES include the MTP layer, invoke it with: --spec-type draft-mtp

Qwen3.8-27B โ€” Mixed Q2

A selectively quantized Qwen3.8-27B GGUF using a mixed-precision Q2 strategy.

This quantization is designed around the observation that not all tensors contribute equally to model quality. Instead of forcing the entire model into the same low-bit representation, sensitive components are retained at higher precision while less sensitive weights are aggressively compressed.

Overview

Property Value
Base model Qwen3.8-27B
Parameters ~27B
Format GGUF
Quantization Mixed Q2
Primary goal Maximum quality per GB
Runtime llama.cpp / compatible GGUF loaders

Why Mixed Q2?

Uniform Q2 quantization treats every tensor as though it has the same tolerance for information loss.

It doesn't.

Some weights are substantially more important to preserving:

  • reasoning quality
  • instruction following
  • long-context behavior
  • mathematical precision
  • attention routing
  • residual information
  • output quality

This release therefore uses selective precision allocation.

The bulk of the model is compressed aggressively, while particularly sensitive tensors are preserved using higher-precision representations.

The result is intended to retain substantially more of the original model's behavior than a naive all-Q2 conversion at a comparable storage budget.

Quantization Philosophy

The guiding principle is:

Spend bits where they matter. Save bits where they don't.

Rather than optimizing solely for an average bits-per-weight number, the quantizer considers the structural role of individual tensors.

This makes the resulting GGUF a mixed quantization, not simply a conventional Q2 file with a different name.

High-sensitivity components

Certain tensors are deliberately excluded from aggressive quantization when their information content or numerical sensitivity makes them disproportionately important.

Intended Use

This model is intended for users who want to run a 27B-class Qwen model locally at an extremely constrained memory footprint without accepting the full quality loss normally associated with uniform Q2 quantization.

It is particularly interesting for:

  • local inference
  • laptops and desktops with limited VRAM
  • CPU inference
  • hybrid CPU/GPU inference
  • experimentation with ultra-low-bit LLMs
  • reasoning and coding workloads

Performance

Actual performance depends heavily on:

  • CPU/GPU
  • llama.cpp build
  • context length
  • batch size
  • GPU offload
  • KV-cache precision
  • operating system
  • backend

This release prioritizes quality retention per unit of storage.

Quality

The purpose of this quantization is not merely to make the model smaller.

The goal is to preserve the behaviors that tend to disappear first when a large model is pushed toward extremely low bitrates.

Credits

Base model: Qwen3.8-27B

Quantization: HPC-Quantize

Format: GGUF

Disclaimer

This is an experimental low-bit quantization.

Ultra-low-bit inference necessarily involves information loss relative to the original model. The mixed-precision strategy is intended to reduce that loss by allocating additional precision selectively, but it does not reproduce the full-precision model.

Results may vary substantially depending on workload and inference configuration.

If you find interesting differences between this release and other Q2/Q3 quantizations, please report the workload and inference configuration along with your results.

Downloads last month
1,360
GGUF
Model size
27B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support