Minicpm-2B-Zero-Refusal-GGUF

Model Description

This model has been fine-tuned on a curated 75,000-row uncensored dataset to drastically reduce and eliminate "mid-refusals". It is explicitly conditioned to fulfill the entirety of the user's request without moralizing or filtering.

Available Files & Quantizations

  • Minicpm-2B-Q4_K_M-Selective-Q8.gguf (๐Ÿš€ NEW & RECOMMENDED): A hybrid quantization marvel. The bulk of the model is compressed to 4-bit (Q4_K_M) using a highly calibrated 3,000-sample importance matrix (imatrix) to save RAM and boost speed. However, the 7 specific projection layers targeted by the uncensoring LoRA (attn_q, attn_k, attn_v, attn_output, ffn_gate, ffn_up, ffn_down) are surgically preserved at Q8_0 (8-bit). This completely prevents "Quantization-Induced Refusal", retaining the full uncensored alignment of the FP16 adapter while maintaining the 2.4GB size of a 4-bit model.
  • Minicpm-2B-Q4_K_M-imatrix.gguf: Standard 4-bit medium quantization calibrated using the Imatrix.
  • Minicpm-2B-Q4_K_M.gguf: Standard 4-bit medium quantization (round-to-nearest) without calibration.
  • Minicpm-2B-FP16.gguf: The unquantized 16-bit base model.
  • minicpm.imatrix: The raw importance matrix data.

Prompt Format

This model uses the standard ChatML format. It supports internal reasoning, meaning assistant responses should contain a <think>...</think> block prior to the final output.

Downloads last month
332
GGUF
Model size
3B params
Architecture
llama
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support