This is a decensored version of prism-ml/Bonsai-27B-unpacked, made using Heretic v1.4.0
Abliteration parameters
| Parameter | Value |
|---|---|
| direction_index | 29.58 |
| attn.o_proj.max_weight | 1.43 |
| attn.o_proj.max_weight_position | 38.82 |
| attn.o_proj.min_weight | 0.91 |
| attn.o_proj.min_weight_distance | 23.43 |
| mlp.down_proj.max_weight | 1.39 |
| mlp.down_proj.max_weight_position | 39.20 |
| mlp.down_proj.min_weight | 1.38 |
| mlp.down_proj.min_weight_distance | 31.17 |
Performance
| Metric | This model | Original model (prism-ml/Bonsai-27B-unpacked) |
|---|---|---|
| KL divergence | 0.0033 | 0 (by definition) |
| Refusals | 6/100 | 81/100 |
1-bit Bonsai 27B โ Unpacked FP16 Safetensors
FP16 safetensors (HuggingFace format) of the 1-bit Bonsai 27B model. This repo exists for users who want to run Bonsai with stock HuggingFace tooling or frameworks that don't yet support 1-bit weights natively. The 1-bit hybrid-attention kernels are currently in our forks of MLX, mlx-swift, and llama.cpp โ once they land upstream, this unpacked version will no longer be needed.
We strongly recommend using the native 1-bit models instead. The 1-bit format is where all the benefits of Bonsai come from โ a 14.2x memory reduction to 3.9 GB, interactive decoding on everyday laptops (44 tok/s on an M5 Pro), and the first 27B-class model that runs on a phone (11 tok/s on iPhone 17 Pro Max). This unpacked FP16 version is full-size (~54 GB) and does not provide any of those advantages.
For the optimized 1-bit release models (recommended):
- Bonsai-27B-mlx-1bit โ 1-bit MLX for Apple Silicon (Mac, iPhone, iPad)
- 1-bit GGUF (Q1_0_g128) for llama.cpp (CUDA, Metal, CPU)
For the quality-oriented variant:
- Ternary-Bonsai-27B-mlx-2bit โ Ternary Bonsai 27B (~7.2 GB, 95% of FP16) for laptops and GPUs
- Downloads last month
- 127