Spaces:
Running
Running
Feature native packed Bonsai 2 with vision
Browse files
README.md
CHANGED
|
@@ -19,7 +19,7 @@ We study how learned systems represent, adapt, and generalize, with experiments
|
|
| 19 |
|
| 20 |
- **[LFM2.5 WebGPU Space · MONARCH](https://huggingface.co/spaces/inductiveML/monarch-webgpu)** — Run LFM2.5-230M in your browser with our WGSL kernels, generate text locally, and measure your GPU's tokens per second.
|
| 21 |
- **[Qwen3.6-35B-A3B · Evolved Mixed-Bit](https://huggingface.co/inductiveML/Qwen3.6-35B-A3B-evolved-mxbit)** — A 12.63 GB MLX quantization with per-module precision selected by evolutionary search. Runs on Apple Silicon with stock `mlx-lm`.
|
| 22 |
-
- **[Ternary Bonsai kernels · Ternel](https://github.com/inductiveML/ternel)** — Custom Metal kernels
|
| 23 |
|
| 24 |
## How we work
|
| 25 |
|
|
|
|
| 19 |
|
| 20 |
- **[LFM2.5 WebGPU Space · MONARCH](https://huggingface.co/spaces/inductiveML/monarch-webgpu)** — Run LFM2.5-230M in your browser with our WGSL kernels, generate text locally, and measure your GPU's tokens per second.
|
| 21 |
- **[Qwen3.6-35B-A3B · Evolved Mixed-Bit](https://huggingface.co/inductiveML/Qwen3.6-35B-A3B-evolved-mxbit)** — A 12.63 GB MLX quantization with per-module precision selected by evolutionary search. Runs on Apple Silicon with stock `mlx-lm`.
|
| 22 |
+
- **[Ternary Bonsai kernels · Ternel](https://github.com/inductiveML/ternel)** — Custom Metal kernels run Bonsai 27B and Bonsai 2 27B directly from losslessly packed 1.75-bit weights on Apple Silicon. Bonsai 2 preserves the original FP16 vision tower and uses 6.92 GB of loaded model memory, 19.5% less than the official MLX checkpoint. [Bonsai 2 with vision](https://huggingface.co/inductiveML/Ternary-Bonsai-2-27B-mlx-lossless-1.75bpw) · [Original Bonsai](https://huggingface.co/inductiveML/Ternary-Bonsai-27B-mlx-lossless-1.75bpw).
|
| 23 |
|
| 24 |
## How we work
|
| 25 |
|