File size: 1,294 Bytes
85f23bf
 
 
 
 
 
c30a200
170174b
85f23bf
 
 
 
170174b
270fa01
 
170174b
 
 
 
85f23bf
 
1478e94
 
 
 
85f23bf
1478e94
85f23bf
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
---
language: en
library_name: mlx
pipeline_tag: text-generation
tags:
- mlx
license: mit
base_model: meituan-longcat/LongCat-2.0
---

# kernelpool/LongCat-2.0-3bit

3-bit quantization of [meituan-longcat/LongCat-2.0](https://huggingface.co/meituan-longcat/LongCat-2.0),
converted with [mlx-lm](https://github.com/ml-explore/mlx-lm).

> **Revision note**: originally converted from the FP8 release
> ([meituan-longcat/LongCat-2.0-FP8](https://huggingface.co/meituan-longcat/LongCat-2.0-FP8)),
> the current revision is re-converted from the bf16 master checkpoint.

## Use with mlx

This model requires LongCat-2.0 support from [mlx-lm PR #1464](https://github.com/ml-explore/mlx-lm/pull/1464),
which has not yet been merged. Until it is included in an mlx-lm release, install
mlx-lm from the PR branch:

```bash
pip install git+https://github.com/ml-explore/mlx-lm.git@refs/pull/1464/head
```

```python
from mlx_lm import load, generate

model, tokenizer = load("kernelpool/LongCat-2.0-3bit")

prompt = "hello"

if tokenizer.chat_template is not None:
    messages = [{"role": "user", "content": prompt}]
    prompt = tokenizer.apply_chat_template(
        messages, add_generation_prompt=True, return_dict=False,
    )

response = generate(model, tokenizer, prompt=prompt, verbose=True)
```