宋小猫
SongXiaoMao
AI & ML interests
None yet
Recent Activity
liked a model 8 days ago
halt95/Swift1.5-Qwen3.8-Flash-Next-W4A16-Merlin liked a model 8 days ago
halt95/Qwen3.8-27B-W4A16-MerlinOrganizations
None yet
Thank you so much for the quantization and for open-sourcing this!
❤️ 2
4
#3 opened 13 days ago
by
SongXiaoMao
VLLM startup error
4
#2 opened 29 days ago
by
SongXiaoMao
ways to squeeze more context?
4
#3 opened about 1 month ago
by
SageHusky
Thanks for the model and 3090GPU launch code
#2 opened about 1 month ago
by
SongXiaoMao
Is 180GB right?
10
#1 opened about 1 month ago
by
sisfabc
New Model based on Qwen 3.8
➕🔥 27
18
#23 opened 2 months ago
by
jaskfsafdjlk
VLLM启动报错
9
#2 opened 4 months ago
by
SongXiaoMao
MTP efficiency without official FP8 ha
👍 1
#1 opened 5 months ago
by
SongXiaoMao
MTP cannot be accelerated
#1 opened 5 months ago
by
SongXiaoMao
The official VLLM example starts normal inference error
#3 opened 5 months ago
by
SongXiaoMao
This model cannot use MTP
4
#2 opened 6 months ago
by
SongXiaoMao
Modify the configuration file
🔥 1
1
#1 opened 6 months ago
by
SongXiaoMao
FP8 work for base model or is 16-bit of 27B required?
17
#2 opened 6 months ago
by
unoid
Is there anyone who can tell me how to run this model with vllm correctly?
😔 3
7
#8 opened 6 months ago
by
beginor
Can the big guy quantify this model into MXFP4? Thank you!!
#3 opened 6 months ago
by
SongXiaoMao
How does the VLLM start this model?
2
#4 opened 7 months ago
by
SongXiaoMao
This quantization model is amzing
👍❤️ 2
5
#1 opened 7 months ago
by
hyunw55
Why is the file size of 4bit similar to FP8?
3
#2 opened 6 months ago
by
SongXiaoMao
Sensitive information is not a question
2
#3 opened 6 months ago
by
SongXiaoMao
VLLM 0.18.0 runs with an error
#2 opened 6 months ago
by
SongXiaoMao