Bunch of good models
Sawyer Bowerman
soyrsoyr
AI & ML interests
cv, optimization
Recent Activity
updated a model 2 days ago
soyrsoyr/GLM-5.3-MXFP4-MTP published a model 2 days ago
soyrsoyr/GLM-5.3-MXFP4-MTP updated a model 5 days ago
soyrsoyr/jev-playground-rlcdOrganizations
Custom Models
Bunch of bad models
Llama-3.2-1B-Instruct GPTQ Quantized
GPTQ quantized across W4A16, W8A8, FP8, NVFP4 using llm-compressor.
-
soyrsoyr/Llama-3.2-1B-Instruct-W4A16-GPTQ
Text Generation • 1B • Updated • 12 -
soyrsoyr/Llama-3.2-1B-Instruct-W8A8-GPTQ
Text Generation • 1B • Updated • 16 -
soyrsoyr/Llama-3.2-1B-Instruct-FP8-GPTQ
Text Generation • 1B • Updated • 13 -
soyrsoyr/Llama-3.2-1B-Instruct-NVFP4-GPTQ
Text Generation • 0.8B • Updated • 83
Gemma 4 12B Quantized
test models
models I'm using for testing
-
soyrsoyr/gemma-4-unified-0.8B-tiny
Image-Text-to-Text • 0.8B • Updated • 107 -
soyrsoyr/Nemotron-3.5-Lightning-1.4B-A0.1B-MTP-FP8-Dynamic
1B • Updated • 53 -
soyrsoyr/Qwen3-VL-Reranker-0.1B-tiny
Image-Text-to-Text • 98.7M • Updated • 12 -
soyrsoyr/Nemotron-3.5-Lightning-1.4B-A0.1B-MTP-NVFP4
Text Generation • 1B • Updated • 404 • 1
Muse-Glimmer-30B
Quantized across W4A16, FP8, NVFP4 using llm-compressor.
DeepSeek-MoE-16B-Chat GPTQ Quantized
DeepSeek-MoE-16B-Chat quantized with GPTQ via llm-compressor: W8A8, W4A16, FP8, NVFP4.
-
soyrsoyr/deepseek-moe-16b-chat-W8A8-GPTQ
Text Generation • 16B • Updated • 36 -
soyrsoyr/deepseek-moe-16b-chat-W4A16-GPTQ
Text Generation • 3B • Updated • 76 -
soyrsoyr/deepseek-moe-16b-chat-FP8-GPTQ
Text Generation • 16B • Updated • 17 -
soyrsoyr/deepseek-moe-16b-chat-NVFP4-GPTQ
Text Generation • 9B • Updated • 29
tiny models
tiny models for development
Quantized Models
Bunch of good models
test models
models I'm using for testing
-
soyrsoyr/gemma-4-unified-0.8B-tiny
Image-Text-to-Text • 0.8B • Updated • 107 -
soyrsoyr/Nemotron-3.5-Lightning-1.4B-A0.1B-MTP-FP8-Dynamic
1B • Updated • 53 -
soyrsoyr/Qwen3-VL-Reranker-0.1B-tiny
Image-Text-to-Text • 98.7M • Updated • 12 -
soyrsoyr/Nemotron-3.5-Lightning-1.4B-A0.1B-MTP-NVFP4
Text Generation • 1B • Updated • 404 • 1
Custom Models
Bunch of bad models
Muse-Glimmer-30B
Quantized across W4A16, FP8, NVFP4 using llm-compressor.
Llama-3.2-1B-Instruct GPTQ Quantized
GPTQ quantized across W4A16, W8A8, FP8, NVFP4 using llm-compressor.
-
soyrsoyr/Llama-3.2-1B-Instruct-W4A16-GPTQ
Text Generation • 1B • Updated • 12 -
soyrsoyr/Llama-3.2-1B-Instruct-W8A8-GPTQ
Text Generation • 1B • Updated • 16 -
soyrsoyr/Llama-3.2-1B-Instruct-FP8-GPTQ
Text Generation • 1B • Updated • 13 -
soyrsoyr/Llama-3.2-1B-Instruct-NVFP4-GPTQ
Text Generation • 0.8B • Updated • 83
DeepSeek-MoE-16B-Chat GPTQ Quantized
DeepSeek-MoE-16B-Chat quantized with GPTQ via llm-compressor: W8A8, W4A16, FP8, NVFP4.
-
soyrsoyr/deepseek-moe-16b-chat-W8A8-GPTQ
Text Generation • 16B • Updated • 36 -
soyrsoyr/deepseek-moe-16b-chat-W4A16-GPTQ
Text Generation • 3B • Updated • 76 -
soyrsoyr/deepseek-moe-16b-chat-FP8-GPTQ
Text Generation • 16B • Updated • 17 -
soyrsoyr/deepseek-moe-16b-chat-NVFP4-GPTQ
Text Generation • 9B • Updated • 29
Gemma 4 12B Quantized
tiny models
tiny models for development