ralph-qwen3-8b-binary

Binary-tier GGUF compression of Qwen/Qwen3-8B โ€” architecture and parameter count unchanged (8,190,735,360 weight-bearing parameters), weights re-stored at reduced bit-width. Submitted to the binary bit-tier on Bittensor subnet 40 (Ralph, model-compression).

License

Released under Apache License 2.0 โ€” see LICENSE. This is a derivative of Qwen3-8B, itself Apache-2.0; the original copyright/attribution notices and a description of what changed (including the third-party quantization tooling used) are preserved in NOTICE.

Model overview

Parent Qwen/Qwen3-8B (Apache-2.0)
Parameters 8,190,735,360
Architecture unchanged from parent
Format GGUF
Bit tier binary
File size 1,928,778,816 bytes (~1.80 GiB)
Measured bits/weight ~1.88 (file size / param count; embedding/output tensors kept at higher precision outside the compressed blocks)
Quantization method llama.cpp, imatrix-calibrated

No retraining or architectural modification โ€” this is a post-training weight re-storage of the parent's own weights.

Downloads last month
127
GGUF
Model size
8B params
Architecture
qwen3
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for tensor-tailor/ralph-qwen3-8b-binary

Finetuned
Qwen/Qwen3-8B
Quantized
(441)
this model