|
Download README.md from godofecht/razor-weights: direct link, hf CLI and curl.
- Browser
- Download file 2.09 kB
-
https://huggingface.co/godofecht/razor-weights/resolve/main/README.md
- Command line
-
hf download hf://godofecht/razor-weights/README.md
-
curl -L -o README.md https://huggingface.co/godofecht/razor-weights/resolve/main/README.md
2.09 kB
| license: apache-2.0 | |
| base_model: | |
| - Qwen/Qwen3-0.6B | |
| - HuggingFaceTB/SmolLM2-135M-Instruct | |
| - google/gemma-3-1b-it | |
| tags: | |
| - razor | |
| - on-device | |
| - quantized | |
| - int8 | |
| # Razor weights | |
| int8 weights for the Razor inference engine, which runs language models on | |
| phone and laptop CPUs. These files are a flat binary that Razor memory-maps | |
| directly; they are not loadable by transformers or llama.cpp. | |
| | File | Base model | Size | Notes | | |
| | --- | --- | --- | --- | | |
| | `razor_gemma3_1b_int8.bin` | Gemma 3 1B IT | 1003 MB | Sliding-window attention, tied int8 embedding | | |
| | `razor_qwen3_06b_int8.bin` | Qwen3 0.6B | 598 MB | Tied int8 embedding | | |
| | `razor_smollm2_135m_int8.bin` | SmolLM2 135M Instruct | 249 MB | fp32 embedding, int8 tied head | | |
| All three store linear weights `[out, in]` and carry the layout marker below. | |
| Tokenizers are not included. Use the tokenizer.json from the base model. | |
| ## Accuracy | |
| Each file was checked against its fp32 HuggingFace reference on five prompts. | |
| The argmax token matched on all five for every model, and the logits over the | |
| reference's top 50 tokens correlate as follows. | |
| | Model | top-1 | top-50 overlap | logit correlation | | |
| | --- | --- | --- | --- | | |
| | Gemma 3 1B | 5/5 | 0.84 | 0.950 | | |
| | Qwen3 0.6B | 5/5 | 0.88 | 0.976 | | |
| Gemma was additionally checked at 354 and 1064 prompt tokens, either side of | |
| its 512-token attention window, since a window bug is invisible on the short | |
| prompts a parity suite usually uses. | |
| ## Layout marker | |
| Files exported after 2026-08-13 contain a `__layout_out_in` tensor, which tells | |
| Razor the linear weights are stored `[out, in]`. Files without it hold | |
| `[in, out]` and are transposed at load. The marker exists because the layout | |
| cannot be inferred from the dimensions: on Qwen3 the K and V projections are | |
| square, so both layouts look identical and a wrong guess produces fluent | |
| nonsense rather than an error. | |
| ## Licence | |
| Apache-2.0, inherited from the base models. Qwen3 is by Alibaba Cloud, | |
| SmolLM2 is by Hugging Face, and Gemma 3 is by Google and additionally subject | |
| to the Gemma Terms of Use. | |