Sakura โ€” Fara1.5-9B (GGUF, three sizes)

Sakura logo

Fara1.5-9B is Microsoft's computer-use agent for web browsing (Qwen3.5-9B base, multimodal: it reads browser screenshots and emits structured tool calls). This repository holds three GGUF files of different size made from the BF16 weights with our measured mixed-codec method. Pick one on the Files page; the table below says what each file is. Independent community quantization, not an official Microsoft release. Part of the Sakura Mini line.

Quality at a glance

Same method, same texts, same machine for every file; lower KL divergence means closer to the BF16 original:

  • Sakura-Fara1.5-9B-3.91GiB.gguf vs bartowski IQ3_XXS (3.98 GiB): KL divergence to the BF16 original 36 % lower on average over all 3 texts (0.07 GiB smaller).
  • Sakura-Fara1.5-9B-4.84GiB.gguf vs bartowski IQ4_XS (4.88 GiB): KL divergence to the BF16 original 18 % lower on average over all 3 texts (0.04 GiB smaller).
  • Sakura-Fara1.5-9B-5.54GiB.gguf vs bartowski Q4_K_M (5.50 GiB): KL divergence to the BF16 original 40 % lower on average over all 3 texts (0.04 GiB larger).

Why: the bit budget per tensor is allocated from measured sensitivity (see How it was made), not from fixed rules. Full table below.

Which file should I take?

File For
Sakura-Fara1.5-9B-3.91GiB.gguf (3.90 GiB, 3.74 bpw) smallest: for tight memory budgets; expect visible quality loss
Sakura-Fara1.5-9B-4.84GiB.gguf (4.83 GiB, 4.64 bpw) middle: best balance of size and quality for most machines
Sakura-Fara1.5-9B-5.54GiB.gguf (5.53 GiB, 5.31 bpw) largest: closest to the BF16 original of the three

The three files

File Size bits/weight KLD en KLD dev KLD wiki same top token (avg) Tensor types by size
Sakura-Fara1.5-9B-3.91GiB.gguf 3.90 GiB 3.74 0.0611 0.0600 0.1071 89.8 % IQ3_S 41%, IQ4_XS 27%, Q3_K 21%, Q4_K 8%, Q5_K 2%
Sakura-Fara1.5-9B-4.84GiB.gguf 4.83 GiB 4.64 0.0158 0.0171 0.0282 94.6 % Q4_K 40%, IQ4_XS 33%, Q5_K 26%
Sakura-Fara1.5-9B-5.54GiB.gguf 5.53 GiB 5.31 0.0082 0.0085 0.0114 96.4 % Q5_K 83%, IQ4_XS 9%, Q4_K 6%

"Tensor types by size" lists the share of the file's bytes per ggml type (all are standard llama.cpp types, so the files run in stock llama.cpp). Quality columns: mean KL divergence of the quantized model's next-token distribution against the BF16 source on held-out text (lower is better) and how often the most likely token is unchanged (higher is better).

  • mmproj-Fara1.5-9B-f16.gguf: vision projector (unchanged from bartowski's F16 conversion).

Comparison at similar size

Same texts, same context, same chunks, measured by us with llama-perplexity (BF16 source = reference; PPL of the BF16 model: en: 1.901, dev: 2.044, wiki: 8.878). The reference quants are bartowski's imatrix quants of the same model, measured the same way. Sorted by size.

Quant Source Size KLD en KLD dev KLD wiki PPL en same top token (avg)
Sakura-Fara1.5-9B-3.91GiB.gguf this repository 3.91 GiB 0.0611 0.0600 0.1071 1.963 89.8 %
IQ3_XXS bartowski (imatrix) 3.98 GiB 0.0937 0.0981 0.1662 1.993 88.0 %
Sakura-Fara1.5-9B-4.84GiB.gguf this repository 4.84 GiB 0.0158 0.0171 0.0282 1.926 94.6 %
IQ4_XS bartowski (imatrix) 4.88 GiB 0.0194 0.0224 0.0325 1.913 94.6 %
Q4_K_M bartowski (imatrix) 5.50 GiB 0.0129 0.0131 0.0218 1.912 95.6 %
Sakura-Fara1.5-9B-5.54GiB.gguf this repository 5.54 GiB 0.0082 0.0085 0.0114 1.918 96.4 %

Method: 12 chunks of 512 tokens per text; texts: English held-out text, developer-text held-out set, English encyclopedia held-out text (not used in any calibration). Single measurement on one machine; small KLD differences at similar size are not a quality ranking.

How it was made

  1. BF16 GGUF of the original model and the importance matrix come from bartowski/Fara1.5-9B-GGUF; we used both unchanged (credit and thanks to bartowski).
  2. Every large weight matrix was quantized once per candidate type (Q2_K, IQ2_S, IQ3_XXS, IQ3_S, Q3_K, IQ4_XS, Q4_K, Q5_K, Q6_K, Q8_0) with llama-quantize and the importance matrix; the error of each choice was estimated per matrix (importance-weighted) and scaled by a sensitivity factor measured with real KL divergence: every group of tensors (for example the FFN down projections, the attention value projections, the first and last layers) was moved alone to a lower-bit type, and the KL divergence it caused was compared with its quantization error.
  3. An exact budget allocation (multiple-choice knapsack over bytes) picked one type per matrix for each target size; the final model was assembled from the stored tensors without re-quantizing. We do not publish per-tensor choices. Norms, small tensors and the multi-token-prediction layers stay at high precision.

Limits

  • Measured only with KL divergence and perplexity on short held-out texts and not with downstream benchmarks; do not read it as a task-quality claim.
  • Quantization always costs some quality; the smallest file costs the most.
  • Fara is a computer-use agent. Microsoft recommends running it only inside a sandbox with monitoring (MagenticLite) and describes safety behaviour (stopping and asking before payments, personal data and irreversible actions). Quantization can change behaviour; we did not test agent safety or screenshot grounding, only text perplexity/KLD.
  • The vision projector mmproj-Fara1.5-9B-f16.gguf is bartowski's conversion (unchanged, F16); it is needed for screenshots and is not quantized by us.

Credits


ไธญๆ–‡่ฏดๆ˜Ž ยท ๆจฑ่Šฑ (Simplified Chinese)

English above. ๆœฌ่Š‚ไธบไธŠๆ–‡็š„ไธญๆ–‡็ฟป่ฏ‘(Sakura = ๆจฑ่Šฑ yฤซnghuฤ);ๅฎŒๆ•ด็š„็‹ฌ็ซ‹ไธญๆ–‡็‰ˆ่ง README_zh.mdใ€‚

Sakura โ€” Fara1.5-9B(GGUF,ไธ‰็งๅคงๅฐ)

Sakura logo

Fara1.5-9B ๆ˜ฏ Microsoft ้ขๅ‘็ฝ‘้กตๆต่งˆ็š„่ฎก็ฎ—ๆœบไฝฟ็”จ(computer-use)ๆ™บ่ƒฝไฝ“(ๅŸบไบŽ Qwen3.5-9B,ๅคšๆจกๆ€:่ฏปๅ–ๆต่งˆๅ™จๆˆชๅ›พๅนถ่พ“ๅ‡บ็ป“ๆž„ๅŒ–็š„ๅทฅๅ…ท่ฐƒ็”จ)ใ€‚ๆœฌไป“ๅบ“ๅŒ…ๅซ ไธ‰ไธชไธๅŒๅคงๅฐ็š„ GGUF ๆ–‡ไปถ,็”ฑ BF16 ๆƒ้‡้€š่ฟ‡ๆˆ‘ไปฌๅฎžๆต‹ๅพ—ๅ‡บ็š„ๆททๅˆ็ผ–็ ๆ–นๆณ•ๅˆถไฝœใ€‚่ฏทๅœจ Files ้กต้ข้€‰ๆ‹ฉๅ…ถไธญไธ€ไธช;ไธ‹่กจ่ฏดๆ˜Žไบ†ๆฏไธชๆ–‡ไปถ็š„็”จ้€”ใ€‚ ่ฟ™ๆ˜ฏ็‹ฌ็ซ‹็š„็คพๅŒบ้‡ๅŒ–็‰ˆๆœฌ,ไธๆ˜ฏ Microsoft ็š„ๅฎ˜ๆ–นๅ‘ๅธƒใ€‚ๅฑžไบŽ Sakura Mini ็ณปๅˆ—ใ€‚

่ดจ้‡ๆฆ‚่งˆ

ๆฏไธชๆ–‡ไปถไฝฟ็”จ็›ธๅŒ็š„ๆ–นๆณ•ใ€็›ธๅŒ็š„ๆ–‡ๆœฌใ€ๅŒไธ€ๅฐๆœบๅ™จ;KL ๆ•ฃๅบฆ่ถŠไฝŽ,่ถŠๆŽฅ่ฟ‘ BF16 ๅŽŸๅง‹ๆจกๅž‹:

  • Sakura-Fara1.5-9B-3.91GiB.gguf ๅฏนๆฏ” bartowski IQ3_XXS (3.98 GiB):็›ธๅฏนไบŽ BF16 ๅŽŸๅง‹ๆจกๅž‹,ๅœจๅ…จ้ƒจ 3 ไธชๆ–‡ๆœฌไธŠๅนณๅ‡ KL ๆ•ฃๅบฆ ไฝŽ 36 %(ๅฐ 0.07 GiB)ใ€‚
  • Sakura-Fara1.5-9B-4.84GiB.gguf ๅฏนๆฏ” bartowski IQ4_XS (4.88 GiB):็›ธๅฏนไบŽ BF16 ๅŽŸๅง‹ๆจกๅž‹,ๅœจๅ…จ้ƒจ 3 ไธชๆ–‡ๆœฌไธŠๅนณๅ‡ KL ๆ•ฃๅบฆ ไฝŽ 18 %(ๅฐ 0.04 GiB)ใ€‚
  • Sakura-Fara1.5-9B-5.54GiB.gguf ๅฏนๆฏ” bartowski Q4_K_M (5.50 GiB):็›ธๅฏนไบŽ BF16 ๅŽŸๅง‹ๆจกๅž‹,ๅœจๅ…จ้ƒจ 3 ไธชๆ–‡ๆœฌไธŠๅนณๅ‡ KL ๆ•ฃๅบฆ ไฝŽ 40 %(ๅคง 0.04 GiB)ใ€‚

ๅŽŸๅ› :ๆฏไธชๅผ ้‡็š„ๆฏ”็‰น้ข„็ฎ—ๆ˜ฏๆ นๆฎๅฎžๆต‹็š„ๆ•ๆ„Ÿๅบฆๅˆ†้…็š„(่ง ๅˆถไฝœๆ–นๅผ),่€Œไธๆ˜ฏๆŒ‰ๅ›บๅฎš่ง„ๅˆ™ใ€‚ๅฎŒๆ•ด่กจๆ ผ่งไธ‹ๆ–‡ใ€‚

ๆˆ‘่ฏฅ้€‰ๅ“ชไธชๆ–‡ไปถ?

ๆ–‡ไปถ ้€‚็”จๅœบๆ™ฏ
Sakura-Fara1.5-9B-3.91GiB.gguf(3.90 GiB,3.74 bpw) ๆœ€ๅฐ:้€‚ๅˆๅ†…ๅญ˜้ข„็ฎ—็ดงๅผ ็š„ๆƒ…ๅ†ต;้ข„่ฎก่ดจ้‡ไผšๆœ‰ๆ˜Žๆ˜พๆŸๅคฑ
Sakura-Fara1.5-9B-4.84GiB.gguf(4.83 GiB,4.64 bpw) ๅฑ…ไธญ:ๅฏนๅคงๅคšๆ•ฐๆœบๅ™จ่€Œ่จ€,ๅคงๅฐไธŽ่ดจ้‡็š„ๆœ€ไฝณๅนณ่กก
Sakura-Fara1.5-9B-5.54GiB.gguf(5.53 GiB,5.31 bpw) ๆœ€ๅคง:ไธ‰่€…ไธญๆœ€ๆŽฅ่ฟ‘ BF16 ๅŽŸๅง‹ๆจกๅž‹

่ฟ™ไธ‰ไธชๆ–‡ไปถ

ๆ–‡ไปถ ๅคงๅฐ bits/weight KLD en KLD dev KLD wiki same top token (avg) ๅ„ๅผ ้‡็ฑปๅž‹(ๆŒ‰ๅคงๅฐ)
Sakura-Fara1.5-9B-3.91GiB.gguf 3.90 GiB 3.74 0.0611 0.0600 0.1071 89.8 % IQ3_S 41%, IQ4_XS 27%, Q3_K 21%, Q4_K 8%, Q5_K 2%
Sakura-Fara1.5-9B-4.84GiB.gguf 4.83 GiB 4.64 0.0158 0.0171 0.0282 94.6 % Q4_K 40%, IQ4_XS 33%, Q5_K 26%
Sakura-Fara1.5-9B-5.54GiB.gguf 5.53 GiB 5.31 0.0082 0.0085 0.0114 96.4 % Q5_K 83%, IQ4_XS 9%, Q4_K 6%

โ€œๅ„ๅผ ้‡็ฑปๅž‹(ๆŒ‰ๅคงๅฐ)โ€ๅˆ—ๅ‡บๆฏไธช ggml ็ฑปๅž‹ๅ ๆ–‡ไปถๅญ—่Š‚ๆ•ฐ็š„ๆฏ”ไพ‹(ๅ…จ้ƒจๆ˜ฏๆ ‡ๅ‡†็š„ llama.cpp ็ฑปๅž‹,ๅ› ๆญค่ฟ™ไบ›ๆ–‡ไปถๅฏๅœจๅŽŸ็‰ˆ llama.cpp ไธญ่ฟ่กŒ)ใ€‚่ดจ้‡ๅˆ—:้‡ๅŒ–ๆจกๅž‹็š„ไธ‹ไธ€ไธช token ๅˆ†ๅธƒ็›ธๅฏนไบŽ BF16 ๆบๆ–‡ไปถๅœจ็•™ๅ‡บๆ–‡ๆœฌไธŠ็š„ๅนณๅ‡ KL ๆ•ฃๅบฆ(่ถŠไฝŽ่ถŠๅฅฝ),ไปฅๅŠๆœ€ๅฏ่ƒฝ็š„ token ไฟๆŒไธๅ˜็š„้ข‘็އ(่ถŠ้ซ˜่ถŠๅฅฝ)ใ€‚

  • mmproj-Fara1.5-9B-f16.gguf:่ง†่ง‰ๆŠ•ๅฝฑๅ™จ(ไธŽ bartowski ็š„ F16 ่ฝฌๆข็‰ˆๆœฌ็›ธๅŒ,ๆœชๆ”นๅŠจ)ใ€‚

็›ธ่ฟ‘ๅคงๅฐ็š„ๆฏ”่พƒ

็›ธๅŒ็š„ๆ–‡ๆœฌใ€็›ธๅŒ็š„ไธŠไธ‹ๆ–‡ใ€็›ธๅŒ็š„ๅˆ†ๅ—,็”ฑๆˆ‘ไปฌไฝฟ็”จ llama-perplexity ๆต‹้‡(BF16 ๆบๆ–‡ไปถ = ๅ‚็…ง;BF16 ๆจกๅž‹็š„ PPL:en: 1.901,dev: 2.044,wiki: 8.878)ใ€‚ๅ‚็…ง้‡ๅŒ–ไธบ bartowski ๅฏนๅŒไธ€ๆจกๅž‹็š„ imatrix ้‡ๅŒ–,ไปฅ็›ธๅŒๆ–นๅผๆต‹้‡ใ€‚ๆŒ‰ๅคงๅฐๆŽ’ๅบใ€‚

้‡ๅŒ– ๆฅๆบ ๅคงๅฐ KLD en KLD dev KLD wiki PPL en same top token (avg)
Sakura-Fara1.5-9B-3.91GiB.gguf ๆœฌไป“ๅบ“ 3.91 GiB 0.0611 0.0600 0.1071 1.963 89.8 %
IQ3_XXS bartowski (imatrix) 3.98 GiB 0.0937 0.0981 0.1662 1.993 88.0 %
Sakura-Fara1.5-9B-4.84GiB.gguf ๆœฌไป“ๅบ“ 4.84 GiB 0.0158 0.0171 0.0282 1.926 94.6 %
IQ4_XS bartowski (imatrix) 4.88 GiB 0.0194 0.0224 0.0325 1.913 94.6 %
Q4_K_M bartowski (imatrix) 5.50 GiB 0.0129 0.0131 0.0218 1.912 95.6 %
Sakura-Fara1.5-9B-5.54GiB.gguf ๆœฌไป“ๅบ“ 5.54 GiB 0.0082 0.0085 0.0114 1.918 96.4 %

ๆ–นๆณ•:ๆฏไธชๆ–‡ๆœฌๅ– 12 ไธชๅ—,ๆฏๅ— 512 ไธช token;ๆ–‡ๆœฌไธบ:่‹ฑๆ–‡็•™ๅ‡บๆ–‡ๆœฌใ€ๅผ€ๅ‘่€…ๆ–‡ๆœฌ็•™ๅ‡บ้›†ใ€่‹ฑๆ–‡็™พ็ง‘็•™ๅ‡บๆ–‡ๆœฌ(ๆœช็”จไบŽไปปไฝ•ๆ กๅ‡†)ใ€‚่ฟ™ๆ˜ฏๅœจไธ€ๅฐๆœบๅ™จไธŠ็š„ๅ•ๆฌกๆต‹้‡;็›ธ่ฟ‘ๅคงๅฐไธ‹ KLD ็š„็ป†ๅฐๅทฎๅผ‚ๅนถไธๆž„ๆˆ่ดจ้‡ๆŽ’ๅใ€‚

ๅˆถไฝœๆ–นๅผ

  1. ๅŽŸๅง‹ๆจกๅž‹็š„ BF16 GGUF ๅ’Œ้‡่ฆๆ€ง็Ÿฉ้˜ตๆฅ่‡ช bartowski/Fara1.5-9B-GGUF;ๆˆ‘ไปฌๅŽŸๆ ทไฝฟ็”จไบ†่ฟ™ไธค่€…(ๆ„Ÿ่ฐขๅนถๅฝ’ๅŠŸไบŽ bartowski)ใ€‚
  2. ๆฏไธชๅคงๅž‹ๆƒ้‡็Ÿฉ้˜ต้’ˆๅฏนๆฏ็งๅ€™้€‰็ฑปๅž‹(Q2_Kใ€IQ2_Sใ€IQ3_XXSใ€IQ3_Sใ€Q3_Kใ€IQ4_XSใ€Q4_Kใ€Q5_Kใ€Q6_Kใ€Q8_0)ไฝฟ็”จ llama-quantize ๅ’Œ้‡่ฆๆ€ง็Ÿฉ้˜ตๅ„้‡ๅŒ–ไธ€ๆฌก;ๆฏ็ง้€‰ๆ‹ฉ็š„่ฏฏๅทฎๆŒ‰็Ÿฉ้˜ตไผฐ่ฎก(ๆŒ‰้‡่ฆๆ€งๅŠ ๆƒ),ๅนถไน˜ไปฅไธ€ไธช ็”จ็œŸๅฎž KL ๆ•ฃๅบฆๆต‹ๅพ—็š„ๆ•ๆ„Ÿๅบฆ็ณปๆ•ฐ:ๆฏ็ป„ๅผ ้‡(ไพ‹ๅฆ‚ FFN down ๆŠ•ๅฝฑใ€ๆณจๆ„ๅŠ› value ๆŠ•ๅฝฑใ€้ฆ–ๅฐพๅ‡ ๅฑ‚)่ขซๅ•็‹ฌ็งปๅˆฐ่พƒไฝŽๆฏ”็‰น็š„็ฑปๅž‹,ๅนถๅฐ†ๅฎƒ้€ ๆˆ็š„ KL ๆ•ฃๅบฆไธŽๅ…ถ้‡ๅŒ–่ฏฏๅทฎ่ฟ›่กŒๆฏ”่พƒใ€‚
  3. ็ฒพ็กฎ็š„้ข„็ฎ—ๅˆ†้…(ไปฅๅญ—่Š‚ไธบๅ•ไฝ็š„ๅคš้€‰่ƒŒๅŒ…้—ฎ้ข˜)ไธบๆฏไธช็›ฎๆ ‡ๅคงๅฐไธบๆฏไธช็Ÿฉ้˜ตๆŒ‘้€‰ไธ€็ง็ฑปๅž‹;ๆœ€็ปˆๆจกๅž‹็”ฑๅทฒๅญ˜ๅ‚จ็š„ๅผ ้‡็ป„่ฃ…่€Œๆˆ,ๆ— ้œ€้‡ๆ–ฐ้‡ๅŒ–ใ€‚ ๆˆ‘ไปฌไธๅ…ฌๅธƒ้€ๅผ ้‡็š„้€‰ๆ‹ฉใ€‚ๅฝ’ไธ€ๅŒ–ๅฑ‚ใ€ๅฐๅผ ้‡ๅ’Œๅคš token ้ข„ๆต‹ๅฑ‚ไฟๆŒ้ซ˜็ฒพๅบฆใ€‚

ๅฑ€้™

  • ไป…็”จ็Ÿญ็š„็•™ๅ‡บๆ–‡ๆœฌไธŠ็š„ KL ๆ•ฃๅบฆๅ’Œๅ›ฐๆƒ‘ๅบฆๆต‹้‡,ๆฒกๆœ‰ไฝฟ็”จไธ‹ๆธธๅŸบๅ‡†ๆต‹่ฏ•;่ฏทๅ‹ฟๅฐ†ๅ…ถ็†่งฃไธบไปปๅŠก่ดจ้‡็š„ๅฃฐๆ˜Žใ€‚
  • ้‡ๅŒ–ๆ€ปไผšๆŸๅคฑไธ€ไบ›่ดจ้‡;ๆœ€ๅฐ็š„ๆ–‡ไปถๆŸๅคฑๆœ€ๅคงใ€‚
  • Fara ๆ˜ฏไธ€ไธช ่ฎก็ฎ—ๆœบไฝฟ็”จๆ™บ่ƒฝไฝ“ใ€‚Microsoft ๅปบ่ฎฎไป…ๅœจๅธฆ็›‘ๆŽง็š„ๆฒ™็ฎฑ(MagenticLite)ไธญ่ฟ่กŒๅฎƒ,ๅนถๆ่ฟฐไบ†ๅ…ถๅฎ‰ๅ…จ่กŒไธบ(ๅœจๆถ‰ๅŠไป˜ๆฌพใ€ไธชไบบๆ•ฐๆฎๅ’Œไธๅฏ้€†ๆ“ไฝœไน‹ๅ‰ๅœไธ‹ๅนถ่ฏข้—ฎ)ใ€‚้‡ๅŒ–ๅฏ่ƒฝๆ”นๅ˜่กŒไธบ;ๆˆ‘ไปฌ ๆฒกๆœ‰ ๆต‹่ฏ•ๆ™บ่ƒฝไฝ“็š„ๅฎ‰ๅ…จๆ€งๆˆ–ๆˆชๅ›พๅฎšไฝ่ƒฝๅŠ›,ๅชๆต‹่ฏ•ไบ†ๆ–‡ๆœฌๅ›ฐๆƒ‘ๅบฆ/KLDใ€‚
  • ่ง†่ง‰ๆŠ•ๅฝฑๅ™จ mmproj-Fara1.5-9B-f16.gguf ๆ˜ฏ bartowski ็š„่ฝฌๆข็‰ˆๆœฌ(ๆœชๆ”นๅŠจ,F16);ๆˆชๅ›พ้œ€่ฆๅฎƒ,ๆˆ‘ไปฌๆฒกๆœ‰ๅฏนๅฎƒ่ฟ›่กŒ้‡ๅŒ–ใ€‚

่‡ด่ฐข

  • ๅŽŸๅง‹ๆจกๅž‹:Fara1.5-9B (Microsoft, MIT),MIT ่ฎธๅฏ่ฏ(ไปฅ LICENSE ้š้™„)ใ€‚
  • BF16 GGUFใ€้‡่ฆๆ€ง็Ÿฉ้˜ตๅ’Œๅ‚็…ง้‡ๅŒ–:bartowskiใ€‚
  • GGUF ๆ ผๅผๅ’Œๅทฅๅ…ท:ggml-org/llama.cpp (MIT)ใ€‚
Downloads last month
4,298
GGUF
Model size
9B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for webmp3/Sakura-Fara1.5-9B-GGUF

Finetuned
Qwen/Qwen3.5-9B
Quantized
(11)
this model

Spaces using webmp3/Sakura-Fara1.5-9B-GGUF 2

Collection including webmp3/Sakura-Fara1.5-9B-GGUF