Instructions to use webmp3/Sakura-Fara1.5-9B-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use webmp3/Sakura-Fara1.5-9B-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf webmp3/Sakura-Fara1.5-9B-GGUF:F16 # Run inference directly in the terminal: llama cli -hf webmp3/Sakura-Fara1.5-9B-GGUF:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf webmp3/Sakura-Fara1.5-9B-GGUF:F16 # Run inference directly in the terminal: llama cli -hf webmp3/Sakura-Fara1.5-9B-GGUF:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf webmp3/Sakura-Fara1.5-9B-GGUF:F16 # Run inference directly in the terminal: ./llama-cli -hf webmp3/Sakura-Fara1.5-9B-GGUF:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf webmp3/Sakura-Fara1.5-9B-GGUF:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf webmp3/Sakura-Fara1.5-9B-GGUF:F16
Use Docker
docker model run hf.co/webmp3/Sakura-Fara1.5-9B-GGUF:F16
- LM Studio
- Jan
- vLLM
How to use webmp3/Sakura-Fara1.5-9B-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "webmp3/Sakura-Fara1.5-9B-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "webmp3/Sakura-Fara1.5-9B-GGUF", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/webmp3/Sakura-Fara1.5-9B-GGUF:F16
- Ollama
How to use webmp3/Sakura-Fara1.5-9B-GGUF with Ollama:
ollama run hf.co/webmp3/Sakura-Fara1.5-9B-GGUF:F16
- Unsloth Desktop
- Pi
How to use webmp3/Sakura-Fara1.5-9B-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf webmp3/Sakura-Fara1.5-9B-GGUF:F16
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "webmp3/Sakura-Fara1.5-9B-GGUF:F16" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use webmp3/Sakura-Fara1.5-9B-GGUF with Docker Model Runner:
docker model run hf.co/webmp3/Sakura-Fara1.5-9B-GGUF:F16
- Lemonade
How to use webmp3/Sakura-Fara1.5-9B-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull webmp3/Sakura-Fara1.5-9B-GGUF:F16
Run and chat with the model
lemonade run user.Sakura-Fara1.5-9B-GGUF-F16
List all available models
lemonade list
- Hermes Agent
How to use webmp3/Sakura-Fara1.5-9B-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf webmp3/Sakura-Fara1.5-9B-GGUF:F16
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default webmp3/Sakura-Fara1.5-9B-GGUF:F16
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use webmp3/Sakura-Fara1.5-9B-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf webmp3/Sakura-Fara1.5-9B-GGUF:F16
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "webmp3/Sakura-Fara1.5-9B-GGUF:F16" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Sakura โ Fara1.5-9B (GGUF, three sizes)
Fara1.5-9B is Microsoft's computer-use agent for web browsing (Qwen3.5-9B base, multimodal: it reads browser screenshots and emits structured tool calls). This repository holds three GGUF files of different size made from the BF16 weights with our measured mixed-codec method. Pick one on the Files page; the table below says what each file is. Independent community quantization, not an official Microsoft release. Part of the Sakura Mini line.
Quality at a glance
Same method, same texts, same machine for every file; lower KL divergence means closer to the BF16 original:
Sakura-Fara1.5-9B-3.91GiB.ggufvs bartowski IQ3_XXS (3.98 GiB): KL divergence to the BF16 original 36 % lower on average over all 3 texts (0.07 GiB smaller).Sakura-Fara1.5-9B-4.84GiB.ggufvs bartowski IQ4_XS (4.88 GiB): KL divergence to the BF16 original 18 % lower on average over all 3 texts (0.04 GiB smaller).Sakura-Fara1.5-9B-5.54GiB.ggufvs bartowski Q4_K_M (5.50 GiB): KL divergence to the BF16 original 40 % lower on average over all 3 texts (0.04 GiB larger).
Why: the bit budget per tensor is allocated from measured sensitivity (see How it was made), not from fixed rules. Full table below.
Which file should I take?
| File | For |
|---|---|
Sakura-Fara1.5-9B-3.91GiB.gguf (3.90 GiB, 3.74 bpw) |
smallest: for tight memory budgets; expect visible quality loss |
Sakura-Fara1.5-9B-4.84GiB.gguf (4.83 GiB, 4.64 bpw) |
middle: best balance of size and quality for most machines |
Sakura-Fara1.5-9B-5.54GiB.gguf (5.53 GiB, 5.31 bpw) |
largest: closest to the BF16 original of the three |
The three files
| File | Size | bits/weight | KLD en | KLD dev | KLD wiki | same top token (avg) | Tensor types by size |
|---|---|---|---|---|---|---|---|
Sakura-Fara1.5-9B-3.91GiB.gguf |
3.90 GiB | 3.74 | 0.0611 | 0.0600 | 0.1071 | 89.8 % | IQ3_S 41%, IQ4_XS 27%, Q3_K 21%, Q4_K 8%, Q5_K 2% |
Sakura-Fara1.5-9B-4.84GiB.gguf |
4.83 GiB | 4.64 | 0.0158 | 0.0171 | 0.0282 | 94.6 % | Q4_K 40%, IQ4_XS 33%, Q5_K 26% |
Sakura-Fara1.5-9B-5.54GiB.gguf |
5.53 GiB | 5.31 | 0.0082 | 0.0085 | 0.0114 | 96.4 % | Q5_K 83%, IQ4_XS 9%, Q4_K 6% |
"Tensor types by size" lists the share of the file's bytes per ggml type (all are standard llama.cpp types, so the files run in stock llama.cpp). Quality columns: mean KL divergence of the quantized model's next-token distribution against the BF16 source on held-out text (lower is better) and how often the most likely token is unchanged (higher is better).
mmproj-Fara1.5-9B-f16.gguf: vision projector (unchanged from bartowski's F16 conversion).
Comparison at similar size
Same texts, same context, same chunks, measured by us with llama-perplexity (BF16 source = reference; PPL of the BF16 model: en: 1.901, dev: 2.044, wiki: 8.878). The reference quants are bartowski's imatrix quants of the same model, measured the same way. Sorted by size.
| Quant | Source | Size | KLD en | KLD dev | KLD wiki | PPL en | same top token (avg) |
|---|---|---|---|---|---|---|---|
Sakura-Fara1.5-9B-3.91GiB.gguf |
this repository | 3.91 GiB | 0.0611 | 0.0600 | 0.1071 | 1.963 | 89.8 % |
| IQ3_XXS | bartowski (imatrix) | 3.98 GiB | 0.0937 | 0.0981 | 0.1662 | 1.993 | 88.0 % |
Sakura-Fara1.5-9B-4.84GiB.gguf |
this repository | 4.84 GiB | 0.0158 | 0.0171 | 0.0282 | 1.926 | 94.6 % |
| IQ4_XS | bartowski (imatrix) | 4.88 GiB | 0.0194 | 0.0224 | 0.0325 | 1.913 | 94.6 % |
| Q4_K_M | bartowski (imatrix) | 5.50 GiB | 0.0129 | 0.0131 | 0.0218 | 1.912 | 95.6 % |
Sakura-Fara1.5-9B-5.54GiB.gguf |
this repository | 5.54 GiB | 0.0082 | 0.0085 | 0.0114 | 1.918 | 96.4 % |
Method: 12 chunks of 512 tokens per text; texts: English held-out text, developer-text held-out set, English encyclopedia held-out text (not used in any calibration). Single measurement on one machine; small KLD differences at similar size are not a quality ranking.
How it was made
- BF16 GGUF of the original model and the importance matrix come from bartowski/Fara1.5-9B-GGUF; we used both unchanged (credit and thanks to bartowski).
- Every large weight matrix was quantized once per candidate type (Q2_K, IQ2_S, IQ3_XXS, IQ3_S, Q3_K, IQ4_XS, Q4_K, Q5_K, Q6_K, Q8_0) with
llama-quantizeand the importance matrix; the error of each choice was estimated per matrix (importance-weighted) and scaled by a sensitivity factor measured with real KL divergence: every group of tensors (for example the FFN down projections, the attention value projections, the first and last layers) was moved alone to a lower-bit type, and the KL divergence it caused was compared with its quantization error. - An exact budget allocation (multiple-choice knapsack over bytes) picked one type per matrix for each target size; the final model was assembled from the stored tensors without re-quantizing. We do not publish per-tensor choices. Norms, small tensors and the multi-token-prediction layers stay at high precision.
Limits
- Measured only with KL divergence and perplexity on short held-out texts and not with downstream benchmarks; do not read it as a task-quality claim.
- Quantization always costs some quality; the smallest file costs the most.
- Fara is a computer-use agent. Microsoft recommends running it only inside a sandbox with monitoring (MagenticLite) and describes safety behaviour (stopping and asking before payments, personal data and irreversible actions). Quantization can change behaviour; we did not test agent safety or screenshot grounding, only text perplexity/KLD.
- The vision projector
mmproj-Fara1.5-9B-f16.ggufis bartowski's conversion (unchanged, F16); it is needed for screenshots and is not quantized by us.
Credits
- Original model: Fara1.5-9B (Microsoft, MIT), MIT license (included as
LICENSE). - BF16 GGUF, importance matrix and the reference quants: bartowski.
- GGUF format and tools: ggml-org/llama.cpp (MIT).
ไธญๆ่ฏดๆ ยท ๆจฑ่ฑ (Simplified Chinese)
English above. ๆฌ่ไธบไธๆ็ไธญๆ็ฟป่ฏ(Sakura = ๆจฑ่ฑ yฤซnghuฤ);ๅฎๆด็็ฌ็ซไธญๆ็่ง README_zh.mdใ
Sakura โ Fara1.5-9B(GGUF,ไธ็งๅคงๅฐ)
Fara1.5-9B ๆฏ Microsoft ้ขๅ็ฝ้กตๆต่ง็่ฎก็ฎๆบไฝฟ็จ(computer-use)ๆบ่ฝไฝ(ๅบไบ Qwen3.5-9B,ๅคๆจกๆ:่ฏปๅๆต่งๅจๆชๅพๅนถ่พๅบ็ปๆๅ็ๅทฅๅ ท่ฐ็จ)ใๆฌไปๅบๅ ๅซ ไธไธชไธๅๅคงๅฐ็ GGUF ๆไปถ,็ฑ BF16 ๆ้้่ฟๆไปฌๅฎๆตๅพๅบ็ๆททๅ็ผ็ ๆนๆณๅถไฝใ่ฏทๅจ Files ้กต้ข้ๆฉๅ ถไธญไธไธช;ไธ่กจ่ฏดๆไบๆฏไธชๆไปถ็็จ้ใ ่ฟๆฏ็ฌ็ซ็็คพๅบ้ๅ็ๆฌ,ไธๆฏ Microsoft ็ๅฎๆนๅๅธใๅฑไบ Sakura Mini ็ณปๅใ
่ดจ้ๆฆ่ง
ๆฏไธชๆไปถไฝฟ็จ็ธๅ็ๆนๆณใ็ธๅ็ๆๆฌใๅไธๅฐๆบๅจ;KL ๆฃๅบฆ่ถไฝ,่ถๆฅ่ฟ BF16 ๅๅงๆจกๅ:
Sakura-Fara1.5-9B-3.91GiB.ggufๅฏนๆฏ bartowski IQ3_XXS (3.98 GiB):็ธๅฏนไบ BF16 ๅๅงๆจกๅ,ๅจๅ จ้จ 3 ไธชๆๆฌไธๅนณๅ KL ๆฃๅบฆ ไฝ 36 %(ๅฐ 0.07 GiB)ใSakura-Fara1.5-9B-4.84GiB.ggufๅฏนๆฏ bartowski IQ4_XS (4.88 GiB):็ธๅฏนไบ BF16 ๅๅงๆจกๅ,ๅจๅ จ้จ 3 ไธชๆๆฌไธๅนณๅ KL ๆฃๅบฆ ไฝ 18 %(ๅฐ 0.04 GiB)ใSakura-Fara1.5-9B-5.54GiB.ggufๅฏนๆฏ bartowski Q4_K_M (5.50 GiB):็ธๅฏนไบ BF16 ๅๅงๆจกๅ,ๅจๅ จ้จ 3 ไธชๆๆฌไธๅนณๅ KL ๆฃๅบฆ ไฝ 40 %(ๅคง 0.04 GiB)ใ
ๅๅ :ๆฏไธชๅผ ้็ๆฏ็น้ข็ฎๆฏๆ นๆฎๅฎๆต็ๆๆๅบฆๅ้ ็(่ง ๅถไฝๆนๅผ),่ไธๆฏๆๅบๅฎ่งๅใๅฎๆด่กจๆ ผ่งไธๆใ
ๆ่ฏฅ้ๅชไธชๆไปถ?
| ๆไปถ | ้็จๅบๆฏ |
|---|---|
Sakura-Fara1.5-9B-3.91GiB.gguf(3.90 GiB,3.74 bpw) |
ๆๅฐ:้ๅๅ ๅญ้ข็ฎ็ดงๅผ ็ๆ ๅต;้ข่ฎก่ดจ้ไผๆๆๆพๆๅคฑ |
Sakura-Fara1.5-9B-4.84GiB.gguf(4.83 GiB,4.64 bpw) |
ๅฑ ไธญ:ๅฏนๅคงๅคๆฐๆบๅจ่่จ,ๅคงๅฐไธ่ดจ้็ๆไฝณๅนณ่กก |
Sakura-Fara1.5-9B-5.54GiB.gguf(5.53 GiB,5.31 bpw) |
ๆๅคง:ไธ่ ไธญๆๆฅ่ฟ BF16 ๅๅงๆจกๅ |
่ฟไธไธชๆไปถ
| ๆไปถ | ๅคงๅฐ | bits/weight | KLD en | KLD dev | KLD wiki | same top token (avg) | ๅๅผ ้็ฑปๅ(ๆๅคงๅฐ) |
|---|---|---|---|---|---|---|---|
Sakura-Fara1.5-9B-3.91GiB.gguf |
3.90 GiB | 3.74 | 0.0611 | 0.0600 | 0.1071 | 89.8 % | IQ3_S 41%, IQ4_XS 27%, Q3_K 21%, Q4_K 8%, Q5_K 2% |
Sakura-Fara1.5-9B-4.84GiB.gguf |
4.83 GiB | 4.64 | 0.0158 | 0.0171 | 0.0282 | 94.6 % | Q4_K 40%, IQ4_XS 33%, Q5_K 26% |
Sakura-Fara1.5-9B-5.54GiB.gguf |
5.53 GiB | 5.31 | 0.0082 | 0.0085 | 0.0114 | 96.4 % | Q5_K 83%, IQ4_XS 9%, Q4_K 6% |
โๅๅผ ้็ฑปๅ(ๆๅคงๅฐ)โๅๅบๆฏไธช ggml ็ฑปๅๅ ๆไปถๅญ่ๆฐ็ๆฏไพ(ๅ จ้จๆฏๆ ๅ็ llama.cpp ็ฑปๅ,ๅ ๆญค่ฟไบๆไปถๅฏๅจๅ็ llama.cpp ไธญ่ฟ่ก)ใ่ดจ้ๅ:้ๅๆจกๅ็ไธไธไธช token ๅๅธ็ธๅฏนไบ BF16 ๆบๆไปถๅจ็ๅบๆๆฌไธ็ๅนณๅ KL ๆฃๅบฆ(่ถไฝ่ถๅฅฝ),ไปฅๅๆๅฏ่ฝ็ token ไฟๆไธๅ็้ข็(่ถ้ซ่ถๅฅฝ)ใ
mmproj-Fara1.5-9B-f16.gguf:่ง่งๆๅฝฑๅจ(ไธ bartowski ็ F16 ่ฝฌๆข็ๆฌ็ธๅ,ๆชๆนๅจ)ใ
็ธ่ฟๅคงๅฐ็ๆฏ่พ
็ธๅ็ๆๆฌใ็ธๅ็ไธไธๆใ็ธๅ็ๅๅ,็ฑๆไปฌไฝฟ็จ llama-perplexity ๆต้(BF16 ๆบๆไปถ = ๅ็
ง;BF16 ๆจกๅ็ PPL:en: 1.901,dev: 2.044,wiki: 8.878)ใๅ็
ง้ๅไธบ bartowski ๅฏนๅไธๆจกๅ็ imatrix ้ๅ,ไปฅ็ธๅๆนๅผๆต้ใๆๅคงๅฐๆๅบใ
| ้ๅ | ๆฅๆบ | ๅคงๅฐ | KLD en | KLD dev | KLD wiki | PPL en | same top token (avg) |
|---|---|---|---|---|---|---|---|
Sakura-Fara1.5-9B-3.91GiB.gguf |
ๆฌไปๅบ | 3.91 GiB | 0.0611 | 0.0600 | 0.1071 | 1.963 | 89.8 % |
| IQ3_XXS | bartowski (imatrix) | 3.98 GiB | 0.0937 | 0.0981 | 0.1662 | 1.993 | 88.0 % |
Sakura-Fara1.5-9B-4.84GiB.gguf |
ๆฌไปๅบ | 4.84 GiB | 0.0158 | 0.0171 | 0.0282 | 1.926 | 94.6 % |
| IQ4_XS | bartowski (imatrix) | 4.88 GiB | 0.0194 | 0.0224 | 0.0325 | 1.913 | 94.6 % |
| Q4_K_M | bartowski (imatrix) | 5.50 GiB | 0.0129 | 0.0131 | 0.0218 | 1.912 | 95.6 % |
Sakura-Fara1.5-9B-5.54GiB.gguf |
ๆฌไปๅบ | 5.54 GiB | 0.0082 | 0.0085 | 0.0114 | 1.918 | 96.4 % |
ๆนๆณ:ๆฏไธชๆๆฌๅ 12 ไธชๅ,ๆฏๅ 512 ไธช token;ๆๆฌไธบ:่ฑๆ็ๅบๆๆฌใๅผๅ่ ๆๆฌ็ๅบ้ใ่ฑๆ็พ็ง็ๅบๆๆฌ(ๆช็จไบไปปไฝๆ กๅ)ใ่ฟๆฏๅจไธๅฐๆบๅจไธ็ๅๆฌกๆต้;็ธ่ฟๅคงๅฐไธ KLD ็็ปๅฐๅทฎๅผๅนถไธๆๆ่ดจ้ๆๅใ
ๅถไฝๆนๅผ
- ๅๅงๆจกๅ็ BF16 GGUF ๅ้่ฆๆง็ฉ้ตๆฅ่ช bartowski/Fara1.5-9B-GGUF;ๆไปฌๅๆ ทไฝฟ็จไบ่ฟไธค่ (ๆ่ฐขๅนถๅฝๅไบ bartowski)ใ
- ๆฏไธชๅคงๅๆ้็ฉ้ต้ๅฏนๆฏ็งๅ้็ฑปๅ(Q2_KใIQ2_SใIQ3_XXSใIQ3_SใQ3_KใIQ4_XSใQ4_KใQ5_KใQ6_KใQ8_0)ไฝฟ็จ
llama-quantizeๅ้่ฆๆง็ฉ้ตๅ้ๅไธๆฌก;ๆฏ็ง้ๆฉ็่ฏฏๅทฎๆ็ฉ้ตไผฐ่ฎก(ๆ้่ฆๆงๅ ๆ),ๅนถไนไปฅไธไธช ็จ็ๅฎ KL ๆฃๅบฆๆตๅพ็ๆๆๅบฆ็ณปๆฐ:ๆฏ็ปๅผ ้(ไพๅฆ FFN down ๆๅฝฑใๆณจๆๅ value ๆๅฝฑใ้ฆๅฐพๅ ๅฑ)่ขซๅ็ฌ็งปๅฐ่พไฝๆฏ็น็็ฑปๅ,ๅนถๅฐๅฎ้ ๆ็ KL ๆฃๅบฆไธๅ ถ้ๅ่ฏฏๅทฎ่ฟ่กๆฏ่พใ - ็ฒพ็กฎ็้ข็ฎๅ้ (ไปฅๅญ่ไธบๅไฝ็ๅค้่ๅ ้ฎ้ข)ไธบๆฏไธช็ฎๆ ๅคงๅฐไธบๆฏไธช็ฉ้ตๆ้ไธ็ง็ฑปๅ;ๆ็ปๆจกๅ็ฑๅทฒๅญๅจ็ๅผ ้็ป่ฃ ่ๆ,ๆ ้้ๆฐ้ๅใ ๆไปฌไธๅ ฌๅธ้ๅผ ้็้ๆฉใๅฝไธๅๅฑใๅฐๅผ ้ๅๅค token ้ขๆตๅฑไฟๆ้ซ็ฒพๅบฆใ
ๅฑ้
- ไป ็จ็ญ็็ๅบๆๆฌไธ็ KL ๆฃๅบฆๅๅฐๆๅบฆๆต้,ๆฒกๆไฝฟ็จไธๆธธๅบๅๆต่ฏ;่ฏทๅฟๅฐๅ ถ็่งฃไธบไปปๅก่ดจ้็ๅฃฐๆใ
- ้ๅๆปไผๆๅคฑไธไบ่ดจ้;ๆๅฐ็ๆไปถๆๅคฑๆๅคงใ
- Fara ๆฏไธไธช ่ฎก็ฎๆบไฝฟ็จๆบ่ฝไฝใMicrosoft ๅปบ่ฎฎไป ๅจๅธฆ็ๆง็ๆฒ็ฎฑ(MagenticLite)ไธญ่ฟ่กๅฎ,ๅนถๆ่ฟฐไบๅ ถๅฎๅ จ่กไธบ(ๅจๆถๅไปๆฌพใไธชไบบๆฐๆฎๅไธๅฏ้ๆไฝไนๅๅไธๅนถ่ฏข้ฎ)ใ้ๅๅฏ่ฝๆนๅ่กไธบ;ๆไปฌ ๆฒกๆ ๆต่ฏๆบ่ฝไฝ็ๅฎๅ จๆงๆๆชๅพๅฎไฝ่ฝๅ,ๅชๆต่ฏไบๆๆฌๅฐๆๅบฆ/KLDใ
- ่ง่งๆๅฝฑๅจ
mmproj-Fara1.5-9B-f16.ggufๆฏ bartowski ็่ฝฌๆข็ๆฌ(ๆชๆนๅจ,F16);ๆชๅพ้่ฆๅฎ,ๆไปฌๆฒกๆๅฏนๅฎ่ฟ่ก้ๅใ
่ด่ฐข
- ๅๅงๆจกๅ:Fara1.5-9B (Microsoft, MIT),MIT ่ฎธๅฏ่ฏ(ไปฅ
LICENSE้้)ใ - BF16 GGUFใ้่ฆๆง็ฉ้ตๅๅ็ ง้ๅ:bartowskiใ
- GGUF ๆ ผๅผๅๅทฅๅ ท:ggml-org/llama.cpp (MIT)ใ
- Downloads last month
- 4,298
We're not able to determine the quantization variants.