Instructions to use webmp3/Sakura-MetaEncoder-30B-Experiment-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use webmp3/Sakura-MetaEncoder-30B-Experiment-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf webmp3/Sakura-MetaEncoder-30B-Experiment-GGUF:Q4_K_XL # Run inference directly in the terminal: llama cli -hf webmp3/Sakura-MetaEncoder-30B-Experiment-GGUF:Q4_K_XL
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf webmp3/Sakura-MetaEncoder-30B-Experiment-GGUF:Q4_K_XL # Run inference directly in the terminal: llama cli -hf webmp3/Sakura-MetaEncoder-30B-Experiment-GGUF:Q4_K_XL
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf webmp3/Sakura-MetaEncoder-30B-Experiment-GGUF:Q4_K_XL # Run inference directly in the terminal: ./llama-cli -hf webmp3/Sakura-MetaEncoder-30B-Experiment-GGUF:Q4_K_XL
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf webmp3/Sakura-MetaEncoder-30B-Experiment-GGUF:Q4_K_XL # Run inference directly in the terminal: ./build/bin/llama-cli -hf webmp3/Sakura-MetaEncoder-30B-Experiment-GGUF:Q4_K_XL
Use Docker
docker model run hf.co/webmp3/Sakura-MetaEncoder-30B-Experiment-GGUF:Q4_K_XL
- LM Studio
- Jan
- Ollama
How to use webmp3/Sakura-MetaEncoder-30B-Experiment-GGUF with Ollama:
ollama run hf.co/webmp3/Sakura-MetaEncoder-30B-Experiment-GGUF:Q4_K_XL
- Unsloth Desktop
- Pi
How to use webmp3/Sakura-MetaEncoder-30B-Experiment-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf webmp3/Sakura-MetaEncoder-30B-Experiment-GGUF:Q4_K_XL
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "webmp3/Sakura-MetaEncoder-30B-Experiment-GGUF:Q4_K_XL" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use webmp3/Sakura-MetaEncoder-30B-Experiment-GGUF with Docker Model Runner:
docker model run hf.co/webmp3/Sakura-MetaEncoder-30B-Experiment-GGUF:Q4_K_XL
- Lemonade
How to use webmp3/Sakura-MetaEncoder-30B-Experiment-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull webmp3/Sakura-MetaEncoder-30B-Experiment-GGUF:Q4_K_XL
Run and chat with the model
lemonade run user.Sakura-MetaEncoder-30B-Experiment-GGUF-Q4_K_XL
List all available models
lemonade list
- Hermes Agent
How to use webmp3/Sakura-MetaEncoder-30B-Experiment-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf webmp3/Sakura-MetaEncoder-30B-Experiment-GGUF:Q4_K_XL
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default webmp3/Sakura-MetaEncoder-30B-Experiment-GGUF:Q4_K_XL
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use webmp3/Sakura-MetaEncoder-30B-Experiment-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf webmp3/Sakura-MetaEncoder-30B-Experiment-GGUF:Q4_K_XL
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "webmp3/Sakura-MetaEncoder-30B-Experiment-GGUF:Q4_K_XL" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf webmp3/Sakura-MetaEncoder-30B-Experiment-GGUF:Q4_K_XL# Run inference directly in the terminal:
llama cli -hf webmp3/Sakura-MetaEncoder-30B-Experiment-GGUF:Q4_K_XLUse pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf webmp3/Sakura-MetaEncoder-30B-Experiment-GGUF:Q4_K_XL# Run inference directly in the terminal:
./llama-cli -hf webmp3/Sakura-MetaEncoder-30B-Experiment-GGUF:Q4_K_XLBuild from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf webmp3/Sakura-MetaEncoder-30B-Experiment-GGUF:Q4_K_XL# Run inference directly in the terminal:
./build/bin/llama-cli -hf webmp3/Sakura-MetaEncoder-30B-Experiment-GGUF:Q4_K_XLUse Docker
docker model run hf.co/webmp3/Sakura-MetaEncoder-30B-Experiment-GGUF:Q4_K_XLSakura — MetaEncoder-30B Experiment (GGUF)
Experiment, not an official release. Independent community work, not made or endorsed by Meta. A Hub search on 2026-10-09 found no quantization or GGUF of MetaEncoder, so we rebuilt one ourselves from Meta's public weights. It is an approximation of facebook/meta-encoder and has not been compared with the original BF16 model. Read What we did not test.
MetaEncoder-30B is Meta's embedding model: a contrastive fine-tune of Muse-Glimmer-30B. You give it a task and candidates in natural language and rank the candidates by an inner product. Because only the language-model projections differ from Muse-Glimmer (vision tower, embeddings, output head and all norms are identical), the difference can be stored as a low-rank adapter. This repository has one model file (the adapter is an intermediate step and is not published):
| File | Size | What it is |
|---|---|---|
Sakura-MetaEncoder-30B-Q4_K_XL-18.30GiB.gguf |
18.30 GiB | Ready to use. Meta's official Muse-Glimmer Q4_K_XL with the adapter merged in and re-quantized to the same per-tensor types (Q4_K/Q5_K/Q6_K). Same size and architecture as Meta's file. |
Quick start (text)
llama-server -m Sakura-MetaEncoder-30B-Q4_K_XL-18.30GiB.gguf --embeddings --pooling last -c 4096 -ngl 99 --port 8080
python metaencoder_prompt.py --query "Aspirin lowers the risk of colon cancer." \
--instruction "Retrieve a scientific abstract that supports or refutes the claim." \
--doc "Regular aspirin use was associated with fewer colorectal adenomas ..." --doc "A study of bicycle helmets ..."
metaencoder_prompt.py builds MetaEncoder's own prompt (system text Represent the user's input., the chat template of the original repo, exactly one BOS token) and ranks the candidates. The embedding is the last-token hidden state, L2-normalised. We evaluated only this prompt format (the one of the original model); we did not test others.
Needs a llama.cpp build that includes the muse-glimmer architecture (upstream master at the time of writing).
What we measured
One small retrieval test, same prompts, same subset, same machine for every row: NanoSciFact (zeta-alpha-ai/NanoSciFact), all 50 queries, corpus reduced to 600 documents (all 55 relevant ones plus 545 random ones), instruction Retrieve a scientific abstract that supports or refutes the claim., cosine similarity.
| Model | nDCG@10 | Recall@10 | MRR@10 |
|---|---|---|---|
| Meta's Muse-Glimmer Q4_K_XL, no adapter (not an embedding model) | 0.019 | 0.04 | 0.012 |
Muse-Glimmer Q4_K_XL + the same adapter applied at run time (--lora; the adapter file is not published) |
0.876 | 0.96 | 0.846 |
Sakura-MetaEncoder-30B-Q4_K_XL (adapter merged and re-quantized) |
0.901 | 0.96 | 0.885 |
The adapter turns the plain language model into a working retriever. Merging it into the Q4 file did not lose anything we could measure on this test (the 0.876 / 0.901 difference is within the noise of 50 queries and should not be read as an improvement).
The adapter is a rank-128 approximation: per tensor it captures at least 99.37 % (median 99.67 %) of the squared norm of the weight difference.
Numbers and conditions are in results.json.
How it was made
- For every one of the 416 language-model projections (
self_attn.{q,k,v,o,gate}_proj,mlp.{gate,up,down}_projin 52 layers) we computed the difference between the published BF16 weights offacebook/meta-encoderandmeta-models/Muse-Glimmer-30Band approximated it by a truncated SVD of rank 128. - The factors were packed as a PEFT adapter and converted with llama.cpp's
convert_lora_to_gguf.py(the Muse converter's RoPE un-permutation of Q/K is applied to the rows of the B factor). - The adapter was merged into Meta's
Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.ggufwithllama-export-loraand re-quantized withllama-quantizeto the original type of every tensor, without an importance matrix. 315 tensors (norms, token embeddings, output head) are byte-identical to Meta's file; the 416 merged ones are new.
What we did not test
- We did not run the original BF16 MetaEncoder on this subset, so we cannot say how close this file is to it. The results table of the original model card (for example NanoBEIR) is not comparable to our subset.
- Images and video (the original is multimodal) were not tested; these files were only used with text. Meta's separate
mmprojwas not tried. - Only one task type (scientific claim retrieval in English) and one instruction were evaluated. Other languages, instructions, the closed-set "criteria" prompts and long inputs are untested.
- The merge re-quantizes 416 tensors without an importance matrix, on top of Meta's own quantization, so small deviations from the original are expected. We did not measure a KL divergence against BF16.
- Speed was not benchmarked.
License and use
Licensed under Apache 2.0, inherited from the base models. Use is additionally subject to the Muse Glimmer Usage Policy (copy included). See NOTICE for what was changed relative to the originals. "Meta", "Muse", "Muse Glimmer" and "MetaEncoder" are names used by Meta Platforms, Inc.; they appear here only to identify the works this derivative is based on. Part of the Sakura Mini line.
Credits
- MetaEncoder-30B: facebook/meta-encoder (Meta Platforms, Inc., Apache-2.0), paper: Cheng et al., MetaEncoder: Exploring the Limit of Bi-Encoders for Multimodal System One Decision Making with Natural Language Interface, arXiv:2610.11316.
- Base model and the Q4_K_XL GGUF: meta-models/Muse-Glimmer-30B, meta-models/Muse-Glimmer-30B-GGUF.
- GGUF format and tools: ggml-org/llama.cpp (MIT).
中文说明 · 樱花 (Simplified Chinese)
English above. 本节为上文的中文翻译(Sakura = 樱花 yīnghuā);完整的独立中文版见 README_zh.md。
Sakura — MetaEncoder-30B 实验版(GGUF)
实验性质,非官方发布。 这是独立的社区工作,并非由 Meta 制作或认可。2026-10-09 在 Hub 上搜索时,我们没有找到 MetaEncoder 的量化版或 GGUF,因此我们用 Meta 公开的权重自行重建了一个。它是 facebook/meta-encoder 的 近似版本,没有与原始 BF16 模型做过对比。请阅读 我们没有测试的内容。
MetaEncoder-30B 是 Meta 的嵌入模型:对 Muse-Glimmer-30B 进行对比学习微调得到。你用自然语言给出一个 任务 和若干 候选项,按内积对候选项排序。 由于与 Muse-Glimmer 相比只有语言模型的投影层不同(视觉塔、词嵌入、输出头和所有归一化层相同),这个差异可以存为低秩适配器。本仓库包含一个模型文件(适配器只是中间步骤,未发布):
| 文件 | 大小 | 说明 |
|---|---|---|
Sakura-MetaEncoder-30B-Q4_K_XL-18.30GiB.gguf |
18.30 GiB | 可直接使用。 在 Meta 官方 Muse-Glimmer Q4_K_XL 的基础上合入适配器,并按每个张量原来的类型(Q4_K/Q5_K/Q6_K)重新量化。大小和架构与 Meta 的文件相同。 |
快速开始(文本)
llama-server -m Sakura-MetaEncoder-30B-Q4_K_XL-18.30GiB.gguf --embeddings --pooling last -c 4096 -ngl 99 --port 8080
python metaencoder_prompt.py --query "Aspirin lowers the risk of colon cancer." \
--instruction "Retrieve a scientific abstract that supports or refutes the claim." \
--doc "Regular aspirin use was associated with fewer colorectal adenomas ..." --doc "A study of bicycle helmets ..."
metaencoder_prompt.py 会构造 MetaEncoder 自己的提示词(系统文本 *Represent the user's input.*、原仓库的聊天模板、恰好一个 BOS 标记)并对候选项排序。嵌入是最后一个 token 的隐藏状态,经 L2 归一化。我们只评测了这一种提示词格式(原模型使用的格式),没有测试其他格式。
需要包含 muse-glimmer 架构的 llama.cpp 构建(撰写时为上游 master)。
我们的测量
一个小型检索测试,每一行使用相同的提示词、相同的子集、同一台机器:NanoSciFact(zeta-alpha-ai/NanoSciFact),全部 50 个查询,语料缩减为 600 篇文档(全部 55 篇相关文档加 545 篇随机文档),指令为 Retrieve a scientific abstract that supports or refutes the claim.,使用余弦相似度。
| 模型 | nDCG@10 | Recall@10 | MRR@10 |
|---|---|---|---|
| Meta 的 Muse-Glimmer Q4_K_XL,无适配器(不是嵌入模型) | 0.019 | 0.04 | 0.012 |
Muse-Glimmer Q4_K_XL + 运行时加载同一个适配器(--lora;适配器文件未发布) |
0.876 | 0.96 | 0.846 |
Sakura-MetaEncoder-30B-Q4_K_XL(适配器合入并重新量化) |
0.901 | 0.96 | 0.885 |
适配器把普通语言模型变成了可用的检索器。把它合入 Q4 文件在这个测试中没有造成我们能测出的损失(0.876 与 0.901 的差异在 50 个查询的噪声范围内,不应理解为提升)。
适配器是秩 128 的近似:每个张量至少捕获权重差异平方范数的 99.37 %(中位数 99.67 %)。
数值和测量条件见 results.json。
制作方法
- 对 416 个语言模型投影(52 层中的
self_attn.{q,k,v,o,gate}_proj、mlp.{gate,up,down}_proj),我们用facebook/meta-encoder与meta-models/Muse-Glimmer-30B公开发布的 BF16 权重计算差异,并用秩 128 的截断 SVD 近似。 - 因子被打包为 PEFT 适配器,并用 llama.cpp 的
convert_lora_to_gguf.py转换(Muse 转换器对 Q/K 的 RoPE 反置换作用在 B 因子的行上)。 - 用
llama-export-lora把适配器合入 Meta 的Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf,再用llama-quantize按每个张量的原始类型重新量化,不使用重要性矩阵。315 个张量(归一化层、词嵌入、输出头)与 Meta 的文件字节完全相同;被合入的 416 个张量是新的。
我们没有测试的内容
- 我们没有在这个子集上运行原始 BF16 的 MetaEncoder,因此无法说明这个文件与它有多接近。原模型卡中的结果表(例如 NanoBEIR)与我们的子集不可比。
- 图像和视频(原模型是多模态的)没有测试;这些文件只用于文本。Meta 单独的
mmproj也没有尝试。 - 只评测了一种任务类型(英文科学论断检索)和一条指令。其他语言、其他指令、封闭集合的"criteria"提示词以及长输入都没有测试。
- 合并过程在 Meta 自己的量化之上,对 416 个张量在没有重要性矩阵的情况下重新量化,因此与原模型相比会有小的偏差。我们没有测量相对于 BF16 的 KL 散度。
- 没有做速度基准测试。
许可证与使用
采用 Apache 2.0,继承自基础模型。使用还需遵守 Muse Glimmer 的 使用政策(附有副本)。相对于原始模型的改动见 NOTICE。"Meta"、"Muse"、"Muse Glimmer" 和 "MetaEncoder" 是 Meta Platforms, Inc. 使用的名称;它们在此仅用于标明本衍生作品所基于的作品。 属于 Sakura Mini 系列。
致谢
- MetaEncoder-30B:facebook/meta-encoder(Meta Platforms, Inc.,Apache-2.0),论文:Cheng et al., MetaEncoder: Exploring the Limit of Bi-Encoders for Multimodal System One Decision Making with Natural Language Interface, arXiv:2610.11316。
- 基础模型和 Q4_K_XL GGUF:meta-models/Muse-Glimmer-30B、meta-models/Muse-Glimmer-30B-GGUF。
- GGUF 格式和工具:ggml-org/llama.cpp(MIT)。
- Downloads last month
- 78
4-bit
Install (macOS, Linux)
# Start a local OpenAI-compatible server with a web UI: llama serve -hf webmp3/Sakura-MetaEncoder-30B-Experiment-GGUF:Q4_K_XL# Run inference directly in the terminal: llama cli -hf webmp3/Sakura-MetaEncoder-30B-Experiment-GGUF:Q4_K_XL