BreadBowl-Embed
Encode once. Retrieve and rerank from the same stored representation.
BreadBowl-Embed is a retrieval model that represents text as 16 routing鈥搗alue slots. Each slot pairs a vector used to find relevant information with a value vector that a query can read. This lets the model retrieve candidates and refine their ranking using document representations computed ahead of time.
The idea is simple: an embedding can support both finding a document and reading query-relevant information from it. Once a document has been encoded, different queries can read different combinations of its stored values. Reranking those cached representations needs no additional backbone pass over document text and no separate cross-encoder.
Paper 路 Source code 路 Example notebook
This is the step-4000 KL+replay research preview. The model weights, architecture implementation, and Python inference toolkit are available under Apache 2.0.
Why this matters
A document is encoded before its future queries are known. BreadBowl keeps a small set of routing and value vectors so that the query can determine how to use that stored information at search time:
- Encode once: turn each passage into a fixed set of slots and cache them.
- Retrieve: compare query routing slots with document routing slots to select candidates.
- Read and rerank: use routing similarity to attend over each candidate's value slots, then combine routing and value scores.
The research question is whether a reusable embedding can do more of the work usually assigned to a separate reranking stage. On the paper's 15-task BEIR evaluation, adding the value read improves macro nDCG@10 from 48.47 to 51.54 over routing alone, on the same 100 candidates. This is evidence for the value component within this model; it is not a claim of state-of-the-art accuracy or a measured speedup over other systems.
Try it
Requires Python 3.12+. Install directly from GitHub; the package is not yet published on PyPI:
python -m pip install "git+https://github.com/BreadBowlAI/breadbowl-embed.git"
from breadbowl_embed import BreadBowl
model = BreadBowl.from_pretrained("dotproductx/breadbowl-embed-v1.2-preview")
passages = [
"Octopuses have three hearts.",
"The Pacific is Earth's largest ocean.",
"Plants use sunlight during photosynthesis.",
]
query = "How many hearts does an octopus have?"
for hit in model.rerank(query, passages):
print(hit.score, passages[hit.document_index])
CPU is the default. On a supported Apple Silicon Mac, pass device="mps" to run the encoder on the Apple GPU. CPU and MPS have passed small inference tests; MPS was tested with native precision and float32, with automatic CPU fallback disabled. CUDA can be requested with device="cuda", but has not been validated for this preview. These are functionality checks, not performance benchmarks.
The high-level encoders return CPU float32 tensors, and the convenience reranking/index APIs score on CPU. Use the default dtype="native" to retain the checkpoint's mixed parameter precision during encoding. The toolkit pins its core dependencies to the validated versions.
Encode documents once and reuse them
Passing strings to rerank encodes those strings on each call. For repeated queries, cache the document representations:
documents = model.encode_documents(passages, batch_size=2)
# Only the new query needs a backbone pass here.
hits = model.rerank(query, documents)
# Small-corpus retrieval, followed by value reranking.
index = model.index(documents, ids=["octopus", "ocean", "plants"])
queries = model.encode_queries([query])
results = index.search(queries, top_k=2, candidate_k=3)[0]
for hit in results:
print(hit.id, hit.routing_score, hit.value_score, hit.score)
The included Index is deliberately simple: it scans all document routing slots in batches, keeps the best candidate_k documents by routing score, then reranks those candidates using the combined score. It uses PyTorch and Python; no FAISS, vector database, or service is required. All document slots stay in memory; the index does not store the original passage text. Use it for small experiments; it is not an approximate-nearest-neighbor index or the paper's per-slot shortlist retrieval pipeline.
Every query scans the whole corpus, so routing retrieval cost grows linearly with the number of documents. The index returns the best top_k results after reranking and requires candidate_k >= top_k >= 1. Set candidate_k to the corpus size to rerank every document. A smaller candidate set can change the results. index.save(...) and Index.load(...) persist the slots, IDs, and optional metadata; keep your own ID-to-text mapping. The two float32 slot tensors occupy 32 KiB per passage, before metadata and runtime overhead.
Access the slots and scores directly
You can use the model's tensors in your own retrieval code:
print(documents.routing.shape) # [3, 16, 256]
print(documents.value.shape) # [3, 16, 256]
scores = model.score(queries, documents)
print(scores.routing) # [number of queries, number of documents]
print(scores.value)
print(scores.combined)
documents.routing contains L2-normalized routing vectors. documents.value contains value vectors without normalization. Both are ordinary PyTorch tensors. Preserve document-value magnitudes before attention aggregation: normalizing each document value first changes the scorer. Representations can be saved with documents.save(...) and loaded with Embeddings.load(...).
For research access to the underlying model, model.network is the actual PyTorch module. A direct forward pass exposes unnormalized routing vectors, values, slot-extraction attention, hidden slot states, and learned slot queries:
import torch
prep = model.config["preprocessing"]
batch = model.tokenizer(
[prep["document_format"].format(text=passages[0])],
padding=True,
truncation=True,
max_length=prep["max_document_length"],
return_tensors="pt",
)
with torch.inference_mode():
raw = model.network(
input_ids=batch["input_ids"].to(model.device),
attention_mask=batch["attention_mask"].to(model.device),
)
raw.routing # unnormalized routing-head outputs
raw.value # value-head outputs
raw.attention # slot-extraction attention over input tokens
raw.slot_states # hidden slot representations
raw.slot_queries # learned slot queries
The direct forward pass leaves tensors on the model device and bypasses the high-level encoding helpers. For query inputs, use the checkpoint's query format and instruction, or call encode_queries to apply them automatically. The attention shown here is the encoder's slot-extraction attention, not query-to-document attention from reranking.
Architecture and scoring
| Component | This checkpoint |
|---|---|
| Backbone | Qwen3.5-0.8B-Base; nominal size excludes pooling and output heads |
| Representation | 16 routing鈥搗alue slots per input |
| Slot dimensions | 256 routing + 256 value |
| Input limits | 128 query tokens, including instruction; 384 document tokens |
| Default query instruction | Retrieve relevant documents. |
| Routing score | Log-mean-exp aggregation, temperature 0.02 |
| Value score | Query-dependent attention read over document values, temperature 0.1 |
| Final score | 0.75 脳 routing + 0.25 脳 value |
This is a custom multi-vector architecture. A single-vector cosine search does not reproduce its scorer. Inputs exceeding the token limits are truncated; split long documents into passages before encoding. Scores are ranking signals, not probabilities.
Evaluation
On 15 BEIR tasks, using the same 100 routing-selected candidates:
| Scoring | BEIR macro nDCG@10 | Recall@100 |
|---|---|---|
| Routing only (QK) | 48.47 | 67.91 |
| Routing + values (QKV) | 51.54 | 67.91 |
Metrics are multiplied by 100. Ten tasks improve and five decline. Recall is identical because both scorers rank the same candidate set. This is an inference component ablation of one trained checkpoint, not a comparison between separately trained architectures. See the paper for the complete task breakdown and external-model comparisons.
The inference export was compared with the research checkpoint on two queries and four documents: representations matched exactly on CPU, and scores agreed within approximately 1.2e-7. See validation.json. Anonymous download and CPU inference from a fresh Hub cache were also verified. These checks establish basic numerical fidelity and usability, not a new BEIR evaluation.
What is open
The public repository includes the architecture, model loader, routing/value scoring, exact document-routing index, representation persistence, checkpoint exporter, tests, and examples. You can inspect and modify the model, access its outputs, or integrate the slots into another retrieval system without a BreadBowl API account.
The current release does not yet include the complete training/data pipeline or the paper's full evaluation and candidate-retrieval pipeline. The hosted service is separate. The toolkit's direct runtime dependencies are PyTorch, Transformers, Hugging Face Hub, and Safetensors; FAISS is not a dependency.
Limitations
- This is research work; no state-of-the-art or general benchmark-parity claim is made.
- End-to-end latency, throughput, and cost advantages over a retriever plus cross-encoder have not been measured. Indexing, storage, and data transfer also contribute to system cost.
- Some BEIR domains occur in adaptation data, and earlier SciFact and NFCorpus evaluations informed development. Weak-training overlap has not been fully audited; the suite is not uniformly zero-shot or wholly untouched.
- BEIR results are point estimates. Development results and their uncertainty are described separately in the paper.
- Cross-version embedding compatibility and domain extension packs are not established by this release.
Author
Ming Xu, Breadbowl AI.
This manuscript is a preprint. It does not claim conference acceptance.
- Downloads last month
- 27
Model tree for dotproductx/breadbowl-embed-v1.2-preview
Base model
Qwen/Qwen3.5-0.8B-Base