mirianokradze's picture
Desearch Embedding 4B
e82acef verified
|
Raw History Blame Contribute Delete
5.19 kB
---
license: apache-2.0
base_model: Qwen/Qwen3-Embedding-4B
library_name: sentence-transformers
pipeline_tag: sentence-similarity
language:
- en
tags:
- sentence-transformers
- text-embeddings
- retrieval
- web-search
- news
- peft
---
# Desearch Embedding 4B
Desearch Embedding 4B is a text embedding model for web and news search, developed by
[Desearch](https://desearch.ai). It is fine-tuned from [Qwen/Qwen3-Embedding-4B](https://huggingface.co/Qwen/Qwen3-Embedding-4B)
and maps search queries and web documents into a shared vector space for first-stage retrieval.
## Highlights
- **Built for news and the open web.** Fine-tuned on recent news coverage, reference pages and
encyclopedic articles, the content a web search engine has to rank every day.
- **Real queries in every form.** Short keyword searches, natural-language questions, noisy
queries, questions about dated news events, and multi-hop questions that combine facts from
linked pages.
- **One model for chunks and whole pages.** Documents are seen as paragraph chunks, page openings
and full pages, the units a search index stores.
- **Binary first-stage ready.** Trained with a binary objective on the leading 256 dimensions, for
search stacks that shortlist with compact binary vectors before rescoring with full vectors.
## Model details
| | |
|---|---|
| Base model | [Qwen/Qwen3-Embedding-4B](https://huggingface.co/Qwen/Qwen3-Embedding-4B) |
| Parameters | 4B |
| Embedding dimension | 2560 |
| Max sequence length | 32K tokens |
| Pooling | Last token, L2-normalized |
| Query instruction | Built-in `query` prompt |
| Language | English |
| Training method | LoRA fine-tuning |
| License | Apache 2.0 |
## Usage
The weights are a LoRA adapter on Qwen/Qwen3-Embedding-4B; the base model downloads automatically.
```bash
pip install -U sentence-transformers peft
```
### Using Sentence Transformers
```python
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("desearch/Desearch-Embedding-4B")
queries = [
"Who became Apple’s chief executive and will lead the Sept. 9, 2026 “Surprise and Shine” product event?",
"How much does the base Mac Mini M6 cost in US dollars as of its August 2026 announcement?",
]
documents = [
"John Ternus, Apple's new chief executive, takes the stage at the September 9, 2026 “Surprise and Shine” event to unveil the latest iPhones and Apple Watches.",
"Apple surprised buyers in late August 2026 with the Mac Mini M6 at $899 and the Mac Mini M5 Pro at $1,699, both shipping on September 22.",
]
query_embeddings = model.encode(queries, prompt_name="query")
document_embeddings = model.encode(documents)
print(model.similarity(query_embeddings, document_embeddings))
```
Queries use the built-in `query` prompt; documents are encoded as they are.
### Using Transformers
```python
import torch
import torch.nn.functional as F
from peft import PeftModel
from transformers import AutoModel, AutoTokenizer
task = "Given a web search query, retrieve relevant passages that answer the query"
questions = [
"Who became Apple’s chief executive and will lead the Sept. 9, 2026 “Surprise and Shine” product event?",
"How much does the base Mac Mini M6 cost in US dollars as of its August 2026 announcement?",
]
queries = [f"Instruct: {task}\nQuery:{q}" for q in questions]
documents = [
"John Ternus, Apple's new chief executive, takes the stage at the September 9, 2026 “Surprise and Shine” event to unveil the latest iPhones and Apple Watches.",
"Apple surprised buyers in late August 2026 with the Mac Mini M6 at $899 and the Mac Mini M5 Pro at $1,699, both shipping on September 22.",
]
tokenizer = AutoTokenizer.from_pretrained("desearch/Desearch-Embedding-4B", padding_side="left")
model = AutoModel.from_pretrained("Qwen/Qwen3-Embedding-4B", dtype=torch.bfloat16)
model = PeftModel.from_pretrained(model, "desearch/Desearch-Embedding-4B").eval()
def encode(texts):
batch = tokenizer(texts, padding=True, truncation=True, max_length=8192, return_tensors="pt")
with torch.no_grad():
hidden = model(**batch).last_hidden_state
return F.normalize(hidden[:, -1], p=2, dim=1)
print(encode(queries) @ encode(documents).T)
```
## Recommended use cases
- First-stage retrieval for web search, news search and retrieval-augmented generation
- Semantic search over news archives, documentation and reference content
- Candidate generation ahead of a reranker
## Limitations
- Tuned on English text; other languages are not a focus of this release.
- Queries need the `query` prompt; documents are encoded without one.
- Trained on documents up to 2,048 tokens; split longer pages into passages.
- Embeddings differ from Qwen/Qwen3-Embedding-4B; re-embed an existing corpus when switching.
## License
This model is licensed under the Apache License 2.0. It is derived from Qwen/Qwen3-Embedding-4B, which is also
licensed under the Apache License 2.0.
## Citation
```bibtex
@misc{desearch2026embedding,
title = {Desearch Embedding 4B},
author = {Desearch},
year = {2026},
url = {https://huggingface.co/desearch/Desearch-Embedding-4B}
}
```