Instructions to use perplexity-ai/pplx-embed-v2-context-9b-preview with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use perplexity-ai/pplx-embed-v2-context-9b-preview with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="perplexity-ai/pplx-embed-v2-context-9b-preview", trust_remote_code=True)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("perplexity-ai/pplx-embed-v2-context-9b-preview", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
pplx-embed-v2-context-9b-preview
pplx-embed-v2-context-9b-preview is a contextual embedding model for document chunks in RAG systems. A document is passed as a list of chunks; the chunks are encoded together, so each chunk's embedding reflects its surrounding context, and one embedding is returned per chunk.
This is a preview release, not a final model. Weights, embeddings, and the interface may change in later versions without backward compatibility, so embeddings produced with this preview should not be mixed with embeddings from a future release.
Queries and documents are encoded with different methods: use
encode_queriesfor queries andencodefor document chunks. The model is trained with separate query and document prefixes, and encoding queries withencodesilently degrades retrieval quality.
Like
pplx-embed-context-v1, the model natively produces unnormalized int8-quantized embeddings. Compare embeddings with cosine similarity, or passnormalize_embeddings=Trueand use the dot product.
Model
| Model | Dimensions | MRL | Quantization | Instruction | Pooling |
|---|---|---|---|---|---|
pplx-embed-v2-context-9b-preview |
2048 | 1024, 2048 | INT8 | No (fixed query/document prefixes) | Mean |
Usage
Requires transformers>=5.4.0, torch, numpy, safetensors, and tqdm. The model uses custom code, so load it with trust_remote_code=True.
from transformers import AutoModel
model = AutoModel.from_pretrained(
"perplexity-ai/pplx-embed-v2-context-9b-preview",
trust_remote_code=True,
).to("cuda")
doc_chunks = [
[
"Curiosity begins in childhood with endless questions about the world.",
"As we grow, curiosity drives us to explore new ideas.",
"Scientific breakthroughs often start with a curious question.",
],
[
"The curiosity rover explores Mars searching for ancient life.",
"Each discovery on Mars sparks new questions about the universe.",
],
]
# One (chunk_count, 2048) array per document:
# doc_embeddings[0].shape == (3, 2048), doc_embeddings[1].shape == (2, 2048)
doc_embeddings = model.encode(doc_chunks, normalize_embeddings=True)
# Each query is a single-chunk row.
queries = [["What drives scientific breakthroughs?"]]
query_embeddings = model.encode_queries(queries, normalize_embeddings=True)
scores = doc_embeddings[0] @ query_embeddings[0][0]
Options
encode(documents, ...) and encode_queries(queries, ...) accept:
| Argument | Default | Description |
|---|---|---|
batch_size |
32 |
Documents (or queries) per forward pass |
normalize_embeddings |
False |
L2-normalize outputs |
convert_to_numpy |
True |
Return NumPy arrays; False returns CPU tensors |
show_progress_bar |
False |
Show a progress bar |
device |
None |
Move the model to this device before encoding |
Matryoshka dimensions
The model was trained with Matryoshka losses at 1024 and 2048 dimensions. To use 1024-dimensional embeddings, take the first 1024 values of each unnormalized embedding and normalize afterwards:
import numpy as np
emb = model.encode(doc_chunks) # unnormalized int8 values
emb_1024 = [e[:, :1024] / np.linalg.norm(e[:, :1024], axis=-1, keepdims=True) for e in emb]
Other truncation sizes were not trained.
- Downloads last month
- 305