Instructions to use perplexity-ai/pplx-embed-v2-late-9b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use perplexity-ai/pplx-embed-v2-late-9b with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("perplexity-ai/pplx-embed-v2-late-9b") sentences = [ "The weather is lovely today.", "It's so sunny outside!", "He drove to the stadium." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [3, 3] - Notebooks
- Google Colab
- Kaggle
pplx-embed-v2-late: Multimodal Late-Interaction Embeddings
pplx-embed-v2-late is a family of multimodal late-interaction (ColBERT) retrievers for text, images, and visual documents, built on Qwen3.5 with bidirectional attention. The models produce one 128-dimensional vector per token and score query–document similarity using MaxSim. The 0.6B and 9B models share an embedding space, allowing the 0.6B model to query an index built with the 9B model.
For benchmarks and details, see our blog post.
Models
| Model | Active Params | Dimension | Public ViDoRe(v3), Image, nDCG@10 | Public ViDoRe(v3), Markdown, nDCG@10 |
|---|---|---|---|---|
pplx-embed-v2-late-0.6b |
340M | 128 | 62.3% | 61.2% |
pplx-embed-v2-late-9b |
7.4B | 128 | 65.2% | 64.7% |
Training
Both models were distilled from an internal 18B ColBERT teacher trained on pair and triplet data. Distillation used a token-level LEAF-style objective. The 0.6B model was fully fine-tuned. For the 9B model, the final eight transformer layers were fully fine-tuned, while the remaining transformer layers and vision encoder were adapted with LoRA.
Usage (Sentence Transformers)
Requires sentence-transformers >= 6.0.0 and transformers >= 5.4.0.
pip install 'sentence-transformers>=6.0.0' 'transformers>=5.4.0'
from PIL import Image
from sentence_transformers import MultiVectorEncoder
model = MultiVectorEncoder(
"perplexity-ai/pplx-embed-v2-late-9b", # or "perplexity-ai/pplx-embed-v2-late-0.6b"
device="cuda",
)
queries = model.encode_query(["what statute governs limitations?"])
documents = model.encode_document(["A document passage"])
scores = model.similarity(queries, documents) # MaxSim
image = Image.open("document.png").convert("RGB")
image_documents = model.encode_document([image])
image_scores = model.similarity(queries, image_documents)
The export uses native Sentence Transformers modules and needs no custom Python code. Use separate encoding calls for text-only and image-only batches. Mixed text+image inputs are not supported.
PyLate inserts Q/D markers at the second position; this model expects them first.
- Downloads last month
- 240