Perplexity Logo

pplx-embed-v2-late: Multimodal Late-Interaction Embeddings

pplx-embed-v2-late is a family of multimodal late-interaction (ColBERT) retrievers for text, images, and visual documents, built on Qwen3.5 with bidirectional attention. The models produce one 128-dimensional vector per token and score query–document similarity using MaxSim. The 0.6B and 9B models share an embedding space, allowing the 0.6B model to query an index built with the 9B model.

For benchmarks and details, see our blog post.

Models

Model Active Params Dimension Public ViDoRe(v3), Image, nDCG@10 Public ViDoRe(v3), Markdown, nDCG@10
pplx-embed-v2-late-0.6b 340M 128 62.3% 61.2%
pplx-embed-v2-late-9b 7.4B 128 65.2% 64.7%

Training

Both models were distilled from an internal 18B ColBERT teacher trained on pair and triplet data. Distillation used a token-level LEAF-style objective. The 0.6B model was fully fine-tuned. For the 9B model, the final eight transformer layers were fully fine-tuned, while the remaining transformer layers and vision encoder were adapted with LoRA.

Usage (Sentence Transformers)

Requires sentence-transformers >= 6.0.0 and transformers >= 5.4.0.

pip install 'sentence-transformers>=6.0.0' 'transformers>=5.4.0'
from PIL import Image
from sentence_transformers import MultiVectorEncoder

model = MultiVectorEncoder(
    "perplexity-ai/pplx-embed-v2-late-9b",  # or "perplexity-ai/pplx-embed-v2-late-0.6b"
    device="cuda",
)
queries = model.encode_query(["what statute governs limitations?"])
documents = model.encode_document(["A document passage"])
scores = model.similarity(queries, documents)  # MaxSim

image = Image.open("document.png").convert("RGB")
image_documents = model.encode_document([image])
image_scores = model.similarity(queries, image_documents)

The export uses native Sentence Transformers modules and needs no custom Python code. Use separate encoding calls for text-only and image-only batches. Mixed text+image inputs are not supported.

PyLate inserts Q/D markers at the second position; this model expects them first.

Downloads last month
240
Safetensors
Model size
8B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including perplexity-ai/pplx-embed-v2-late-9b