Papers
arxiv:2506.18902

jina-embeddings-v4: Universal Embeddings for Multimodal Multilingual Retrieval

Published on Jun 23, 2025
Authors:
,
,
,
,
,
,

Abstract

jina-embeddings-v4, a multimodal embedding model with LoRA adapters, achieves state-of-the-art performance in single-modal and cross-modal retrieval, especially for visually rich content, supported by Jina-VDR benchmark.

We introduce jina-embeddings-v4, a 3.8 billion parameter multimodal embedding model that unifies text and image representations through a novel architecture supporting both single-vector and multi-vector embeddings in the late interaction style. The model incorporates task-specific Low-Rank Adaptation (LoRA) adapters to optimize performance across diverse retrieval scenarios, including query-based information retrieval, cross-modal semantic similarity, and programming code search. Comprehensive evaluations demonstrate that jina-embeddings-v4 achieves state-of-the-art performance on both single- modal and cross-modal retrieval tasks, with particular strength in processing visually rich content such as tables, charts, diagrams, and mixed-media formats. To facilitate evaluation of this capability, we also introduce Jina-VDR, a novel benchmark specifically designed for visually rich image retrieval.

Community

Very nice model!

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2506.18902
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 14

Browse 14 models citing this paper

Datasets citing this paper 50

Browse 50 datasets citing this paper

Spaces citing this paper 18

Browse 18 spaces citing this paper

Collections including this paper 2