--- license: mit datasets: - TIGER-Lab/MMEB-train language: - en base_model: - Qwen/Qwen3-VL-4B-Instruct library_name: transformers tags: - Retrieval - Multimodal - Embedding - Chain-of-Thought - Reinforcement-Learning pipeline_tag: image-text-to-text ---

UniME-R1-4B: Learning from Failures for Unified Multimodal Retrieval

Zelong Sun*, Jun Wang*, Kaicheng Yang, Tiancheng Gu, Ziyong Feng, Zhiwu Lu Glint Lab [![GitHub](https://img.shields.io/badge/⭐-GitHub-black?logo=github)](https://github.com/deepglint/UniME-R1) [![Paper](https://img.shields.io/badge/πŸ“„-Paper-b31b1b.svg)](https://arxiv.org/abs/2608.06060) [![Model](https://img.shields.io/badge/πŸ€—-UniME--R1_Models-yellow)](https://huggingface.co/DeepGlint-AI)
UniME-R1 is an **Embedder–Adviser** framework that learns to reason over *retrieved candidates* (not the query alone) and generate **Retrieval-Centric Chain-of-Thought (RC-CoT)** to correct retrieval failures. This repository ships the **4B-scale** pair: a Qwen3-VL-4B dual-mode Embedder and the Qwen3-VL-4B retrieval-aware Adviser β€” the best-performing configuration in the paper. ## πŸ’‘ Highlights - **Retrieval-Centric CoT (RC-CoT)** β€” The Adviser analyzes the *actual* top-k retrieved candidates to diagnose model-specific confusion, then emits `` (which discriminative cues are missing) and `` (a concise corrected query) to redirect retrieval.
- **Dual-Mode Embedder** β€” A single VLM backbone produces two embeddings via special tokens: `` for fast discriminative retrieval and `` for RC-CoT-enhanced re-retrieval. Candidates are encoded **once** with `` and reused across both paths β€” no candidate-side CoT, no index rebuilding. - **Adaptive Rerank-or-Retrieve** β€” The Adviser predicts whether a match exists in the top-k set. If yes, it reranks; if not, it appends RC-CoT to the query and re-retrieves over the full corpus. - **Retrieval-Oriented GRPO** β€” The Adviser is optimized with a 4-component reward (format / NDCG rerank / CoT-embedding quality / judge decision) that calls a **frozen Embedder API** to score the Adviser's CoT against mined hard negatives, so the RL signal reflects real end-to-end retrieval quality.
## 🧱 Model Components This release contains two sub-directories: | Component | Backbone | Format | Notes | |-----------|----------|--------|-------| | **Embedder** (`embedder/`) | Qwen3-VL-4B-Instruct | LoRA (DoRA, r=16, Ξ±=64) + `new_token_embeddings.pt` | Adds `` / `` tokens; pooling at the special-token position | | **Adviser** (`adviser/`) | Qwen3-VL-4B-Instruct | Full merged weights (bf16) | Outputs 5 structured XML fields; served with vLLM | ``` β”œβ”€β”€ adviser/ # Qwen3-VL-4B Adviser (full weights, vLLM-ready) β”‚ β”œβ”€β”€ model-0000{1,2}-of-00002.safetensors β”‚ β”œβ”€β”€ config.json / generation_config.json β”‚ β”œβ”€β”€ preprocessor_config.json / video_preprocessor_config.json β”‚ └── tokenizer.json / chat_template.jinja └── embedder/ # Qwen3-VL-4B Embedder (LoRA adapter) β”œβ”€β”€ adapter_config.json β”œβ”€β”€ adapter_model.safetensors β”œβ”€β”€ new_token_embeddings.pt # trained / embeddings β”œβ”€β”€ config.json └── preprocessor_config.json / video_preprocessor_config.json ``` > The Embedder is released as a **PEFT/LoRA adapter** β€” load it on top of `Qwen/Qwen3-VL-4B-Instruct`. The two special tokens (``=151670, ``=151669) and their embeddings are provided via `added_tokens.json` + `new_token_embeddings.pt`. The Adviser always emits five structured fields: | Field | Role | |-------|------| | `` | Candidate-by-candidate relevance analysis | | `` | Predicted candidate ordering (JSON array, 0-based) | | `` | Path decision: best-match ID, or `-1` if none matches | | `` | Retrieval-failure diagnosis β€” which discriminative cues are confused | | `` | Concise refined query text for re-retrieval | ## πŸš€ Quick Start ```bash git clone https://github.com/deepglint/UniME-R1.git cd UniME-R1 ``` ```bash conda create -n vlm2vec python=3.10 -y conda activate vlm2vec # Install torch matching your CUDA first, e.g.: # pip install torch==2.5.1 torchvision --index-url https://download.pytorch.org/whl/cu121 # pip install flash-attn==2.7.3 --no-build-isolation pip install -r Embedder/requirements.txt ``` ### πŸ” Embedder-only evaluation (direct `` retrieval) ```bash cd Embedder bash shell/eval/eval.sh image "../models/UniME-R1-4B/embedder" ``` ### 🎯 Full Adviser evaluation (rerank + RC-CoT iterative retrieval) Serve the Adviser via vLLM, then run the unified evaluation: ```bash vllm serve models/UniME-R1-4B/adviser --tensor-parallel-size 8 --port 9000 cd RL/eval export EMBEDDER_CHECKPOINT="../../models/UniME-R1-4B/embedder" export ADVISER_MODEL="Qwen3-VL-4B-Instruct" export ADVISER_URL="http://127.0.0.1:9000/v1" bash eval.sh image # image | visdoc | video | uvrb | image_caption ``` ## πŸ“Š Results ### πŸ† MMEB-V2 UniME-R1 achieves the best overall performance at both model scales. Notably, the 2B model already outperforms all medium-size (4B–7B) baselines, indicating the gains stem from the framework rather than model scale alone.
### 🌈 Zero-shot General Retrieval
## πŸ–ŠοΈ Citation If you find this repository useful, please use the following BibTeX entry for citation. ```bibtex @misc{unimer1, title={Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval}, author={Zelong Sun and Jun Wang and Kaicheng Yang and Tiancheng Gu and Ziyong Feng and Zhiwu Lu}, year={2026}, eprint={2608.06060}, archivePrefix={arXiv}, primaryClass={cs.CV}, url={https://arxiv.org/abs/2608.06060}, } ```
⭐ Don't forget to star this repository if you find it helpful!