Papers
arxiv:2609.32699

RandSlot: Learning Compact Visual Document Representations with Random Soft Tokens

Published on Sep 26
Authors:
,
,
,
,
,

Abstract

Visual document retrieval requires expressive representations to match queries with evidence distributed across text, tables, and page layouts. Multi-vector representations capture fine-grained information, but storing and comparing many vectors introduces substantial retrieval costs. In this paper, we introduce RandSlot, a simple approach to learn compact visual document representations with random soft tokens. During training, we append independently-sampled random unit vectors to query and document input sequences and resample them at every use, without introducing learnable soft-token parameters. The encoder contextualizes these auxiliary inputs with the original content to produce a small set of retrieval vectors. A standard late-interaction objective trains the encoder to extract relevant information under varying input conditions. Experiments with different backbone models show that RandSlot improves retrieval quality over alternative readout strategies under the same vector budget. Further analysis shows that these gains can persist when random soft tokens are replaced with zeros at inference, demonstrating that random inputs during training can improve compact retrieval representations even when inference no longer requires sampling.

Community

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.32699
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.32699 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.32699 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.32699 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.