VisCo: Leveraging Large Language Models as Intrinsic Encoders for Visual Token Compression
Abstract
Vision-language models (VLMs) process large numbers of visual tokens, resulting in substantial inference latency and memory overhead. This has motivated extensive research on visual token compression. While training-free strategies rely on heuristic metrics and suffer significant performance degradation under high compression ratios, many training-based methods introduce external compression modules that force the VLM backbone to adapt, incurring substantial retraining cost and compromising VLMs' priors. Effective visual token compression hinges on strong information encoding, a capability already present in pretrained VLMs but underutilized by existing approaches. Motivated by this, we propose VisCo, a training-efficient self-compression framework that reuses the pretrained VLM itself as an intrinsic compressor. VisCo is a parameter-sharing autoencoder that compresses visual information using a small set of memory tokens and transfers hierarchical information from encoding to decoding. Experiments show that VisCo surpasses prior methods across all evaluated compression ratios, with larger gains under more aggressive compression, and remains stable even in the extreme single-token setting. Moreover, when combined with the original visual tokens, the learned memory tokens can even improve the base model, suggesting that VisCo captures complementary representations beyond compression.
Community
Accepted by ACM MM 2026
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models (2026)
- Spectral Evolution-Guided Token Pruning in Multimodal Large Language Models (2026)
- ETC: Extreme Token Compression via Task-aware Visual Information Distillation in VLMs (2026)
- Accelerating Multimodal Large Language Models with Prior-Corrected Token Reduction (2026)
- EvoCut: Multi-Layer Evolution-Aware Visual Token Compression for Efficient Large Vision-Language Models (2026)
- CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference (2026)
- ERA: Entropy-Guided Visual Token Pruning with Rectified Attention for Efficient MLLMs (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper