--- license: apache-2.0 tags: - multimodal - discrete-diffusion - vision-language - unified-model ---

Gestalt

Gestalt: Large Multimodal Interplay Model

Website Paper Code Model

Zequn Yang†  Yu Miao†  Haotian Ni†  Ziheng Chen†  Chengxiang Huang†
Dongzhan Zhou  Kai Chen  Qi Zhang  Ji-Rong Wen  Yake Wei‡  Di Hu‡,✉

† Equal contribution ‡ Team leader ✉ Corresponding author

> **Gestalt** is a new paradigm of large multimodal model built around **multimodal interplay**. Guided by a multimodal interplay pyramid — from modality-specific modeling, through cross-modal alignment, to multimodal synergy — Gestalt adopts a unified **discrete diffusion** framework with an interplay-partitioned architecture, where learnable interplay tokens mediate cross-modal exchange and integration. The name is inspired by Gestalt psychology: *the whole is greater than the sum of its parts.*

Gestalt teaser

## Model Description **Gestalt** is an interplay-centric large multimodal model built on a unified discrete diffusion framework. Through multimodal pretraining, continual pretraining, and interplay-oriented supervised fine-tuning,Gestalt supports both understanding and generation. | Property | Value | |:--|:--| | Architecture | GestaltModelLM (discrete diffusion with interplay tokens) | | Parameters | ~8B | | Hidden size | 4096 | | Layers | 32 | | Attention heads | 32 | | Vocab size | 142,848 | | Max sequence length | 4,096 | | Precision | bfloat16 | ## Download ```bash # Gestalt model huggingface-cli download GeWuLab/Gestalt --local-dir /path/to/gestalt # IBQ vision tokenizer (required for encoding/decoding images) huggingface-cli download TencentARC/IBQ-Tokenizer-16384 --local-dir /path/to/ibq ``` The IBQ tokenizer is from [TencentARC/IBQ-Tokenizer-16384](https://huggingface.co/TencentARC/IBQ-Tokenizer-16384), used to convert between images and discrete visual tokens. ## Training Pipeline | Stage | Data | Objective | |:--|:--:|:--| | **Multimodal Pretraining** | 70M | Joint masked prediction | | **Continual Pretraining** ← this checkpoint | 8M | Conditional masked prediction | | **Supervised Fine-Tuning** | 13.7M (+≈2.5M T2I) | Four interplay categories, three-phase curriculum | ## Citation If you find Gestalt useful for your research, please cite: ```bibtex @article{gestalt2026, title = {Gestalt: Large Multimodal Interplay Model}, author = {Yang, Zequn and Miao, Yu and Ni, Haotian and Chen, Ziheng and Huang, Chengxiang and Zhou, Dongzhan and Chen, Kai and Zhang, Qi and Wen, Ji-Rong and Wei, Yake and Hu, Di}, year = {2026}, url = {https://github.com/GeWu-Lab/Gestalt} } ``` ## License This project is released under the Apache 2.0 license. ## Author Contributions Zequn Yang, Yake Wei, and Di Hu drove the overall advancement of the project. Zequn Yang, Yu Miao, Haotian Ni, Ziheng Chen, and Chengxiang Huang contributed equally to this work. Zequn Yang conducted model pretraining and supervised fine-tuning. Zequn Yang, Yu Miao, and Haotian Ni developed the model architecture and conducted the core experiments. Yu Miao, Ziheng Chen, and Chengxiang Huang contributed to data processing and organization. Yu Miao conducted the evaluation of image generation capabilities, Haotian Ni conducted the multimodal understanding evaluation, and Ziheng Chen conducted the text-only evaluation. Dongzhan Zhou, Kai Chen, Qi Zhang, and Ji-Rong Wen contributed to discussions on the technical design and methodology. Yake Wei, Zequn Yang, Yu Miao, Haotian Ni, Ziheng Chen, and Di Hu contributed to writing and revising the manuscript. Di Hu initiated the project. Yake Wei and Di Hu supervised and advised the project.