Text-to-Speech
Transformers
ONNX
Safetensors
multimodal
audio
video
conversational
full-duplex
function-calling
gander
Instructions to use lingyuxing/Gander with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use lingyuxing/Gander with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-to-speech", model="lingyuxing/Gander")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("lingyuxing/Gander", device_map="auto") - Notebooks
- Google Colab
- Kaggle
|
Download README.md from lingyuxing/Gander: direct link, hf CLI and curl.
- Browser
- Download file 6.24 kB
-
https://huggingface.co/lingyuxing/Gander/resolve/main/README.md
- Command line
-
hf download hf://lingyuxing/Gander/README.md
-
curl -L -o README.md https://huggingface.co/lingyuxing/Gander/resolve/main/README.md
6.24 kB
| license: apache-2.0 | |
| base_model: openbmb/MiniCPM-o-4_5 | |
| library_name: transformers | |
| tags: | |
| - multimodal | |
| - audio | |
| - video | |
| - conversational | |
| - full-duplex | |
| - function-calling | |
| - text-to-speech | |
| - gander | |
| <h1 align="center">Gander</h1> | |
| <p align="center"><strong>Streaming audio-visual interaction and agentic task orchestration</strong></p> | |
| <!-- Replace the remaining #resources targets with public URLs before publication. --> | |
| <p align="center"> | |
| <a href="https://github.com/Omni-Interaction-Gander/Omni-Interaction-Agent"><img src="https://img.shields.io/badge/GitHub-Repository-181717?style=for-the-badge&logo=github" alt="GitHub repository"></a> | |
| <a href="https://arxiv.org/abs/2609.08977"><img src="https://img.shields.io/badge/Paper-Arxiv-C62828?style=for-the-badge" alt="Paper"></a> | |
| <a href=" https://Omni-Interaction-Gander.github.io/Omni-Interaction-Agent"><img src="https://img.shields.io/badge/Demo-Project_Page-2563EB?style=for-the-badge" alt="Demo"></a> | |
| <a href="#resources"><img src="https://img.shields.io/badge/Dataset-Coming_Soon-059669?style=for-the-badge" alt="Dataset coming soon"></a> | |
| </p> | |
| Gander is an omni-interaction model designed for continuous spoken and visual interaction | |
| while long-running tasks execute asynchronously. It combines a realtime multimodal Thinker, | |
| a streaming speech Talker, and the Gander runtime for task lifecycle management. | |
| ## Model Overview | |
| Gander organizes interaction into causal one-second units. In each unit, newly available audio | |
| and visual observations precede the model decision. The Thinker predicts whether to listen, | |
| speak, interrupt, or invoke a task operation, and produces text or structured calls when | |
| required. The Talker conditions on Thinker representations to synthesize speech incrementally. | |
| The model supports: | |
| - continuous speech interaction, interruption handling, and conversational backchannels; | |
| - streaming image and video understanding with temporally grounded responses; | |
| - structured task delegation, follow-up instructions, progress interaction, and completion; | |
| - incremental speech generation synchronized with the Thinker output. | |
| ## Components | |
| | Component | Description | | |
| | --- | --- | | |
| | `thinker/` | Multimodal perception, interaction control, language generation, and task operations | | |
| | `talker/` | Streaming S3 speech-token generation and Token2wav waveform decoding | | |
| | `release_manifest.json` | Model provenance and the shared temporal-unit contract | | |
| ## Model Specification | |
| | Property | Value | | |
| | --- | --- | | |
| | Base model | `openbmb/MiniCPM-o-4_5` | | |
| | Input modalities | Text, audio, image, and video frames | | |
| | Temporal unit | 1 second | | |
| | Thinker output budget | Up to 8 lexical tokens per speaking unit | | |
| | Talker output budget | 50 S3 tokens per speaking unit | | |
| | Streaming speech chunk | 25 S3 tokens, approximately 0.5 seconds | | |
| | Audio input sample rate | 16 kHz | | |
| | Speech output sample rate | 24 kHz | | |
| | Model precision | BF16 | | |
| ## Usage | |
| ### Thinker | |
| ```python | |
| import torch | |
| from transformers import AutoModel, AutoProcessor, AutoTokenizer | |
| path = "/path/to/gander/thinker" | |
| tokenizer = AutoTokenizer.from_pretrained(path, trust_remote_code=True) | |
| processor = AutoProcessor.from_pretrained(path, trust_remote_code=True) | |
| model = AutoModel.from_pretrained( | |
| path, | |
| trust_remote_code=True, | |
| torch_dtype=torch.bfloat16, | |
| init_vision=True, | |
| init_audio=True, | |
| init_tts=False, | |
| ).eval() | |
| ``` | |
| ### Full Realtime Inference | |
| The complete audio-visual and speech experience uses the | |
| [Gander runtime](https://github.com/Omni-Interaction-Gander/Omni-Interaction-Agent). Start from | |
| its serving example and configure the base model together with both released components: | |
| ```yaml | |
| model: | |
| model_name_or_path: /path/to/MiniCPM-o-4_5 | |
| processor_name_or_path: /path/to/MiniCPM-o-4_5 | |
| init_vision: true | |
| init_audio: true | |
| init_tts: true | |
| token2wav_dir: /path/to/gander/talker/assets/token2wav | |
| duplex: | |
| checkpoint: /path/to/gander/thinker | |
| talker_checkpoint: /path/to/gander/talker | |
| detached_talker_device: cuda:1 | |
| ref_audio_path: /path/to/gander/talker/assets/ref_audio.wav | |
| speak_text_tokens_per_unit: 8 | |
| talker_speech_tokens_per_unit: 50 | |
| talker_emit_speech_tokens: 25 | |
| ``` | |
| ```bash | |
| git clone https://github.com/Omni-Interaction-Gander/Omni-Interaction-Agent.git | |
| cd Omni-Interaction-Agent | |
| cp gander_runtime/configs/serve.example.yaml gander_runtime/configs/serve.local.yaml | |
| # Apply the paths above to serve.local.yaml. | |
| ./scripts/serve.sh gander_runtime/configs/serve.local.yaml | |
| ``` | |
| This configuration loads the released Thinker and Talker as a matched Unit-8/50 pair and runs | |
| the Talker asynchronously on the second GPU. | |
| ## Training | |
| Both components are initialized from MiniCPM-o 4.5 and trained in two stages: | |
| | Stage | Trainable modules | Schedule | | |
| | --- | --- | --- | | |
| | Thinker | Language model and audio projection | 1 epoch, 8,407 optimization steps | | |
| | Talker | Semantic projection and speech decoder | 2 epochs, 3,246 optimization steps | | |
| Training uses a multimodal mixture spanning spoken dialogue, full-duplex interaction, | |
| streaming video understanding, and agent task lifecycles. The training representation follows | |
| the same causal one-second unit structure used during online inference. | |
| ## Intended Use | |
| Gander-Unit8 is intended for research on realtime multimodal assistants, full-duplex spoken | |
| interaction, streaming video-language modeling, and interactive agent systems. External task | |
| execution is coordinated by the Gander runtime and its configured worker provider. | |
| ## Limitations | |
| Model outputs may contain factual or perceptual errors, particularly under ambiguous audio, | |
| rapid visual changes, or long temporal dependencies. Applications should validate consequential | |
| external actions and follow deployment requirements appropriate to their domain. | |
| ## Resources | |
| | Resource | Link | | |
| | --- | --- | | |
| | GitHub | [Omni-Interaction-Agent](https://github.com/Omni-Interaction-Gander/Omni-Interaction-Agent) | | |
| | Paper | [Arxiv Link](https://arxiv.org/abs/2609.08977) | | |
| | Demo | [Project Page]( https://Omni-Interaction-Gander.github.io/Omni-Interaction-Agent)| | |
| | Dataset | Coming soon | | |
| ## License | |
| Gander-Unit8 is released under the Apache License 2.0 and is derived from | |
| `openbmb/MiniCPM-o-4_5`. | |