Image-Text-to-Text
Transformers
Safetensors
English
Chinese
groundinganything
text-generation
visual-grounding
object-detection
referring-expression-comprehension
pointing
ocr
document-layout
custom-code
diffusion-language-model
conversational
custom_code
Instructions to use GroundingPI/GroundAnything with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use GroundingPI/GroundAnything with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="GroundingPI/GroundAnything", trust_remote_code=True) messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("GroundingPI/GroundAnything", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use GroundingPI/GroundAnything with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "GroundingPI/GroundAnything" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "GroundingPI/GroundAnything", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/GroundingPI/GroundAnything
- SGLang
How to use GroundingPI/GroundAnything with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "GroundingPI/GroundAnything" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "GroundingPI/GroundAnything", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "GroundingPI/GroundAnything" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "GroundingPI/GroundAnything", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use GroundingPI/GroundAnything with Docker Model Runner:
docker model run hf.co/GroundingPI/GroundAnything
Update main demo video and prepared cover
#4
by Skywalker0410 - opened
- README.md +11 -15
- assets/demo-poster.jpg +2 -2
- assets/demo.mp4 +2 -2
- checksums.sha256 +3 -1
README.md
CHANGED
|
@@ -19,20 +19,20 @@ tags:
|
|
| 19 |
inference: false
|
| 20 |
---
|
| 21 |
|
| 22 |
-
# <img src="https://huggingface.co/GroundingPI/GroundAnything/resolve/
|
| 23 |
|
| 24 |
**Model family:** [GroundAnything β DLM / parallel decoding](https://huggingface.co/GroundingPI/GroundAnything) Β· [GroundAnything-VLM β autoregressive](https://huggingface.co/GroundingPI/GroundAnything-VLM).
|
| 25 |
|
| 26 |
**This repository contains the GroundAnything DLM checkpoint.** Use entropy-guided decoding for the main benchmark setting, or select optional self-speculative decoding.
|
| 27 |
|
| 28 |
-
<p align="center"><img src="https://huggingface.co/GroundingPI/GroundAnything/resolve/
|
| 29 |
|
| 30 |
## π Quick Links
|
| 31 |
|
| 32 |
-
- π **Online Demo:**
|
|
|
|
| 33 |
- π» **GitHub Code:** [groundingpi/GroundAnything](https://github.com/groundingpi/GroundAnything).
|
| 34 |
- π **Paper:** [arXiv:2609.39600](https://arxiv.org/abs/2609.39600).
|
| 35 |
-
- π§ͺ **Evaluation data:** Coming soon β XXX.
|
| 36 |
|
| 37 |
# Model Overview
|
| 38 |
|
|
@@ -46,11 +46,11 @@ An optional **self-speculative mode** achieves a **4.51Γ speedup** over the AR
|
|
| 46 |
|
| 47 |
### Demo Videos
|
| 48 |
|
| 49 |
-
<video controls playsinline preload="none" width="100%" poster="https://huggingface.co/GroundingPI/GroundAnything/resolve/
|
| 50 |
|
| 51 |
**Parallel Decoding**
|
| 52 |
|
| 53 |
-
<video controls playsinline preload="none" width="100%" src="https://huggingface.co/GroundingPI/GroundAnything/resolve/
|
| 54 |
|
| 55 |
### License/Terms of Use:
|
| 56 |
|
|
@@ -72,17 +72,13 @@ Global.
|
|
| 72 |
|
| 73 |
- **Paper [09/30/2026]:** [GroundAnything](https://arxiv.org/abs/2609.39600).
|
| 74 |
|
| 75 |
-
## References(s):
|
| 76 |
-
|
| 77 |
-
- [GroundAnything paper and supplementary material](https://arxiv.org/abs/2609.39600).
|
| 78 |
-
- [GroundingPI: grounding with visual primitives](https://arxiv.org/abs/2609.39601).
|
| 79 |
|
| 80 |
<details>
|
| 81 |
<summary>Citation</summary>
|
| 82 |
|
| 83 |
```bibtex
|
| 84 |
-
@misc{
|
| 85 |
-
title = {GroundAnything: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed},
|
| 86 |
author = {Qize Yu and Lianrui Fan and Bowen Ping and Xini Ding and Zetian Song and Junbo Niu and Kaixuan Wang and Tianxing Chen and Yue Chen and Minghua He and Yuran Wang and Jie Huang and Haojun Zhang and Min Chen and Hao Li and Wenxuan Song and Ruihai Wu and Xianming Liu and Shilong Liu and Shuchang Zhou and Ping Luo and Shiyu Huang},
|
| 87 |
year = {2026},
|
| 88 |
eprint = {2609.39600},
|
|
@@ -104,7 +100,7 @@ Global.
|
|
| 104 |
- **Spatial vocabulary:** 1,000 coordinate tokens shared with semantic labels and protocol markers.
|
| 105 |
- **DLM conversion:** the shared decoder and vocabulary head support both causal prediction and bidirectional response-block denoising. A mask token is added for diffusion generation.
|
| 106 |
|
| 107 |
-
<p align="center"><img src="https://huggingface.co/GroundingPI/GroundAnything/resolve/
|
| 108 |
|
| 109 |
## Input(s):
|
| 110 |
|
|
@@ -173,7 +169,7 @@ Evaluation code and instructions: [GitHub](https://github.com/groundingpi/Ground
|
|
| 173 |
|
| 174 |
## Quantitative Evaluation Benchmarks
|
| 175 |
|
| 176 |
-
<p align="center"><img src="https://huggingface.co/GroundingPI/GroundAnything/resolve/
|
| 177 |
|
| 178 |
## Inference:
|
| 179 |
|
|
@@ -278,7 +274,7 @@ After a block is complete, a causal forward reconstructs its authoritative KV ca
|
|
| 278 |
|
| 279 |
The model uses its own shared weights to draft tokens with bidirectional attention and verify them with causal attention. Verification accepts the **longest consecutive matching prefix**, stops at the first mismatch, applies the causal correction, and discards the rejected suffix cache states. The shipped speculative route uses greedy verification; it is not a general stochastic speculative sampler.
|
| 280 |
|
| 281 |
-
<p align="center"><img src="https://huggingface.co/GroundingPI/GroundAnything/resolve/
|
| 282 |
|
| 283 |
The documented `--decoder speculative` service is the linear shared-weight route. Exact greedy verification is relative to the converted model's causal branch; it does not imply identical outputs to the separately trained GroundAnything-VLM checkpoint.
|
| 284 |
|
|
|
|
| 19 |
inference: false
|
| 20 |
---
|
| 21 |
|
| 22 |
+
# <img src="https://huggingface.co/GroundingPI/GroundAnything/resolve/20fda7044a5850df0c773831bd3b94ce9cf1b127/assets/logo.png" width="40" style="display: inline-block; vertical-align: middle; margin: 0;" alt="GroundAnything logo" /> GroundAnything: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed
|
| 23 |
|
| 24 |
**Model family:** [GroundAnything β DLM / parallel decoding](https://huggingface.co/GroundingPI/GroundAnything) Β· [GroundAnything-VLM β autoregressive](https://huggingface.co/GroundingPI/GroundAnything-VLM).
|
| 25 |
|
| 26 |
**This repository contains the GroundAnything DLM checkpoint.** Use entropy-guided decoding for the main benchmark setting, or select optional self-speculative decoding.
|
| 27 |
|
| 28 |
+
<p align="center"><img src="https://huggingface.co/GroundingPI/GroundAnything/resolve/20fda7044a5850df0c773831bd3b94ce9cf1b127/assets/fig1-teaser.png" width="100%" alt="GroundAnything: broad visual grounding and parallel visual evidence extraction" /></p>
|
| 29 |
|
| 30 |
## π Quick Links
|
| 31 |
|
| 32 |
+
- π **Online Demo:** The GroundAnything DLM runtime is not supported on ZeroGPU.
|
| 33 |
+
- π **Project Page:** [GroundAnything](https://groundingpi.github.io/groundanything/).
|
| 34 |
- π» **GitHub Code:** [groundingpi/GroundAnything](https://github.com/groundingpi/GroundAnything).
|
| 35 |
- π **Paper:** [arXiv:2609.39600](https://arxiv.org/abs/2609.39600).
|
|
|
|
| 36 |
|
| 37 |
# Model Overview
|
| 38 |
|
|
|
|
| 46 |
|
| 47 |
### Demo Videos
|
| 48 |
|
| 49 |
+
<video controls playsinline preload="none" width="100%" poster="https://huggingface.co/GroundingPI/GroundAnything/resolve/20fda7044a5850df0c773831bd3b94ce9cf1b127/assets/demo-poster.jpg" src="https://huggingface.co/GroundingPI/GroundAnything/resolve/20fda7044a5850df0c773831bd3b94ce9cf1b127/assets/demo.mp4"></video>
|
| 50 |
|
| 51 |
**Parallel Decoding**
|
| 52 |
|
| 53 |
+
<video controls playsinline preload="none" width="100%" src="https://huggingface.co/GroundingPI/GroundAnything/resolve/20fda7044a5850df0c773831bd3b94ce9cf1b127/assets/decoding.mp4"></video>
|
| 54 |
|
| 55 |
### License/Terms of Use:
|
| 56 |
|
|
|
|
| 72 |
|
| 73 |
- **Paper [09/30/2026]:** [GroundAnything](https://arxiv.org/abs/2609.39600).
|
| 74 |
|
|
|
|
|
|
|
|
|
|
|
|
|
| 75 |
|
| 76 |
<details>
|
| 77 |
<summary>Citation</summary>
|
| 78 |
|
| 79 |
```bibtex
|
| 80 |
+
@misc{yu2026groundanythingreconcilingparalleldecoding,
|
| 81 |
+
title = {{GroundAnything}: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed},
|
| 82 |
author = {Qize Yu and Lianrui Fan and Bowen Ping and Xini Ding and Zetian Song and Junbo Niu and Kaixuan Wang and Tianxing Chen and Yue Chen and Minghua He and Yuran Wang and Jie Huang and Haojun Zhang and Min Chen and Hao Li and Wenxuan Song and Ruihai Wu and Xianming Liu and Shilong Liu and Shuchang Zhou and Ping Luo and Shiyu Huang},
|
| 83 |
year = {2026},
|
| 84 |
eprint = {2609.39600},
|
|
|
|
| 100 |
- **Spatial vocabulary:** 1,000 coordinate tokens shared with semantic labels and protocol markers.
|
| 101 |
- **DLM conversion:** the shared decoder and vocabulary head support both causal prediction and bidirectional response-block denoising. A mask token is added for diffusion generation.
|
| 102 |
|
| 103 |
+
<p align="center"><img src="https://huggingface.co/GroundingPI/GroundAnything/resolve/20fda7044a5850df0c773831bd3b94ce9cf1b127/assets/fig2-architecture.png" width="100%" alt="GroundAnything: vision-language architecture and autoregressive-to-diffusion conversion" /></p>
|
| 104 |
|
| 105 |
## Input(s):
|
| 106 |
|
|
|
|
| 169 |
|
| 170 |
## Quantitative Evaluation Benchmarks
|
| 171 |
|
| 172 |
+
<p align="center"><img src="https://huggingface.co/GroundingPI/GroundAnything/resolve/20fda7044a5850df0c773831bd3b94ce9cf1b127/assets/fig7-grounding-performance.png" width="100%" alt="GroundAnything: GroundAnything and GroundAnything-VLM benchmark overview" /></p>
|
| 173 |
|
| 174 |
## Inference:
|
| 175 |
|
|
|
|
| 274 |
|
| 275 |
The model uses its own shared weights to draft tokens with bidirectional attention and verify them with causal attention. Verification accepts the **longest consecutive matching prefix**, stops at the first mismatch, applies the causal correction, and discards the rejected suffix cache states. The shipped speculative route uses greedy verification; it is not a general stochastic speculative sampler.
|
| 276 |
|
| 277 |
+
<p align="center"><img src="https://huggingface.co/GroundingPI/GroundAnything/resolve/20fda7044a5850df0c773831bd3b94ce9cf1b127/assets/fig6-self-speculative-decoding.png" width="100%" alt="GroundAnything: linear and quadratic self-speculative schedules with shared model weights" /></p>
|
| 278 |
|
| 279 |
The documented `--decoder speculative` service is the linear shared-weight route. Exact greedy verification is relative to the converted model's causal branch; it does not imply identical outputs to the separately trained GroundAnything-VLM checkpoint.
|
| 280 |
|
assets/demo-poster.jpg
CHANGED
|
Git LFS Details
|
|
Git LFS Details
|
assets/demo.mp4
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
-
size
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:cddaec9f528f9ebf5ac627850eefcb3f14ee55e168776a1b650f91353ce8293e
|
| 3 |
+
size 45075106
|
checksums.sha256
CHANGED
|
@@ -1,5 +1,5 @@
|
|
| 1 |
c419c191a33774f4e1f834b3fb36033dc84e8ba316a3d1935eed1fa529884a74 LICENSE
|
| 2 |
-
|
| 3 |
669bb095cca2e86ddc821926f4c3b42dd7389f2ad3ec49c8edca8fe6294aae7f added_tokens.json
|
| 4 |
a0bc6f6fc7a29a80017a433e8f03a1cc1236e838a944a2d034295a60c4f2fddb chat_template.jinja
|
| 5 |
22c369e2bdac19723fe7db7e4c64b224462744b1f1aae7bd5daf05307db3b9fd config.json
|
|
@@ -22,3 +22,5 @@ d948513463b339ae85d6be6c99ecee787cc9c7a96be09180aa521a32a216eff4 special_tokens
|
|
| 22 |
ca10d7e9fb3ed18575dd1e277a2579c16d108e32f27439684afa0e10b1440910 vocab.json
|
| 23 |
cfc7749b96f63bd31c3c42b5c471bf756814053e847c10f3eb003417bc523d30 LICENSE-Apache-2.0
|
| 24 |
20c797ce19af0c17de52c6afb144644768a591c521655f5ebf5712c9850f2887 LICENSE-Kimi-K3
|
|
|
|
|
|
|
|
|
| 1 |
c419c191a33774f4e1f834b3fb36033dc84e8ba316a3d1935eed1fa529884a74 LICENSE
|
| 2 |
+
331eeebd16f922886141ce4562f3b2504cddb0cb8f0adc7350bc93e72b207638 README.md
|
| 3 |
669bb095cca2e86ddc821926f4c3b42dd7389f2ad3ec49c8edca8fe6294aae7f added_tokens.json
|
| 4 |
a0bc6f6fc7a29a80017a433e8f03a1cc1236e838a944a2d034295a60c4f2fddb chat_template.jinja
|
| 5 |
22c369e2bdac19723fe7db7e4c64b224462744b1f1aae7bd5daf05307db3b9fd config.json
|
|
|
|
| 22 |
ca10d7e9fb3ed18575dd1e277a2579c16d108e32f27439684afa0e10b1440910 vocab.json
|
| 23 |
cfc7749b96f63bd31c3c42b5c471bf756814053e847c10f3eb003417bc523d30 LICENSE-Apache-2.0
|
| 24 |
20c797ce19af0c17de52c6afb144644768a591c521655f5ebf5712c9850f2887 LICENSE-Kimi-K3
|
| 25 |
+
cddaec9f528f9ebf5ac627850eefcb3f14ee55e168776a1b650f91353ce8293e assets/demo.mp4
|
| 26 |
+
3c282238bb1598093c078e57dcc1db4b3ce8fb1f39db6c09ca90e0686f1043a2 assets/demo-poster.jpg
|