Update main demo video and prepared cover

#4
by Skywalker0410 - opened
Files changed (4) hide show
  1. README.md +11 -15
  2. assets/demo-poster.jpg +2 -2
  3. assets/demo.mp4 +2 -2
  4. checksums.sha256 +3 -1
README.md CHANGED
@@ -19,20 +19,20 @@ tags:
19
  inference: false
20
  ---
21
 
22
- # <img src="https://huggingface.co/GroundingPI/GroundAnything/resolve/0f8e30894c3ca86378d01ae51ec69c217c78151b/assets/logo.png" width="40" style="display: inline-block; vertical-align: middle; margin: 0;" alt="GroundAnything logo" /> GroundAnything: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed
23
 
24
  **Model family:** [GroundAnything β€” DLM / parallel decoding](https://huggingface.co/GroundingPI/GroundAnything) Β· [GroundAnything-VLM β€” autoregressive](https://huggingface.co/GroundingPI/GroundAnything-VLM).
25
 
26
  **This repository contains the GroundAnything DLM checkpoint.** Use entropy-guided decoding for the main benchmark setting, or select optional self-speculative decoding.
27
 
28
- <p align="center"><img src="https://huggingface.co/GroundingPI/GroundAnything/resolve/0f8e30894c3ca86378d01ae51ec69c217c78151b/assets/fig1-teaser.png" width="100%" alt="GroundAnything: broad visual grounding and parallel visual evidence extraction" /></p>
29
 
30
  ## πŸ”— Quick Links
31
 
32
- - πŸš€ **Online Demo:** Coming soon β€” XXX.
 
33
  - πŸ’» **GitHub Code:** [groundingpi/GroundAnything](https://github.com/groundingpi/GroundAnything).
34
  - πŸ“„ **Paper:** [arXiv:2609.39600](https://arxiv.org/abs/2609.39600).
35
- - πŸ§ͺ **Evaluation data:** Coming soon β€” XXX.
36
 
37
  # Model Overview
38
 
@@ -46,11 +46,11 @@ An optional **self-speculative mode** achieves a **4.51Γ— speedup** over the AR
46
 
47
  ### Demo Videos
48
 
49
- <video controls playsinline preload="none" width="100%" poster="https://huggingface.co/GroundingPI/GroundAnything/resolve/0f8e30894c3ca86378d01ae51ec69c217c78151b/assets/demo-poster.jpg" src="https://huggingface.co/GroundingPI/GroundAnything/resolve/0f8e30894c3ca86378d01ae51ec69c217c78151b/assets/demo.mp4"></video>
50
 
51
  **Parallel Decoding**
52
 
53
- <video controls playsinline preload="none" width="100%" src="https://huggingface.co/GroundingPI/GroundAnything/resolve/0f8e30894c3ca86378d01ae51ec69c217c78151b/assets/decoding.mp4"></video>
54
 
55
  ### License/Terms of Use:
56
 
@@ -72,17 +72,13 @@ Global.
72
 
73
  - **Paper [09/30/2026]:** [GroundAnything](https://arxiv.org/abs/2609.39600).
74
 
75
- ## References(s):
76
-
77
- - [GroundAnything paper and supplementary material](https://arxiv.org/abs/2609.39600).
78
- - [GroundingPI: grounding with visual primitives](https://arxiv.org/abs/2609.39601).
79
 
80
  <details>
81
  <summary>Citation</summary>
82
 
83
  ```bibtex
84
- @misc{yu2026groundanything,
85
- title = {GroundAnything: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed},
86
  author = {Qize Yu and Lianrui Fan and Bowen Ping and Xini Ding and Zetian Song and Junbo Niu and Kaixuan Wang and Tianxing Chen and Yue Chen and Minghua He and Yuran Wang and Jie Huang and Haojun Zhang and Min Chen and Hao Li and Wenxuan Song and Ruihai Wu and Xianming Liu and Shilong Liu and Shuchang Zhou and Ping Luo and Shiyu Huang},
87
  year = {2026},
88
  eprint = {2609.39600},
@@ -104,7 +100,7 @@ Global.
104
  - **Spatial vocabulary:** 1,000 coordinate tokens shared with semantic labels and protocol markers.
105
  - **DLM conversion:** the shared decoder and vocabulary head support both causal prediction and bidirectional response-block denoising. A mask token is added for diffusion generation.
106
 
107
- <p align="center"><img src="https://huggingface.co/GroundingPI/GroundAnything/resolve/0f8e30894c3ca86378d01ae51ec69c217c78151b/assets/fig2-architecture.png" width="100%" alt="GroundAnything: vision-language architecture and autoregressive-to-diffusion conversion" /></p>
108
 
109
  ## Input(s):
110
 
@@ -173,7 +169,7 @@ Evaluation code and instructions: [GitHub](https://github.com/groundingpi/Ground
173
 
174
  ## Quantitative Evaluation Benchmarks
175
 
176
- <p align="center"><img src="https://huggingface.co/GroundingPI/GroundAnything/resolve/0f8e30894c3ca86378d01ae51ec69c217c78151b/assets/fig7-grounding-performance.png" width="100%" alt="GroundAnything: GroundAnything and GroundAnything-VLM benchmark overview" /></p>
177
 
178
  ## Inference:
179
 
@@ -278,7 +274,7 @@ After a block is complete, a causal forward reconstructs its authoritative KV ca
278
 
279
  The model uses its own shared weights to draft tokens with bidirectional attention and verify them with causal attention. Verification accepts the **longest consecutive matching prefix**, stops at the first mismatch, applies the causal correction, and discards the rejected suffix cache states. The shipped speculative route uses greedy verification; it is not a general stochastic speculative sampler.
280
 
281
- <p align="center"><img src="https://huggingface.co/GroundingPI/GroundAnything/resolve/0f8e30894c3ca86378d01ae51ec69c217c78151b/assets/fig6-self-speculative-decoding.png" width="100%" alt="GroundAnything: linear and quadratic self-speculative schedules with shared model weights" /></p>
282
 
283
  The documented `--decoder speculative` service is the linear shared-weight route. Exact greedy verification is relative to the converted model's causal branch; it does not imply identical outputs to the separately trained GroundAnything-VLM checkpoint.
284
 
 
19
  inference: false
20
  ---
21
 
22
+ # <img src="https://huggingface.co/GroundingPI/GroundAnything/resolve/20fda7044a5850df0c773831bd3b94ce9cf1b127/assets/logo.png" width="40" style="display: inline-block; vertical-align: middle; margin: 0;" alt="GroundAnything logo" /> GroundAnything: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed
23
 
24
  **Model family:** [GroundAnything β€” DLM / parallel decoding](https://huggingface.co/GroundingPI/GroundAnything) Β· [GroundAnything-VLM β€” autoregressive](https://huggingface.co/GroundingPI/GroundAnything-VLM).
25
 
26
  **This repository contains the GroundAnything DLM checkpoint.** Use entropy-guided decoding for the main benchmark setting, or select optional self-speculative decoding.
27
 
28
+ <p align="center"><img src="https://huggingface.co/GroundingPI/GroundAnything/resolve/20fda7044a5850df0c773831bd3b94ce9cf1b127/assets/fig1-teaser.png" width="100%" alt="GroundAnything: broad visual grounding and parallel visual evidence extraction" /></p>
29
 
30
  ## πŸ”— Quick Links
31
 
32
+ - πŸš€ **Online Demo:** The GroundAnything DLM runtime is not supported on ZeroGPU.
33
+ - 🌐 **Project Page:** [GroundAnything](https://groundingpi.github.io/groundanything/).
34
  - πŸ’» **GitHub Code:** [groundingpi/GroundAnything](https://github.com/groundingpi/GroundAnything).
35
  - πŸ“„ **Paper:** [arXiv:2609.39600](https://arxiv.org/abs/2609.39600).
 
36
 
37
  # Model Overview
38
 
 
46
 
47
  ### Demo Videos
48
 
49
+ <video controls playsinline preload="none" width="100%" poster="https://huggingface.co/GroundingPI/GroundAnything/resolve/20fda7044a5850df0c773831bd3b94ce9cf1b127/assets/demo-poster.jpg" src="https://huggingface.co/GroundingPI/GroundAnything/resolve/20fda7044a5850df0c773831bd3b94ce9cf1b127/assets/demo.mp4"></video>
50
 
51
  **Parallel Decoding**
52
 
53
+ <video controls playsinline preload="none" width="100%" src="https://huggingface.co/GroundingPI/GroundAnything/resolve/20fda7044a5850df0c773831bd3b94ce9cf1b127/assets/decoding.mp4"></video>
54
 
55
  ### License/Terms of Use:
56
 
 
72
 
73
  - **Paper [09/30/2026]:** [GroundAnything](https://arxiv.org/abs/2609.39600).
74
 
 
 
 
 
75
 
76
  <details>
77
  <summary>Citation</summary>
78
 
79
  ```bibtex
80
+ @misc{yu2026groundanythingreconcilingparalleldecoding,
81
+ title = {{GroundAnything}: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed},
82
  author = {Qize Yu and Lianrui Fan and Bowen Ping and Xini Ding and Zetian Song and Junbo Niu and Kaixuan Wang and Tianxing Chen and Yue Chen and Minghua He and Yuran Wang and Jie Huang and Haojun Zhang and Min Chen and Hao Li and Wenxuan Song and Ruihai Wu and Xianming Liu and Shilong Liu and Shuchang Zhou and Ping Luo and Shiyu Huang},
83
  year = {2026},
84
  eprint = {2609.39600},
 
100
  - **Spatial vocabulary:** 1,000 coordinate tokens shared with semantic labels and protocol markers.
101
  - **DLM conversion:** the shared decoder and vocabulary head support both causal prediction and bidirectional response-block denoising. A mask token is added for diffusion generation.
102
 
103
+ <p align="center"><img src="https://huggingface.co/GroundingPI/GroundAnything/resolve/20fda7044a5850df0c773831bd3b94ce9cf1b127/assets/fig2-architecture.png" width="100%" alt="GroundAnything: vision-language architecture and autoregressive-to-diffusion conversion" /></p>
104
 
105
  ## Input(s):
106
 
 
169
 
170
  ## Quantitative Evaluation Benchmarks
171
 
172
+ <p align="center"><img src="https://huggingface.co/GroundingPI/GroundAnything/resolve/20fda7044a5850df0c773831bd3b94ce9cf1b127/assets/fig7-grounding-performance.png" width="100%" alt="GroundAnything: GroundAnything and GroundAnything-VLM benchmark overview" /></p>
173
 
174
  ## Inference:
175
 
 
274
 
275
  The model uses its own shared weights to draft tokens with bidirectional attention and verify them with causal attention. Verification accepts the **longest consecutive matching prefix**, stops at the first mismatch, applies the causal correction, and discards the rejected suffix cache states. The shipped speculative route uses greedy verification; it is not a general stochastic speculative sampler.
276
 
277
+ <p align="center"><img src="https://huggingface.co/GroundingPI/GroundAnything/resolve/20fda7044a5850df0c773831bd3b94ce9cf1b127/assets/fig6-self-speculative-decoding.png" width="100%" alt="GroundAnything: linear and quadratic self-speculative schedules with shared model weights" /></p>
278
 
279
  The documented `--decoder speculative` service is the linear shared-weight route. Exact greedy verification is relative to the converted model's causal branch; it does not imply identical outputs to the separately trained GroundAnything-VLM checkpoint.
280
 
assets/demo-poster.jpg CHANGED

Git LFS Details

  • SHA256: 288c6121ec41e3114fe98221cf795381b1aace7c99660da945164be88af331f0
  • Pointer size: 131 Bytes
  • Size of remote file: 217 kB

Git LFS Details

  • SHA256: 3c282238bb1598093c078e57dcc1db4b3ce8fb1f39db6c09ca90e0686f1043a2
  • Pointer size: 131 Bytes
  • Size of remote file: 725 kB
assets/demo.mp4 CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:1d987aac49b390c0411faecc33892873fb0b8b811ff1f961d8a87cb39c406285
3
- size 22955151
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:cddaec9f528f9ebf5ac627850eefcb3f14ee55e168776a1b650f91353ce8293e
3
+ size 45075106
checksums.sha256 CHANGED
@@ -1,5 +1,5 @@
1
  c419c191a33774f4e1f834b3fb36033dc84e8ba316a3d1935eed1fa529884a74 LICENSE
2
- 7daa965f60851d4504e1cfc5ceb4cfc6ff40e083c855b27b300d8e346a3157b3 README.md
3
  669bb095cca2e86ddc821926f4c3b42dd7389f2ad3ec49c8edca8fe6294aae7f added_tokens.json
4
  a0bc6f6fc7a29a80017a433e8f03a1cc1236e838a944a2d034295a60c4f2fddb chat_template.jinja
5
  22c369e2bdac19723fe7db7e4c64b224462744b1f1aae7bd5daf05307db3b9fd config.json
@@ -22,3 +22,5 @@ d948513463b339ae85d6be6c99ecee787cc9c7a96be09180aa521a32a216eff4 special_tokens
22
  ca10d7e9fb3ed18575dd1e277a2579c16d108e32f27439684afa0e10b1440910 vocab.json
23
  cfc7749b96f63bd31c3c42b5c471bf756814053e847c10f3eb003417bc523d30 LICENSE-Apache-2.0
24
  20c797ce19af0c17de52c6afb144644768a591c521655f5ebf5712c9850f2887 LICENSE-Kimi-K3
 
 
 
1
  c419c191a33774f4e1f834b3fb36033dc84e8ba316a3d1935eed1fa529884a74 LICENSE
2
+ 331eeebd16f922886141ce4562f3b2504cddb0cb8f0adc7350bc93e72b207638 README.md
3
  669bb095cca2e86ddc821926f4c3b42dd7389f2ad3ec49c8edca8fe6294aae7f added_tokens.json
4
  a0bc6f6fc7a29a80017a433e8f03a1cc1236e838a944a2d034295a60c4f2fddb chat_template.jinja
5
  22c369e2bdac19723fe7db7e4c64b224462744b1f1aae7bd5daf05307db3b9fd config.json
 
22
  ca10d7e9fb3ed18575dd1e277a2579c16d108e32f27439684afa0e10b1440910 vocab.json
23
  cfc7749b96f63bd31c3c42b5c471bf756814053e847c10f3eb003417bc523d30 LICENSE-Apache-2.0
24
  20c797ce19af0c17de52c6afb144644768a591c521655f5ebf5712c9850f2887 LICENSE-Kimi-K3
25
+ cddaec9f528f9ebf5ac627850eefcb3f14ee55e168776a1b650f91353ce8293e assets/demo.mp4
26
+ 3c282238bb1598093c078e57dcc1db4b3ce8fb1f39db6c09ca90e0686f1043a2 assets/demo-poster.jpg