Image-Text-to-Text
Transformers
Safetensors
English
Chinese
groundinganything
text-generation
visual-grounding
object-detection
referring-expression-comprehension
pointing
ocr
document-layout
custom-code
diffusion-language-model
conversational
custom_code
Instructions to use GroundingPI/GroundAnything with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use GroundingPI/GroundAnything with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="GroundingPI/GroundAnything", trust_remote_code=True) messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("GroundingPI/GroundAnything", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use GroundingPI/GroundAnything with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "GroundingPI/GroundAnything" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "GroundingPI/GroundAnything", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/GroundingPI/GroundAnything
- SGLang
How to use GroundingPI/GroundAnything with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "GroundingPI/GroundAnything" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "GroundingPI/GroundAnything", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "GroundingPI/GroundAnything" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "GroundingPI/GroundAnything", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use GroundingPI/GroundAnything with Docker Model Runner:
docker model run hf.co/GroundingPI/GroundAnything
Document architecture, inference, evaluation, and paper results
Browse files- README.md +379 -0
- checksums.sha256 +1 -1
README.md
CHANGED
|
@@ -0,0 +1,379 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: apache-2.0
|
| 3 |
+
language:
|
| 4 |
+
- en
|
| 5 |
+
- zh
|
| 6 |
+
library_name: transformers
|
| 7 |
+
pipeline_tag: image-text-to-text
|
| 8 |
+
tags:
|
| 9 |
+
- visual-grounding
|
| 10 |
+
- object-detection
|
| 11 |
+
- referring-expression-comprehension
|
| 12 |
+
- pointing
|
| 13 |
+
- ocr
|
| 14 |
+
- document-layout
|
| 15 |
+
- custom-code
|
| 16 |
+
- diffusion-language-model
|
| 17 |
+
inference: false
|
| 18 |
+
---
|
| 19 |
+
|
| 20 |
+
<p align="center"><img src="assets/logo.png" width="150" alt="GroundAnything logo" /></p>
|
| 21 |
+
|
| 22 |
+
# GroundAnything: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed
|
| 23 |
+
|
| 24 |
+
**Model family:** [GroundAnything — DLM / parallel decoding](https://huggingface.co/GroundingPI/GroundAnything) · [GroundAnything-VLM — autoregressive](https://huggingface.co/GroundingPI/GroundAnything-VLM).
|
| 25 |
+
|
| 26 |
+
**This repository contains the GroundAnything DLM checkpoint.** Use entropy-guided decoding for the main benchmark setting, or select optional self-speculative decoding.
|
| 27 |
+
|
| 28 |
+
<p align="center"><img src="assets/fig1-teaser.png" width="100%" alt="GroundAnything Figure 1: broad visual grounding and parallel visual evidence extraction" /></p>
|
| 29 |
+
|
| 30 |
+
## 🔗 Quick Links
|
| 31 |
+
|
| 32 |
+
- 🚀 **Online Demo:** Coming soon — XXX.
|
| 33 |
+
- 💻 **GitHub Code:** Coming soon — XXX.
|
| 34 |
+
- 📄 **Paper:** [arXiv:2609.39600](https://arxiv.org/abs/2609.39600).
|
| 35 |
+
- 🧪 **Evaluation data:** Coming soon — XXX.
|
| 36 |
+
|
| 37 |
+
# Model Overview
|
| 38 |
+
|
| 39 |
+
### Description:
|
| 40 |
+
|
| 41 |
+
Autoregressive (AR) grounding models serialize spatial predictions, introducing sequential latency and imposing a causal order on output tokens. We view grounding as visual evidence extraction: objects, locations, and spatial relations are jointly constrained by the image and query, yet their dependencies do not imply an intrinsic left-to-right generation order. This distinction makes bidirectional diffusion a natural fit, allowing spatial hypotheses to emerge in parallel and be jointly refined through iterative denoising.
|
| 42 |
+
|
| 43 |
+
We introduce **GroundAnything**, a 4B-parameter grounding foundation model that reconciles fast parallel decoding with precise localization through blockwise denoising. Across 30 grounding benchmarks, the autoregressive variant, **GroundAnything-VLM**, establishes a new overall state of the art among similarly sized models at **72.42%**, remaining competitive with GPT-6 Astra (**71.35%**). With **entropy-guided decoding**, GroundAnything also surpasses the prior state of the art at this scale, averaging **61.75%** versus **53.32%** for the fast MTP-based LocateAnything model.
|
| 44 |
+
|
| 45 |
+
An optional **self-speculative mode** achieves a **4.51× speedup** over the AR counterpart with a **0.74 percentage-point drop in COCO F1mIoU** at the paper's reported operating point. Infrastructure experiments show that progressive inference optimizations translate parallel decoding into practical speedups. These support efficient visual grounding in latency-sensitive real-world systems.
|
| 46 |
+
|
| 47 |
+
### Demo Videos
|
| 48 |
+
|
| 49 |
+
<video controls playsinline preload="none" width="100%" poster="https://huggingface.co/GroundingPI/GroundAnything/resolve/0f8e30894c3ca86378d01ae51ec69c217c78151b/assets/demo-poster.jpg" src="https://huggingface.co/GroundingPI/GroundAnything/resolve/0f8e30894c3ca86378d01ae51ec69c217c78151b/assets/demo.mp4"></video>
|
| 50 |
+
|
| 51 |
+
**Parallel decoding in action**
|
| 52 |
+
|
| 53 |
+
<video controls playsinline preload="none" width="100%" src="https://huggingface.co/GroundingPI/GroundAnything/resolve/0f8e30894c3ca86378d01ae51ec69c217c78151b/assets/decoding.mp4"></video>
|
| 54 |
+
|
| 55 |
+
### License/Terms of Use:
|
| 56 |
+
|
| 57 |
+
This model is released under the [Apache License 2.0](https://www.apache.org/licenses/LICENSE-2.0).
|
| 58 |
+
|
| 59 |
+
### Deployment Geography:
|
| 60 |
+
|
| 61 |
+
Global.
|
| 62 |
+
|
| 63 |
+
### Use Case:
|
| 64 |
+
|
| 65 |
+
- Open-vocabulary object localization and dense-scene grounding.
|
| 66 |
+
- Referring-expression comprehension and visual-prompt-based localization.
|
| 67 |
+
- Point-based localization, spatial reasoning, and GUI element grounding.
|
| 68 |
+
- OCR with text localization and document layout understanding.
|
| 69 |
+
- Perception research for robotics, embodied agents, and autonomous systems.
|
| 70 |
+
|
| 71 |
+
### Release Date:
|
| 72 |
+
|
| 73 |
+
- **Paper [09/30/2026]:** [GroundAnything](https://arxiv.org/abs/2609.39600).
|
| 74 |
+
|
| 75 |
+
## References(s):
|
| 76 |
+
|
| 77 |
+
- [GroundAnything paper and supplementary material](https://arxiv.org/abs/2609.39600).
|
| 78 |
+
- [GroundingPI: grounding with visual primitives](https://arxiv.org/abs/2609.39601).
|
| 79 |
+
|
| 80 |
+
<details>
|
| 81 |
+
<summary>Citation</summary>
|
| 82 |
+
|
| 83 |
+
```bibtex
|
| 84 |
+
@misc{yu2026groundanything,
|
| 85 |
+
title = {GroundAnything: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed},
|
| 86 |
+
author = {Qize Yu and Lianrui Fan and Bowen Ping and Xini Ding and Zetian Song and Junbo Niu and Kaixuan Wang and Tianxing Chen and Yue Chen and Minghua He and Yuran Wang and Jie Huang and Haojun Zhang and Min Chen and Hao Li and Wenxuan Song and Ruihai Wu and Xianming Liu and Shilong Liu and Shuchang Zhou and Ping Luo and Shiyu Huang},
|
| 87 |
+
year = {2026},
|
| 88 |
+
eprint = {2609.39600},
|
| 89 |
+
archivePrefix = {arXiv},
|
| 90 |
+
primaryClass = {cs.CV},
|
| 91 |
+
url = {https://arxiv.org/abs/2609.39600},
|
| 92 |
+
}
|
| 93 |
+
```
|
| 94 |
+
|
| 95 |
+
</details>
|
| 96 |
+
|
| 97 |
+
## Model Architecture:
|
| 98 |
+
|
| 99 |
+
**Architecture type:** a shared vision-language backbone supporting an autoregressive checkpoint and a blockwise diffusion checkpoint.
|
| 100 |
+
|
| 101 |
+
- **Vision encoder:** MoonViT-V2 / Kimi-K3 vision backbone.
|
| 102 |
+
- **Language backbone:** Qwen3-4B-Instruct-2507.
|
| 103 |
+
- **Multimodal projector:** 2 × 2 spatial aggregation and a two-layer MLP.
|
| 104 |
+
- **Spatial vocabulary:** 1,000 coordinate tokens shared with semantic labels and protocol markers.
|
| 105 |
+
- **DLM conversion:** the shared decoder and vocabulary head support both causal prediction and bidirectional response-block denoising. A mask token is added for diffusion generation.
|
| 106 |
+
|
| 107 |
+
<p align="center"><img src="assets/fig2-architecture.png" width="100%" alt="GroundAnything Figure 2: vision-language architecture and autoregressive-to-diffusion conversion" /></p>
|
| 108 |
+
|
| 109 |
+
*Figure 2. Shared model architecture and the conversion to parallel grounding.*
|
| 110 |
+
|
| 111 |
+
<p align="center"><img src="assets/fig4-attention-mask.png" width="640" alt="GroundAnything Figure 4: clean-stream causal and noisy-response block attention" /></p>
|
| 112 |
+
|
| 113 |
+
*Figure 4. The attention mask used during diffusion conversion. The clean stream uses causal attention. A noisy response block attends bidirectionally within its block and reads the clean conditioning context and strictly preceding clean response blocks. Conversion uses B=32; the illustration uses two-token blocks.*
|
| 114 |
+
|
| 115 |
+
## Input(s):
|
| 116 |
+
|
| 117 |
+
**Input types:** image and text.
|
| 118 |
+
|
| 119 |
+
- **Image:** one RGB image per request. The Python client accepts JPEG, PNG, and WebP files; the supplied processor handles image preparation.
|
| 120 |
+
- **Text:** a natural-language instruction, category list, referring expression, OCR/layout query, or a prompt containing example boxes.
|
| 121 |
+
- **Image encoding for the HTTP API:** an `image_url` content part containing a base64 data URI, alongside a `text` content part.
|
| 122 |
+
|
| 123 |
+
Use the checkpoint's own tokenizer, processor, and chat template. Multiple categories are separated by `</c>`. Reference boxes use the same 0–999 spatial-token vocabulary as outputs.
|
| 124 |
+
|
| 125 |
+
## Output(s):
|
| 126 |
+
|
| 127 |
+
**Output type:** text containing semantic labels and quantized spatial coordinates.
|
| 128 |
+
|
| 129 |
+
All three released models use the **GAM protocol**: integer coordinate tokens `<0>` through `<999>`, object-reference delimiters, and box delimiters. A bounding box contains `(x1, y1, x2, y2)`; a point contains `(x, y)`. Adjacent coordinate tokens have no intervening spaces. Multiple instances of the same label are comma-separated inside one box wrapper. A missing target is represented by `None`.
|
| 130 |
+
|
| 131 |
+
Illustrative syntax:
|
| 132 |
+
|
| 133 |
+
```text
|
| 134 |
+
<|object_ref_start|>car<|object_ref_end|><|box_start|><100><200><500><650><|box_end|>
|
| 135 |
+
<|object_ref_start|>car center<|object_ref_end|><|box_start|><300><425><|box_end|>
|
| 136 |
+
<|object_ref_start|>absent object<|object_ref_end|><|box_start|>None<|box_end|>
|
| 137 |
+
```
|
| 138 |
+
|
| 139 |
+
The client returns parsed predictions together with `raw_output`, `finish_reason`, `usage`, `parse_error`, and `valid`. It maps coordinates back to image pixels for visualization. A truncated or malformed response is marked invalid. For custom API clients, preserve spatial tokens with `skip_special_tokens=false` and avoid inserting spaces between them.
|
| 140 |
+
|
| 141 |
+
## Software Integration:
|
| 142 |
+
|
| 143 |
+
**Runtime engines:** the repository's custom **SGLang** integration for the DLM and VLM services, plus a native Transformers reference route for the DLM.
|
| 144 |
+
|
| 145 |
+
**Default serving environment:** Linux and Python 3.12 with the source package's serving profile. Use `python3 run.py setup serve` to install the bundled custom engine and its pinned dependencies; installing upstream SGLang alone does not provide the same model and decoding integration.
|
| 146 |
+
|
| 147 |
+
The serving profile pins Torch **2.9.1**, Transformers **5.5.4**, Triton **3.5.1**, sgl-kernel **0.3.20**, and FlashInfer **0.5.3**. Serving and evaluation use separate environments.
|
| 148 |
+
|
| 149 |
+
**Tested hardware:** NVIDIA **B300, B200, H200, H800**, and **PPU**. The provided SGLang serving installer is a CUDA/GPU recipe; PPU requires its matching platform runtime and is not selected by this GPU installer.
|
| 150 |
+
|
| 151 |
+
The default DLM service uses **BF16, Triton attention, eager execution, one GPU, one active request, and two queued requests**. Client concurrency queues requests; it does not imply a multi-request model batch. CUDA Graph and selective FP8 are discussed below as separate infrastructure experiments.
|
| 152 |
+
|
| 153 |
+
## Model Version(s):
|
| 154 |
+
|
| 155 |
+
| Checkpoint | Generation | Evaluation mode |
|
| 156 |
+
|:---|:---|:---|
|
| 157 |
+
| [GroundAnything](https://huggingface.co/GroundingPI/GroundAnything) | Entropy-guided blockwise diffusion; optional self-speculation | **GAM** |
|
| 158 |
+
| [GroundAnything-VLM](https://huggingface.co/GroundingPI/GroundAnything-VLM) | Autoregressive generation | **GAM** |
|
| 159 |
+
|
| 160 |
+
The main GroundAnything benchmark results use **entropy-guided diffusion without autoregressive verification**. GroundAnything-VLM uses autoregressive decoding. The self-speculative speed–quality operating point is a separate experiment and should not be substituted for the main diffusion benchmark setting.
|
| 161 |
+
|
| 162 |
+
## Testing and Evaluation Datasets:
|
| 163 |
+
|
| 164 |
+
### Data Modality:
|
| 165 |
+
|
| 166 |
+
Image and text, with task-specific box, point, text-region, or interaction annotations.
|
| 167 |
+
|
| 168 |
+
## Evaluation Dataset:
|
| 169 |
+
|
| 170 |
+
**Evaluation data: XXX — download URL coming soon.**
|
| 171 |
+
|
| 172 |
+
The shared evaluation toolkit provides **42 task recipes across eight task families**: Grounding, Referring, Dense, OCR, Layout, GUI, Pointing, and VisualPrompt. These are executable task/split recipes, not a count of distinct datasets. Dataset paths are registered in `configs/datasets.yaml`; task IDs are listed in `configs/eval/tasks.json`.
|
| 173 |
+
|
| 174 |
+
### Evaluation Modes
|
| 175 |
+
|
| 176 |
+
The evaluator accepts seven mode identifiers. A mode selects the **prompt and output parser** independently of the inference engine and decoding algorithm.
|
| 177 |
+
|
| 178 |
+
| Mode | Supported evaluation interface | Output / coordinates |
|
| 179 |
+
|:---|:---|:---|
|
| 180 |
+
| **`GAM`** | **GroundingPI, GroundAnything, and GroundAnything-VLM** | Native spatial tokens, **0–999** |
|
| 181 |
+
| `DLM` | Legacy compatible diffusion-checkpoint alias | Same GAM spatial-token protocol |
|
| 182 |
+
| `RLV2` | Compatible RL-checkpoint alias | Same GAM spatial-token protocol |
|
| 183 |
+
| `VLM` | Generic VLM baselines | Explicit pixel, 0–1000, or 0–1 coordinate mode |
|
| 184 |
+
| `REXOMNI` | Rex-Omni adapter | Model-specific spatial-token parser |
|
| 185 |
+
| `LOCATEANYTHING` | LocateAnything adapter | Its own output parser and generation-mode setting |
|
| 186 |
+
| `GROUNDINGDINO` | External compatible GroundingDINO bridge | JSON coordinate responses |
|
| 187 |
+
|
| 188 |
+
**Use `mode: GAM` for all three of our checkpoints.** The `-VLM` suffix identifies an autoregressive checkpoint; it does not select the evaluator's generic `VLM` mode. The DLM's entropy-guided or self-speculative algorithm is selected separately through `decoder`.
|
| 189 |
+
|
| 190 |
+
External baseline weights and model services are separate from the evaluation toolkit. Metrics remain task-specific: box localization quality, point accuracy, OCR text-and-region matching, GUI grounding accuracy, and counting error. Full protocols and benchmark tables are provided in the paper and supplementary material.
|
| 191 |
+
|
| 192 |
+
## Quantitative Evaluation Benchmarks
|
| 193 |
+
|
| 194 |
+
<p align="center"><img src="assets/fig7-grounding-performance.png" width="100%" alt="GroundAnything Figure 7: GroundAnything and GroundAnything-VLM benchmark overview" /></p>
|
| 195 |
+
|
| 196 |
+
*Figure 7. Paper-reported capability overview. Detailed task-level scores, baseline settings, and evaluation protocols are provided in the paper and supplementary material.*
|
| 197 |
+
|
| 198 |
+
## Inference:
|
| 199 |
+
|
| 200 |
+
### Installation
|
| 201 |
+
|
| 202 |
+
Run the commands from the **GroundAnything source repository root** after obtaining the code package, using **Linux x86_64 and Python 3.12**. The code link above is reserved for the public release. Use the serving profile in the supplied source bundle so that the custom model adapter, decoding implementation, and dependencies remain aligned.
|
| 203 |
+
|
| 204 |
+
```bash
|
| 205 |
+
python3 -m pip install -r requirements.txt huggingface_hub
|
| 206 |
+
python3 run.py setup serve
|
| 207 |
+
```
|
| 208 |
+
|
| 209 |
+
The installer verifies and extracts its bundled frameworks, creates `.venv-serve`, and records resolved packages. It prepares the custom SGLang implementation and applies the serving profile's dependency order. The profile includes the required cuDNN compatibility selection; its dependency-check report records the known Torch/cuDNN metadata exception.
|
| 210 |
+
|
| 211 |
+
**Choose the checkpoint-specific recipe below.** GroundAnything-VLM users need only the VLM download and launch commands; the DLM commands require the separate GroundAnything weights.
|
| 212 |
+
|
| 213 |
+
### GroundAnything: Entropy-guided Service
|
| 214 |
+
|
| 215 |
+
Download the **complete DLM model bundle**, including its custom code and tokenizer, then launch:
|
| 216 |
+
|
| 217 |
+
```bash
|
| 218 |
+
hf download GroundingPI/GroundAnything --local-dir weights/dlm_bundle
|
| 219 |
+
python3 run.py serve --decoder denoise
|
| 220 |
+
```
|
| 221 |
+
|
| 222 |
+
The published model package is already a DLM bundle, so `prepare-model` is unnecessary for this download. The endpoint is `http://127.0.0.1:8101/v1`, with model ID `groundinganything`.
|
| 223 |
+
|
| 224 |
+
### GroundAnything: Self-speculative Service
|
| 225 |
+
|
| 226 |
+
Stop the existing DLM service before switching its decoder:
|
| 227 |
+
|
| 228 |
+
```bash
|
| 229 |
+
python3 run.py serve --decoder speculative
|
| 230 |
+
```
|
| 231 |
+
|
| 232 |
+
This reuses the same DLM weights and endpoint, with diffusion proposals and greedy causal verification.
|
| 233 |
+
|
| 234 |
+
### GroundAnything-VLM: Autoregressive Service
|
| 235 |
+
|
| 236 |
+
Use the separate autoregressive checkpoint and its service recipe:
|
| 237 |
+
|
| 238 |
+
```bash
|
| 239 |
+
hf download GroundingPI/GroundAnything-VLM --local-dir weights/vlm
|
| 240 |
+
python3 run.py serve --config configs/release/vlm_sglang.yaml
|
| 241 |
+
```
|
| 242 |
+
|
| 243 |
+
This service uses `http://127.0.0.1:8102/v1`, with model ID `groundinganything-vlm`. The two repositories share the spatial interface but have different loading and generation paths.
|
| 244 |
+
|
| 245 |
+
### Worker (recommended)
|
| 246 |
+
|
| 247 |
+
Start the appropriate service once, then reuse the client:
|
| 248 |
+
|
| 249 |
+
```python
|
| 250 |
+
from grounding_anything import GroundingAnything, visualize
|
| 251 |
+
|
| 252 |
+
client = GroundingAnything(
|
| 253 |
+
base_url="http://127.0.0.1:8101/v1",
|
| 254 |
+
model="groundinganything",
|
| 255 |
+
)
|
| 256 |
+
result = client.predict("example.jpg", "the red car", task="bbox")
|
| 257 |
+
print(result.to_dict())
|
| 258 |
+
if result.valid:
|
| 259 |
+
visualize("example.jpg", result).save("prediction.png")
|
| 260 |
+
|
| 261 |
+
point_result = client.predict(
|
| 262 |
+
"example.jpg", "the center of the red car", task="point"
|
| 263 |
+
)
|
| 264 |
+
```
|
| 265 |
+
|
| 266 |
+
The HTTP client retains no model weights. A generic interactive request has no benchmark task identity; it does not automatically receive the benchmark's task-specific sampling and stopping policy. Use the evaluator below to reproduce the reported protocol.
|
| 267 |
+
|
| 268 |
+
### Supported Tasks & Prompt Templates
|
| 269 |
+
|
| 270 |
+
| Task | Prompt example |
|
| 271 |
+
|:---|:---|
|
| 272 |
+
| Category / dense grounding | `Locate all the instances that match the following categories: car</c>person.` |
|
| 273 |
+
| Referring boxes | `Locate the target referred to by the following description: the red car.` |
|
| 274 |
+
| Category points | `Point to: car</c>person.` |
|
| 275 |
+
| Referring points | `Point to the target referred to by the following description: the red car.` |
|
| 276 |
+
| OCR | `OCR task detect all the text in box format.` |
|
| 277 |
+
| Layout | `Detect all document layout elements that match the following categories: title</c>text.` |
|
| 278 |
+
| GUI | `Point to the UI element to click for the following instruction: open the settings menu.` |
|
| 279 |
+
| Visual prompting | Provide reference boxes in the native spatial-token format, then request similar objects. |
|
| 280 |
+
|
| 281 |
+
```text
|
| 282 |
+
Given reference boxes <|box_start|><100><200><500><650><|box_end|> indicating one or more objects, find all similar objects in the image and output their bounding boxes.
|
| 283 |
+
```
|
| 284 |
+
|
| 285 |
+
The convenience client's `predict()` method wraps referring-box and referring-point prompts. Use an OpenAI-compatible request to `/chat/completions` for the other task templates; keep the image and prompt in the same user message.
|
| 286 |
+
|
| 287 |
+
### Generation Modes
|
| 288 |
+
|
| 289 |
+
#### Entropy-guided decoding
|
| 290 |
+
|
| 291 |
+
Image/query prefill creates a causal prefix cache and a known anchor. The release recipe uses **block size 32, sub-block size 4, and entropy threshold 0.8**. A physical block contains the known anchor and 31 masked positions; sub-blocks are completed from left to right while the forward pass evaluates the physical block.
|
| 292 |
+
|
| 293 |
+
For each still-masked position in the active sub-block, the decoder measures entropy from the unmodified token distribution over the generatable vocabulary, excluding the mask token. Positions at or below the threshold are committed together. If no position qualifies, the lowest-entropy position is committed so decoding makes progress. Committed tokens remain fixed.
|
| 294 |
+
|
| 295 |
+
After a block is complete, a causal forward reconstructs its authoritative KV cache and supplies the next anchor. This cache-building pass **does not verify or reject the generated block**. A block requiring D denoising passes therefore uses D + 1 model forwards, excluding the initial prefill.
|
| 296 |
+
|
| 297 |
+
#### Self-speculative decoding
|
| 298 |
+
|
| 299 |
+
The model uses its own shared weights to draft tokens with bidirectional attention and verify them with causal attention. Verification accepts the **longest consecutive matching prefix**, stops at the first mismatch, applies the causal correction, and discards the rejected suffix cache states. The shipped speculative route uses greedy verification; it is not a general stochastic speculative sampler.
|
| 300 |
+
|
| 301 |
+
<p align="center"><img src="assets/fig6-self-speculative-decoding.png" width="100%" alt="GroundAnything Figure 6: linear and quadratic self-speculative schedules with shared model weights" /></p>
|
| 302 |
+
|
| 303 |
+
*Figure 6. The paper studies two schedules: linear drafting and verification use two model forwards and 2B query tokens per round; quadratic fusion uses one forward after initialization with B(B + 1) query tokens. These counts describe queries and model calls, not total Transformer FLOPs.*
|
| 304 |
+
|
| 305 |
+
The documented `--decoder speculative` service is the linear shared-weight route. Exact greedy verification is relative to the converted model's causal branch; it does not imply identical outputs to the separately trained GroundAnything-VLM checkpoint.
|
| 306 |
+
|
| 307 |
+
### Benchmark Sampling Policy
|
| 308 |
+
|
| 309 |
+
Entropy-guided evaluation applies the five task profiles from `infer/decode/configs/task_profiles.json`:
|
| 310 |
+
|
| 311 |
+
| Task profile | Recipes | Temperature | Top-p | Output-token budget |
|
| 312 |
+
|:---|---:|---:|---:|---:|
|
| 313 |
+
| Strict single target | 14 | 0 | 1 | 512 |
|
| 314 |
+
| Medium, non-OCR | 12 | 0.1 | 0.95 | 4,096 |
|
| 315 |
+
| Dense, non-OCR | 8 | 0.1 | 0.95 | 8,192 |
|
| 316 |
+
| Medium OCR | 4 | 0.3 | 0.95 | 4,096 |
|
| 317 |
+
| Dense OCR | 4 | 0 | 1 | 4,096 |
|
| 318 |
+
|
| 319 |
+
The evaluator sends outer sampling parameters and the nested `custom_params.gam_dlm_decode` contract, adds the strict-single-target policy when required, and checks the server's actual decoder before sending requests. Unknown task profiles and mismatched decoders fail early. The entropy-guided benchmark policy rejects a global `max_tokens` override because it would replace the task budgets.
|
| 320 |
+
|
| 321 |
+
Self-speculative evaluation uses greedy sampling (`temperature=0`, `top_p=1`, repetition penalty 1). GroundAnything-VLM retains its own autoregressive task recipes. All of these routes use **GAM** prompts and spatial-token parsing.
|
| 322 |
+
|
| 323 |
+
### Run Evaluation
|
| 324 |
+
|
| 325 |
+
```bash
|
| 326 |
+
python3 run.py setup eval
|
| 327 |
+
|
| 328 |
+
# Match the command to the service that is already running.
|
| 329 |
+
python3 run.py eval --decoder denoise
|
| 330 |
+
python3 run.py eval --decoder speculative
|
| 331 |
+
python3 run.py eval --config configs/eval/vlm.yaml
|
| 332 |
+
```
|
| 333 |
+
|
| 334 |
+
Choose one evaluation command for the intended checkpoint/decoder:
|
| 335 |
+
|
| 336 |
+
| Checkpoint / decoder | Recipe | Mode | Endpoint model ID |
|
| 337 |
+
|:---|:---|:---|:---|
|
| 338 |
+
| GroundAnything / entropy-guided | `configs/eval/dlm.yaml` | **GAM** | `groundinganything` |
|
| 339 |
+
| GroundAnything / self-speculative | `configs/eval/dlm_speculative.yaml` | **GAM** | `groundinganything` |
|
| 340 |
+
| GroundAnything-VLM / autoregressive | `configs/eval/vlm.yaml` | **GAM** | `groundinganything-vlm` |
|
| 341 |
+
|
| 342 |
+
Use `service_contract: openai` and the corresponding port (8101 for DLM, 8102 for VLM). For a preflight that checks configured inputs without sending inference requests:
|
| 343 |
+
|
| 344 |
+
```bash
|
| 345 |
+
.venv-eval/bin/python scripts/evaluate.py configs/eval/dlm.yaml --dry-run
|
| 346 |
+
```
|
| 347 |
+
|
| 348 |
+
Configure the running service, `data_root`, dataset registry, task list, and a fresh `run_id` before launching. The default `limit: 8` is a smoke test; set **`limit: null`** for full evaluation. Evaluation connects to an existing service and does not start or change the decoder.
|
| 349 |
+
|
| 350 |
+
Each run saves `run.json` (configuration and provenance), `responses.jsonl` (raw responses, finish reasons, and token usage), task logs, and `summary.json` (metrics and completion state). Record the checkpoint revision and task configuration when comparing results.
|
| 351 |
+
|
| 352 |
+
## Inference Infrastructure
|
| 353 |
+
|
| 354 |
+
### SGLang Execution
|
| 355 |
+
|
| 356 |
+
The custom SGLang integration coordinates model loading, request scheduling, attention kernels, and KV-cache ownership for blockwise generation. Denoising uses bidirectional attention inside the active block; completed history is retained as causal KV. Self-speculation additionally verifies proposals and removes rejected suffix states. These cache semantics must be preserved when optimizing execution.
|
| 357 |
+
|
| 358 |
+
The supplied recipe selects **BF16 + Triton attention + eager execution**. DLM recipes record the engine source, decoder, effective settings, profile digest, and package versions in `outputs/sglang/<decoder>/engine_runtime.json`; DLM evaluation checks decoder alignment before starting workers. The VLM service records its causal model/runtime configuration separately in `outputs/sglang/vlm/engine_runtime.json`. For higher service concurrency, use independent replicas with distinct devices, ports, and output directories; the default queue is not continuous multi-request batching.
|
| 359 |
+
|
| 360 |
+
### CUDA Graph and Selective FP8
|
| 361 |
+
|
| 362 |
+
The paper evaluates a progressive infrastructure sequence:
|
| 363 |
+
|
| 364 |
+
| Layer | Purpose | Status in the supplied default recipe |
|
| 365 |
+
|:---|:---|:---|
|
| 366 |
+
| Native PyTorch eager | Reference model execution | Reference implementation |
|
| 367 |
+
| SGLang eager | Integrated scheduling, attention, and cache execution | **Default serving path** |
|
| 368 |
+
| CUDA Graph replay | Replay compatible captured GPU work to reduce repeated launch overhead | Evaluated in the paper; **disabled by the default launcher** |
|
| 369 |
+
| Selective FP8 | Reduce arithmetic cost in eligible language-model linear operations while retaining other components in BF16 | Evaluated in the paper; **default serving remains BF16** |
|
| 370 |
+
|
| 371 |
+
CUDA Graph changes how compatible GPU work is submitted; it does not define a new token-commitment or verification rule. Captured shapes and state updates must remain compatible with the active decoding path. Selective FP8 can change logits and subsequent decoding decisions, so it is a distinct numerical configuration.
|
| 372 |
+
|
| 373 |
+
The optional graph implementation captures fixed-shape block work. Its FlashInfer path uses persistent attention masks and device-resident buffer updates for bidirectional drafting and causal verification. The Triton verifier path separates draft/verification metadata and input buffers while sharing parameters and the real KV pool. Its supported capture case is B32 with one request and no tensor/pipeline/data parallel expansion; prefill and unsupported shapes use eager execution. Optional shadow checks compare cache states, logits, and token choices against eager execution. These implementation paths are not exposed as a public `run.py serve --cuda-graph` switch in this release.
|
| 374 |
+
|
| 375 |
+
The paper's cumulative infrastructure speedups compare execution implementations **within the same decoding mode**. They are separate from the headline comparison of self-speculation against the autoregressive checkpoint. The default release commands above do not enable Graph replay or FP8, and no unsupported activation flags are implied.
|
| 376 |
+
|
| 377 |
+
## Ethical Considerations:
|
| 378 |
+
|
| 379 |
+
Grounding predictions can miss small or occluded objects, repeat instances, or produce inaccurate text and coordinates. Validate localization quality on the intended task and inspect incomplete responses. GUI points describe image locations; the model does not execute interface actions. Perception outputs require task-specific validation before integration into physical systems.
|
checksums.sha256
CHANGED
|
@@ -1,5 +1,5 @@
|
|
| 1 |
20c797ce19af0c17de52c6afb144644768a591c521655f5ebf5712c9850f2887 LICENSE
|
| 2 |
-
|
| 3 |
669bb095cca2e86ddc821926f4c3b42dd7389f2ad3ec49c8edca8fe6294aae7f added_tokens.json
|
| 4 |
a0bc6f6fc7a29a80017a433e8f03a1cc1236e838a944a2d034295a60c4f2fddb chat_template.jinja
|
| 5 |
22c369e2bdac19723fe7db7e4c64b224462744b1f1aae7bd5daf05307db3b9fd config.json
|
|
|
|
| 1 |
20c797ce19af0c17de52c6afb144644768a591c521655f5ebf5712c9850f2887 LICENSE
|
| 2 |
+
2c09ea3d520a5ed5d06ed0fd5d02c57df4d9be4b05ed469ad30c231c5c1221c5 README.md
|
| 3 |
669bb095cca2e86ddc821926f4c3b42dd7389f2ad3ec49c8edca8fe6294aae7f added_tokens.json
|
| 4 |
a0bc6f6fc7a29a80017a433e8f03a1cc1236e838a944a2d034295a60c4f2fddb chat_template.jinja
|
| 5 |
22c369e2bdac19723fe7db7e4c64b224462744b1f1aae7bd5daf05307db3b9fd config.json
|