Image-Text-to-Text
Transformers
Safetensors
English
Chinese
groundinganything
text-generation
visual-grounding
object-detection
referring-expression-comprehension
pointing
ocr
document-layout
custom-code
diffusion-language-model
conversational
custom_code
Instructions to use GroundingPI/GroundAnything with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use GroundingPI/GroundAnything with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="GroundingPI/GroundAnything", trust_remote_code=True) messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("GroundingPI/GroundAnything", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use GroundingPI/GroundAnything with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "GroundingPI/GroundAnything" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "GroundingPI/GroundAnything", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/GroundingPI/GroundAnything
- SGLang
How to use GroundingPI/GroundAnything with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "GroundingPI/GroundAnything" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "GroundingPI/GroundAnything", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "GroundingPI/GroundAnything" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "GroundingPI/GroundAnything", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use GroundingPI/GroundAnything with Docker Model Runner:
docker model run hf.co/GroundingPI/GroundAnything
|
Download README.md from GroundingPI/GroundAnything: direct link, hf CLI and curl.
- Browser
- Download file 20.2 kB
-
https://huggingface.co/GroundingPI/GroundAnything/resolve/main/README.md
- Command line
-
hf download hf://GroundingPI/GroundAnything/README.md
-
curl -L -o README.md https://huggingface.co/GroundingPI/GroundAnything/resolve/main/README.md
20.2 kB
| license: other | |
| license_name: apache-2.0-with-upstream-kimi-k3 | |
| license_link: LICENSE | |
| language: | |
| - en | |
| - zh | |
| library_name: transformers | |
| pipeline_tag: image-text-to-text | |
| tags: | |
| - visual-grounding | |
| - object-detection | |
| - referring-expression-comprehension | |
| - pointing | |
| - ocr | |
| - document-layout | |
| - custom-code | |
| - diffusion-language-model | |
| inference: false | |
| # <img src="https://huggingface.co/GroundingPI/GroundAnything/resolve/0f8e30894c3ca86378d01ae51ec69c217c78151b/assets/logo.png" width="40" style="display: inline-block; vertical-align: middle; margin: 0;" alt="GroundAnything logo" /> GroundAnything: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed | |
| **Model family:** [GroundAnything — DLM / parallel decoding](https://huggingface.co/GroundingPI/GroundAnything) · [GroundAnything-VLM — autoregressive](https://huggingface.co/GroundingPI/GroundAnything-VLM). | |
| **This repository contains the GroundAnything DLM checkpoint.** Use entropy-guided decoding for the main benchmark setting, or select optional self-speculative decoding. | |
| <p align="center"><img src="https://huggingface.co/GroundingPI/GroundAnything/resolve/0f8e30894c3ca86378d01ae51ec69c217c78151b/assets/fig1-teaser.png" width="100%" alt="GroundAnything: broad visual grounding and parallel visual evidence extraction" /></p> | |
| ## 🔗 Quick Links | |
| - 🚀 **Online Demo:** Coming soon — XXX. | |
| - 💻 **GitHub Code:** [groundingpi/GroundAnything](https://github.com/groundingpi/GroundAnything). | |
| - 📄 **Paper:** [arXiv:2609.39600](https://arxiv.org/abs/2609.39600). | |
| - 🧪 **Evaluation data:** Coming soon — XXX. | |
| # Model Overview | |
| ### Description: | |
| Autoregressive (AR) grounding models serialize spatial predictions, introducing sequential latency and imposing a causal order on output tokens. We view grounding as visual evidence extraction: objects, locations, and spatial relations are jointly constrained by the image and query, yet their dependencies do not imply an intrinsic left-to-right generation order. This distinction makes bidirectional diffusion a natural fit, allowing spatial hypotheses to emerge in parallel and be jointly refined through iterative denoising. | |
| We introduce **GroundAnything**, a 4B-parameter grounding foundation model that reconciles fast parallel decoding with precise localization through blockwise denoising. Across 30 grounding benchmarks, the autoregressive variant, **GroundAnything-VLM**, establishes a new overall state of the art among similarly sized models at **72.42%**, remaining competitive with GPT-6 Astra (**71.35%**). With **entropy-guided decoding**, GroundAnything also surpasses the prior state of the art at this scale, averaging **61.75%** versus **53.32%** for the fast MTP-based LocateAnything model. | |
| An optional **self-speculative mode** achieves a **4.51× speedup** over the AR counterpart with a **0.74 percentage-point drop in COCO F1mIoU** at the paper's reported operating point. Infrastructure experiments show that progressive inference optimizations translate parallel decoding into practical speedups. These support efficient visual grounding in latency-sensitive real-world systems. | |
| ### Demo Videos | |
| <video controls playsinline preload="none" width="100%" poster="https://huggingface.co/GroundingPI/GroundAnything/resolve/0f8e30894c3ca86378d01ae51ec69c217c78151b/assets/demo-poster.jpg" src="https://huggingface.co/GroundingPI/GroundAnything/resolve/0f8e30894c3ca86378d01ae51ec69c217c78151b/assets/demo.mp4"></video> | |
| **Parallel Decoding** | |
| <video controls playsinline preload="none" width="100%" src="https://huggingface.co/GroundingPI/GroundAnything/resolve/0f8e30894c3ca86378d01ae51ec69c217c78151b/assets/decoding.mp4"></video> | |
| ### License/Terms of Use: | |
| Our original contributions are available under **Apache 2.0**, with no additional restrictions imposed by this project. Third-party material retains its applicable licenses, including the **Kimi K3 License** for Kimi-derived material and applicable derivative works. Its conditions continue to apply when using or redistributing the combined model package. See [LICENSE](LICENSE) for the scope and full license texts. | |
| ### Deployment Geography: | |
| Global. | |
| ### Use Case: | |
| - Open-vocabulary object localization and dense-scene grounding. | |
| - Referring-expression comprehension and visual-prompt-based localization. | |
| - Point-based localization, spatial reasoning, and GUI element grounding. | |
| - OCR with text localization and document layout understanding. | |
| - Perception research for robotics, embodied agents, and autonomous systems. | |
| ### Release Date: | |
| - **Paper [09/30/2026]:** [GroundAnything](https://arxiv.org/abs/2609.39600). | |
| ## References(s): | |
| - [GroundAnything paper and supplementary material](https://arxiv.org/abs/2609.39600). | |
| - [GroundingPI: grounding with visual primitives](https://arxiv.org/abs/2609.39601). | |
| <details> | |
| <summary>Citation</summary> | |
| ```bibtex | |
| @misc{yu2026groundanything, | |
| title = {GroundAnything: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed}, | |
| author = {Qize Yu and Lianrui Fan and Bowen Ping and Xini Ding and Zetian Song and Junbo Niu and Kaixuan Wang and Tianxing Chen and Yue Chen and Minghua He and Yuran Wang and Jie Huang and Haojun Zhang and Min Chen and Hao Li and Wenxuan Song and Ruihai Wu and Xianming Liu and Shilong Liu and Shuchang Zhou and Ping Luo and Shiyu Huang}, | |
| year = {2026}, | |
| eprint = {2609.39600}, | |
| archivePrefix = {arXiv}, | |
| primaryClass = {cs.CV}, | |
| url = {https://arxiv.org/abs/2609.39600}, | |
| } | |
| ``` | |
| </details> | |
| ## Model Architecture: | |
| **Architecture type:** a shared vision-language backbone supporting an autoregressive checkpoint and a blockwise diffusion checkpoint. | |
| - **Vision encoder:** MoonViT-V2 / Kimi-K3 vision backbone. | |
| - **Language backbone:** Qwen3-4B-Instruct-2507. | |
| - **Multimodal projector:** 2 × 2 spatial aggregation and a two-layer MLP. | |
| - **Spatial vocabulary:** 1,000 coordinate tokens shared with semantic labels and protocol markers. | |
| - **DLM conversion:** the shared decoder and vocabulary head support both causal prediction and bidirectional response-block denoising. A mask token is added for diffusion generation. | |
| <p align="center"><img src="https://huggingface.co/GroundingPI/GroundAnything/resolve/0f8e30894c3ca86378d01ae51ec69c217c78151b/assets/fig2-architecture.png" width="100%" alt="GroundAnything: vision-language architecture and autoregressive-to-diffusion conversion" /></p> | |
| ## Input(s): | |
| **Input types:** image and text. | |
| - **Image:** one RGB image per request. The Python client accepts JPEG, PNG, and WebP files; the supplied processor handles image preparation. | |
| - **Text:** a natural-language instruction, category list, referring expression, OCR/layout query, or a prompt containing example boxes. | |
| - **Image encoding for the HTTP API:** an `image_url` content part containing a base64 data URI, alongside a `text` content part. | |
| Use the checkpoint's own tokenizer, processor, and chat template. Multiple categories are separated by `</c>`. Reference boxes use the same 0–999 spatial-token vocabulary as outputs. | |
| ## Output(s): | |
| **Output type:** text containing semantic labels and quantized spatial coordinates. | |
| All three released models use the **GAM protocol**: integer coordinate tokens `<0>` through `<999>`, object-reference delimiters, and box delimiters. A bounding box contains `(x1, y1, x2, y2)`; a point contains `(x, y)`. Adjacent coordinate tokens have no intervening spaces. Multiple instances of the same label are comma-separated inside one box wrapper. A missing target is represented by `None`. | |
| Illustrative syntax: | |
| ```text | |
| <|object_ref_start|>car<|object_ref_end|><|box_start|><100><200><500><650><|box_end|> | |
| <|object_ref_start|>car center<|object_ref_end|><|box_start|><300><425><|box_end|> | |
| <|object_ref_start|>absent object<|object_ref_end|><|box_start|>None<|box_end|> | |
| ``` | |
| The client returns parsed predictions together with `raw_output`, `finish_reason`, `usage`, `parse_error`, and `valid`. It maps coordinates back to image pixels for visualization. A truncated or malformed response is marked invalid. For custom API clients, preserve spatial tokens with `skip_special_tokens=false` and avoid inserting spaces between them. | |
| ## Software Integration: | |
| **Runtime engines:** the repository's custom **SGLang** integration for the DLM and VLM services, plus a native Transformers reference route for the DLM. | |
| **Default serving environment:** Linux and Python 3.12 with the source package's serving profile. Use `python3 run.py setup serve` to install the bundled custom engine and its pinned dependencies; installing upstream SGLang alone does not provide the same model and decoding integration. | |
| The serving profile pins Torch **2.9.1**, Transformers **5.5.4**, Triton **3.5.1**, sgl-kernel **0.3.20**, and FlashInfer **0.5.3**. Serving and evaluation use separate environments. | |
| **Tested hardware:** NVIDIA **B300, B200, H200, H800**, and **PPU**. The provided SGLang serving installer is a CUDA/GPU recipe; PPU requires its matching platform runtime and is not selected by this GPU installer. | |
| The default DLM service uses **BF16, Triton attention, eager execution, one GPU, one active request, and two queued requests**. Client concurrency queues requests; it does not imply a multi-request model batch. CUDA Graph and selective FP8 are discussed below as separate infrastructure experiments. | |
| ## Model Version(s): | |
| | Checkpoint | Generation | Evaluation mode | | |
| |:---|:---|:---| | |
| | [GroundAnything](https://huggingface.co/GroundingPI/GroundAnything) | Entropy-guided blockwise diffusion; optional self-speculation | **GAM** | | |
| | [GroundAnything-VLM](https://huggingface.co/GroundingPI/GroundAnything-VLM) | Autoregressive generation | **GAM** | | |
| GroundAnything's main results use **entropy-guided decoding**; GroundAnything-VLM uses **autoregressive decoding**. Self-speculative decoding is an optional acceleration mode. | |
| ## Evaluation | |
| The evaluation toolkit supports **7 modes**: | |
| | Mode | Supported models | | |
| |:---|:---| | |
| | **`GAM`** | **GroundingPI, GroundAnything, GroundAnything-VLM** | | |
| | `VLM` | Generic vision-language baselines | | |
| | `REXOMNI` | Rex-Omni | | |
| | `LOCATEANYTHING` | LocateAnything | | |
| | `GROUNDINGDINO` | GroundingDINO through a compatible service | | |
| | `DLM` | Legacy diffusion checkpoints using the GAM protocol | | |
| | `RLV2` | Legacy RL checkpoints using the GAM protocol | | |
| **All three released checkpoints use GAM mode.** | |
| Evaluation code and instructions: [GitHub](https://github.com/groundingpi/GroundAnything). | |
| ## Quantitative Evaluation Benchmarks | |
| <p align="center"><img src="https://huggingface.co/GroundingPI/GroundAnything/resolve/0f8e30894c3ca86378d01ae51ec69c217c78151b/assets/fig7-grounding-performance.png" width="100%" alt="GroundAnything: GroundAnything and GroundAnything-VLM benchmark overview" /></p> | |
| ## Inference: | |
| ### Installation | |
| Run the commands from the **GroundAnything source repository root** after obtaining the code package, using **Linux x86_64 and Python 3.12**. Use the serving profile in the supplied source bundle so that the custom model adapter, decoding implementation, and dependencies remain aligned. | |
| ```bash | |
| python3 -m pip install -r requirements.txt huggingface_hub | |
| python3 run.py setup serve | |
| ``` | |
| The installer verifies and extracts its bundled frameworks, creates `.venv-serve`, and records resolved packages. It prepares the custom SGLang implementation and applies the serving profile's dependency order. The profile includes the required cuDNN compatibility selection; its dependency-check report records the known Torch/cuDNN metadata exception. | |
| **Choose the checkpoint-specific recipe below.** GroundAnything-VLM users need only the VLM download and launch commands; the DLM commands require the separate GroundAnything weights. | |
| ### GroundAnything: Entropy-guided Service | |
| Download the **complete DLM model bundle**, including its custom code and tokenizer, then launch: | |
| ```bash | |
| hf download GroundingPI/GroundAnything --local-dir weights/dlm_bundle | |
| python3 run.py serve --decoder denoise | |
| ``` | |
| The published model package is already a DLM bundle, so `prepare-model` is unnecessary for this download. The endpoint is `http://127.0.0.1:8101/v1`, with model ID `groundinganything`. | |
| ### GroundAnything: Self-speculative Service | |
| Stop the existing DLM service before switching its decoder: | |
| ```bash | |
| python3 run.py serve --decoder speculative | |
| ``` | |
| This reuses the same DLM weights and endpoint, with diffusion proposals and greedy causal verification. | |
| ### GroundAnything-VLM: Autoregressive Service | |
| Use the separate autoregressive checkpoint and its service recipe: | |
| ```bash | |
| hf download GroundingPI/GroundAnything-VLM --local-dir weights/vlm | |
| python3 run.py serve --config configs/release/vlm_sglang.yaml | |
| ``` | |
| This service uses `http://127.0.0.1:8102/v1`, with model ID `groundinganything-vlm`. The two repositories share the spatial interface but have different loading and generation paths. | |
| ### Worker (recommended) | |
| Start the appropriate service once, then reuse the client: | |
| ```python | |
| from grounding_anything import GroundingAnything, visualize | |
| client = GroundingAnything( | |
| base_url="http://127.0.0.1:8101/v1", | |
| model="groundinganything", | |
| ) | |
| result = client.predict("example.jpg", "the red car", task="bbox") | |
| print(result.to_dict()) | |
| if result.valid: | |
| visualize("example.jpg", result).save("prediction.png") | |
| point_result = client.predict( | |
| "example.jpg", "the center of the red car", task="point" | |
| ) | |
| ``` | |
| The HTTP client retains no model weights. | |
| ### Supported Tasks & Prompt Templates | |
| | Task | Prompt example | | |
| |:---|:---| | |
| | Category / dense grounding | `Locate all the instances that match the following categories: car</c>person.` | | |
| | Referring boxes | `Locate the target referred to by the following description: the red car.` | | |
| | Category points | `Point to: car</c>person.` | | |
| | Referring points | `Point to the target referred to by the following description: the red car.` | | |
| | OCR | `OCR task detect all the text in box format.` | | |
| | Layout | `Detect all document layout elements that match the following categories: title</c>text.` | | |
| | GUI | `Point to the UI element to click for the following instruction: open the settings menu.` | | |
| | Visual prompting | Provide reference boxes in the native spatial-token format, then request similar objects. | | |
| ```text | |
| Given reference boxes <|box_start|><100><200><500><650><|box_end|> indicating one or more objects, find all similar objects in the image and output their bounding boxes. | |
| ``` | |
| The convenience client's `predict()` method wraps referring-box and referring-point prompts. Use an OpenAI-compatible request to `/chat/completions` for the other task templates; keep the image and prompt in the same user message. | |
| ### Generation Modes | |
| #### Entropy-guided decoding | |
| Image/query prefill creates a causal prefix cache and a known anchor. The release recipe uses **block size 32, sub-block size 4, and entropy threshold 0.8**. A physical block contains the known anchor and 31 masked positions; sub-blocks are completed from left to right while the forward pass evaluates the physical block. | |
| For each still-masked position in the active sub-block, the decoder measures entropy from the unmodified token distribution over the generatable vocabulary, excluding the mask token. Positions at or below the threshold are committed together. If no position qualifies, the lowest-entropy position is committed so decoding makes progress. Committed tokens remain fixed. | |
| After a block is complete, a causal forward reconstructs its authoritative KV cache and supplies the next anchor. This cache-building pass **does not verify or reject the generated block**. A block requiring D denoising passes therefore uses D + 1 model forwards, excluding the initial prefill. | |
| #### Self-speculative decoding | |
| The model uses its own shared weights to draft tokens with bidirectional attention and verify them with causal attention. Verification accepts the **longest consecutive matching prefix**, stops at the first mismatch, applies the causal correction, and discards the rejected suffix cache states. The shipped speculative route uses greedy verification; it is not a general stochastic speculative sampler. | |
| <p align="center"><img src="https://huggingface.co/GroundingPI/GroundAnything/resolve/0f8e30894c3ca86378d01ae51ec69c217c78151b/assets/fig6-self-speculative-decoding.png" width="100%" alt="GroundAnything: linear and quadratic self-speculative schedules with shared model weights" /></p> | |
| The documented `--decoder speculative` service is the linear shared-weight route. Exact greedy verification is relative to the converted model's causal branch; it does not imply identical outputs to the separately trained GroundAnything-VLM checkpoint. | |
| ## Inference Infrastructure | |
| ### SGLang Execution | |
| The custom SGLang integration coordinates model loading, request scheduling, attention kernels, and KV-cache ownership for blockwise generation. Denoising uses bidirectional attention inside the active block; completed history is retained as causal KV. Self-speculation additionally verifies proposals and removes rejected suffix states. These cache semantics must be preserved when optimizing execution. | |
| The supplied recipe selects **BF16 + Triton attention + eager execution**. DLM recipes record the engine source, decoder, effective settings, and package versions in `outputs/sglang/<decoder>/engine_runtime.json`. The VLM service records its causal model/runtime configuration separately in `outputs/sglang/vlm/engine_runtime.json`. For higher service concurrency, use independent replicas with distinct devices, ports, and output directories; the default queue is not continuous multi-request batching. | |
| ### CUDA Graph and Selective FP8 | |
| The paper evaluates a progressive infrastructure sequence: | |
| | Layer | Purpose | Status in the supplied default recipe | | |
| |:---|:---|:---| | |
| | Native PyTorch eager | Reference model execution | Reference implementation | | |
| | SGLang eager | Integrated scheduling, attention, and cache execution | **Default serving path** | | |
| | CUDA Graph replay | Replay compatible captured GPU work to reduce repeated launch overhead | Evaluated in the paper; **disabled by the default launcher** | | |
| | Selective FP8 | Reduce arithmetic cost in eligible language-model linear operations while retaining other components in BF16 | Evaluated in the paper; **default serving remains BF16** | | |
| CUDA Graph changes how compatible GPU work is submitted; it does not define a new token-commitment or verification rule. Captured shapes and state updates must remain compatible with the active decoding path. Selective FP8 can change logits and subsequent decoding decisions, so it is a distinct numerical configuration. | |
| The optional graph implementation captures fixed-shape block work. Its FlashInfer path uses persistent attention masks and device-resident buffer updates for bidirectional drafting and causal verification. The Triton verifier path separates draft/verification metadata and input buffers while sharing parameters and the real KV pool. Its supported capture case is B32 with one request and no tensor/pipeline/data parallel expansion; prefill and unsupported shapes use eager execution. Optional shadow checks compare cache states, logits, and token choices against eager execution. These implementation paths are not exposed as a public `run.py serve --cuda-graph` switch in this release. | |
| The paper's cumulative infrastructure speedups compare execution implementations **within the same decoding mode**. They are separate from the headline comparison of self-speculation against the autoregressive checkpoint. The default release commands above do not enable Graph replay or FP8, and no unsupported activation flags are implied. | |
| ## Ethical Considerations: | |
| Grounding predictions can miss small or occluded objects, repeat instances, or produce inaccurate text and coordinates. Validate localization quality on the intended task and inspect incomplete responses. GUI points describe image locations; the model does not execute interface actions. Perception outputs require task-specific validation before integration into physical systems. | |