Instructions to use GroundingPI/GroundAnything with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use GroundingPI/GroundAnything with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="GroundingPI/GroundAnything", trust_remote_code=True) messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("GroundingPI/GroundAnything", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use GroundingPI/GroundAnything with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "GroundingPI/GroundAnything" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "GroundingPI/GroundAnything", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/GroundingPI/GroundAnything
- SGLang
How to use GroundingPI/GroundAnything with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "GroundingPI/GroundAnything" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "GroundingPI/GroundAnything", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "GroundingPI/GroundAnything" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "GroundingPI/GroundAnything", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use GroundingPI/GroundAnything with Docker Model Runner:
docker model run hf.co/GroundingPI/GroundAnything
GroundAnything: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed - Model Overview
GroundAnything: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed
Model family: GroundAnything — DLM / parallel decoding · GroundAnything-VLM — autoregressive.
This repository contains the GroundAnything DLM checkpoint. Use entropy-guided decoding for the main benchmark setting, or select optional self-speculative decoding.

🔗 Quick Links
- 🚀 Online Demo: Coming soon — XXX.
- 💻 GitHub Code: groundingpi/GroundAnything.
- 📄 Paper: arXiv:2609.39600.
- 🧪 Evaluation data: Coming soon — XXX.
Model Overview
Description:
Autoregressive (AR) grounding models serialize spatial predictions, introducing sequential latency and imposing a causal order on output tokens. We view grounding as visual evidence extraction: objects, locations, and spatial relations are jointly constrained by the image and query, yet their dependencies do not imply an intrinsic left-to-right generation order. This distinction makes bidirectional diffusion a natural fit, allowing spatial hypotheses to emerge in parallel and be jointly refined through iterative denoising.
We introduce GroundAnything, a 4B-parameter grounding foundation model that reconciles fast parallel decoding with precise localization through blockwise denoising. Across 30 grounding benchmarks, the autoregressive variant, GroundAnything-VLM, establishes a new overall state of the art among similarly sized models at 72.42%, remaining competitive with GPT-6 Astra (71.35%). With entropy-guided decoding, GroundAnything also surpasses the prior state of the art at this scale, averaging 61.75% versus 53.32% for the fast MTP-based LocateAnything model.
An optional self-speculative mode achieves a 4.51× speedup over the AR counterpart with a 0.74 percentage-point drop in COCO F1mIoU at the paper's reported operating point. Infrastructure experiments show that progressive inference optimizations translate parallel decoding into practical speedups. These support efficient visual grounding in latency-sensitive real-world systems.
Demo Videos
Parallel Decoding
License/Terms of Use:
Our original contributions are available under Apache 2.0, with no additional restrictions imposed by this project. Third-party material retains its applicable licenses, including the Kimi K3 License for Kimi-derived material and applicable derivative works. Its conditions continue to apply when using or redistributing the combined model package. See LICENSE for the scope and full license texts.
Deployment Geography:
Global.
Use Case:
- Open-vocabulary object localization and dense-scene grounding.
- Referring-expression comprehension and visual-prompt-based localization.
- Point-based localization, spatial reasoning, and GUI element grounding.
- OCR with text localization and document layout understanding.
- Perception research for robotics, embodied agents, and autonomous systems.
Release Date:
- Paper [09/30/2026]: GroundAnything.
References(s):
Citation
@misc{yu2026groundanything,
title = {GroundAnything: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed},
author = {Qize Yu and Lianrui Fan and Bowen Ping and Xini Ding and Zetian Song and Junbo Niu and Kaixuan Wang and Tianxing Chen and Yue Chen and Minghua He and Yuran Wang and Jie Huang and Haojun Zhang and Min Chen and Hao Li and Wenxuan Song and Ruihai Wu and Xianming Liu and Shilong Liu and Shuchang Zhou and Ping Luo and Shiyu Huang},
year = {2026},
eprint = {2609.39600},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
url = {https://arxiv.org/abs/2609.39600},
}
Model Architecture:
Architecture type: a shared vision-language backbone supporting an autoregressive checkpoint and a blockwise diffusion checkpoint.
- Vision encoder: MoonViT-V2 / Kimi-K3 vision backbone.
- Language backbone: Qwen3-4B-Instruct-2507.
- Multimodal projector: 2 × 2 spatial aggregation and a two-layer MLP.
- Spatial vocabulary: 1,000 coordinate tokens shared with semantic labels and protocol markers.
- DLM conversion: the shared decoder and vocabulary head support both causal prediction and bidirectional response-block denoising. A mask token is added for diffusion generation.

Input(s):
Input types: image and text.
- Image: one RGB image per request. The Python client accepts JPEG, PNG, and WebP files; the supplied processor handles image preparation.
- Text: a natural-language instruction, category list, referring expression, OCR/layout query, or a prompt containing example boxes.
- Image encoding for the HTTP API: an
image_urlcontent part containing a base64 data URI, alongside atextcontent part.
Use the checkpoint's own tokenizer, processor, and chat template. Multiple categories are separated by </c>. Reference boxes use the same 0–999 spatial-token vocabulary as outputs.
Output(s):
Output type: text containing semantic labels and quantized spatial coordinates.
All three released models use the GAM protocol: integer coordinate tokens <0> through <999>, object-reference delimiters, and box delimiters. A bounding box contains (x1, y1, x2, y2); a point contains (x, y). Adjacent coordinate tokens have no intervening spaces. Multiple instances of the same label are comma-separated inside one box wrapper. A missing target is represented by None.
Illustrative syntax:
<|object_ref_start|>car<|object_ref_end|><|box_start|><100><200><500><650><|box_end|>
<|object_ref_start|>car center<|object_ref_end|><|box_start|><300><425><|box_end|>
<|object_ref_start|>absent object<|object_ref_end|><|box_start|>None<|box_end|>
The client returns parsed predictions together with raw_output, finish_reason, usage, parse_error, and valid. It maps coordinates back to image pixels for visualization. A truncated or malformed response is marked invalid. For custom API clients, preserve spatial tokens with skip_special_tokens=false and avoid inserting spaces between them.
Software Integration:
Runtime engines: the repository's custom SGLang integration for the DLM and VLM services, plus a native Transformers reference route for the DLM.
Default serving environment: Linux and Python 3.12 with the source package's serving profile. Use python3 run.py setup serve to install the bundled custom engine and its pinned dependencies; installing upstream SGLang alone does not provide the same model and decoding integration.
The serving profile pins Torch 2.9.1, Transformers 5.5.4, Triton 3.5.1, sgl-kernel 0.3.20, and FlashInfer 0.5.3. Serving and evaluation use separate environments.
Tested hardware: NVIDIA B300, B200, H200, H800, and PPU. The provided SGLang serving installer is a CUDA/GPU recipe; PPU requires its matching platform runtime and is not selected by this GPU installer.
The default DLM service uses BF16, Triton attention, eager execution, one GPU, one active request, and two queued requests. Client concurrency queues requests; it does not imply a multi-request model batch. CUDA Graph and selective FP8 are discussed below as separate infrastructure experiments.
Model Version(s):
| Checkpoint | Generation | Evaluation mode |
|---|---|---|
| GroundAnything | Entropy-guided blockwise diffusion; optional self-speculation | GAM |
| GroundAnything-VLM | Autoregressive generation | GAM |
GroundAnything's main results use entropy-guided decoding; GroundAnything-VLM uses autoregressive decoding. Self-speculative decoding is an optional acceleration mode.
Evaluation
The evaluation toolkit supports 7 modes:
| Mode | Supported models |
|---|---|
GAM |
GroundingPI, GroundAnything, GroundAnything-VLM |
VLM |
Generic vision-language baselines |
REXOMNI |
Rex-Omni |
LOCATEANYTHING |
LocateAnything |
GROUNDINGDINO |
GroundingDINO through a compatible service |
DLM |
Legacy diffusion checkpoints using the GAM protocol |
RLV2 |
Legacy RL checkpoints using the GAM protocol |
All three released checkpoints use GAM mode.
Evaluation code and instructions: GitHub.
Quantitative Evaluation Benchmarks

Inference:
Installation
Run the commands from the GroundAnything source repository root after obtaining the code package, using Linux x86_64 and Python 3.12. Use the serving profile in the supplied source bundle so that the custom model adapter, decoding implementation, and dependencies remain aligned.
python3 -m pip install -r requirements.txt huggingface_hub
python3 run.py setup serve
The installer verifies and extracts its bundled frameworks, creates .venv-serve, and records resolved packages. It prepares the custom SGLang implementation and applies the serving profile's dependency order. The profile includes the required cuDNN compatibility selection; its dependency-check report records the known Torch/cuDNN metadata exception.
Choose the checkpoint-specific recipe below. GroundAnything-VLM users need only the VLM download and launch commands; the DLM commands require the separate GroundAnything weights.
GroundAnything: Entropy-guided Service
Download the complete DLM model bundle, including its custom code and tokenizer, then launch:
hf download GroundingPI/GroundAnything --local-dir weights/dlm_bundle
python3 run.py serve --decoder denoise
The published model package is already a DLM bundle, so prepare-model is unnecessary for this download. The endpoint is http://127.0.0.1:8101/v1, with model ID groundinganything.
GroundAnything: Self-speculative Service
Stop the existing DLM service before switching its decoder:
python3 run.py serve --decoder speculative
This reuses the same DLM weights and endpoint, with diffusion proposals and greedy causal verification.
GroundAnything-VLM: Autoregressive Service
Use the separate autoregressive checkpoint and its service recipe:
hf download GroundingPI/GroundAnything-VLM --local-dir weights/vlm
python3 run.py serve --config configs/release/vlm_sglang.yaml
This service uses http://127.0.0.1:8102/v1, with model ID groundinganything-vlm. The two repositories share the spatial interface but have different loading and generation paths.
Worker (recommended)
Start the appropriate service once, then reuse the client:
from grounding_anything import GroundingAnything, visualize
client = GroundingAnything(
base_url="http://127.0.0.1:8101/v1",
model="groundinganything",
)
result = client.predict("example.jpg", "the red car", task="bbox")
print(result.to_dict())
if result.valid:
visualize("example.jpg", result).save("prediction.png")
point_result = client.predict(
"example.jpg", "the center of the red car", task="point"
)
The HTTP client retains no model weights.
Supported Tasks & Prompt Templates
| Task | Prompt example |
|---|---|
| Category / dense grounding | Locate all the instances that match the following categories: car</c>person. |
| Referring boxes | Locate the target referred to by the following description: the red car. |
| Category points | Point to: car</c>person. |
| Referring points | Point to the target referred to by the following description: the red car. |
| OCR | OCR task detect all the text in box format. |
| Layout | Detect all document layout elements that match the following categories: title</c>text. |
| GUI | Point to the UI element to click for the following instruction: open the settings menu. |
| Visual prompting | Provide reference boxes in the native spatial-token format, then request similar objects. |
Given reference boxes <|box_start|><100><200><500><650><|box_end|> indicating one or more objects, find all similar objects in the image and output their bounding boxes.
The convenience client's predict() method wraps referring-box and referring-point prompts. Use an OpenAI-compatible request to /chat/completions for the other task templates; keep the image and prompt in the same user message.
Generation Modes
Entropy-guided decoding
Image/query prefill creates a causal prefix cache and a known anchor. The release recipe uses block size 32, sub-block size 4, and entropy threshold 0.8. A physical block contains the known anchor and 31 masked positions; sub-blocks are completed from left to right while the forward pass evaluates the physical block.
For each still-masked position in the active sub-block, the decoder measures entropy from the unmodified token distribution over the generatable vocabulary, excluding the mask token. Positions at or below the threshold are committed together. If no position qualifies, the lowest-entropy position is committed so decoding makes progress. Committed tokens remain fixed.
After a block is complete, a causal forward reconstructs its authoritative KV cache and supplies the next anchor. This cache-building pass does not verify or reject the generated block. A block requiring D denoising passes therefore uses D + 1 model forwards, excluding the initial prefill.
Self-speculative decoding
The model uses its own shared weights to draft tokens with bidirectional attention and verify them with causal attention. Verification accepts the longest consecutive matching prefix, stops at the first mismatch, applies the causal correction, and discards the rejected suffix cache states. The shipped speculative route uses greedy verification; it is not a general stochastic speculative sampler.

The documented --decoder speculative service is the linear shared-weight route. Exact greedy verification is relative to the converted model's causal branch; it does not imply identical outputs to the separately trained GroundAnything-VLM checkpoint.
Inference Infrastructure
SGLang Execution
The custom SGLang integration coordinates model loading, request scheduling, attention kernels, and KV-cache ownership for blockwise generation. Denoising uses bidirectional attention inside the active block; completed history is retained as causal KV. Self-speculation additionally verifies proposals and removes rejected suffix states. These cache semantics must be preserved when optimizing execution.
The supplied recipe selects BF16 + Triton attention + eager execution. DLM recipes record the engine source, decoder, effective settings, and package versions in outputs/sglang/<decoder>/engine_runtime.json. The VLM service records its causal model/runtime configuration separately in outputs/sglang/vlm/engine_runtime.json. For higher service concurrency, use independent replicas with distinct devices, ports, and output directories; the default queue is not continuous multi-request batching.
CUDA Graph and Selective FP8
The paper evaluates a progressive infrastructure sequence:
| Layer | Purpose | Status in the supplied default recipe |
|---|---|---|
| Native PyTorch eager | Reference model execution | Reference implementation |
| SGLang eager | Integrated scheduling, attention, and cache execution | Default serving path |
| CUDA Graph replay | Replay compatible captured GPU work to reduce repeated launch overhead | Evaluated in the paper; disabled by the default launcher |
| Selective FP8 | Reduce arithmetic cost in eligible language-model linear operations while retaining other components in BF16 | Evaluated in the paper; default serving remains BF16 |
CUDA Graph changes how compatible GPU work is submitted; it does not define a new token-commitment or verification rule. Captured shapes and state updates must remain compatible with the active decoding path. Selective FP8 can change logits and subsequent decoding decisions, so it is a distinct numerical configuration.
The optional graph implementation captures fixed-shape block work. Its FlashInfer path uses persistent attention masks and device-resident buffer updates for bidirectional drafting and causal verification. The Triton verifier path separates draft/verification metadata and input buffers while sharing parameters and the real KV pool. Its supported capture case is B32 with one request and no tensor/pipeline/data parallel expansion; prefill and unsupported shapes use eager execution. Optional shadow checks compare cache states, logits, and token choices against eager execution. These implementation paths are not exposed as a public run.py serve --cuda-graph switch in this release.
The paper's cumulative infrastructure speedups compare execution implementations within the same decoding mode. They are separate from the headline comparison of self-speculation against the autoregressive checkpoint. The default release commands above do not enable Graph replay or FP8, and no unsupported activation flags are implied.
Ethical Considerations:
Grounding predictions can miss small or occluded objects, repeat instances, or produce inaccurate text and coordinates. Validate localization quality on the intended task and inspect incomplete responses. GUI points describe image locations; the model does not execute interface actions. Perception outputs require task-specific validation before integration into physical systems.
- Downloads last month
- -
docker model run hf.co/GroundingPI/GroundAnything