How to use from
Docker Model Runner
docker model run hf.co/GroundingPI/GroundAnything
Quick Links

GroundAnything logo GroundAnything: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed

Model family: GroundAnything — DLM / parallel decoding · GroundAnything-VLM — autoregressive.

This repository contains the GroundAnything DLM checkpoint. Use entropy-guided decoding for the main benchmark setting, or select optional self-speculative decoding.

GroundAnything: broad visual grounding and parallel visual evidence extraction

🔗 Quick Links

Model Overview

Description:

Autoregressive (AR) grounding models serialize spatial predictions, introducing sequential latency and imposing a causal order on output tokens. We view grounding as visual evidence extraction: objects, locations, and spatial relations are jointly constrained by the image and query, yet their dependencies do not imply an intrinsic left-to-right generation order. This distinction makes bidirectional diffusion a natural fit, allowing spatial hypotheses to emerge in parallel and be jointly refined through iterative denoising.

We introduce GroundAnything, a 4B-parameter grounding foundation model that reconciles fast parallel decoding with precise localization through blockwise denoising. Across 30 grounding benchmarks, the autoregressive variant, GroundAnything-VLM, establishes a new overall state of the art among similarly sized models at 72.42%, remaining competitive with GPT-6 Astra (71.35%). With entropy-guided decoding, GroundAnything also surpasses the prior state of the art at this scale, averaging 61.75% versus 53.32% for the fast MTP-based LocateAnything model.

An optional self-speculative mode achieves a 4.51× speedup over the AR counterpart with a 0.74 percentage-point drop in COCO F1mIoU at the paper's reported operating point. Infrastructure experiments show that progressive inference optimizations translate parallel decoding into practical speedups. These support efficient visual grounding in latency-sensitive real-world systems.

Demo Videos

Parallel Decoding

License/Terms of Use:

Our original contributions are available under Apache 2.0, with no additional restrictions imposed by this project. Third-party material retains its applicable licenses, including the Kimi K3 License for Kimi-derived material and applicable derivative works. Its conditions continue to apply when using or redistributing the combined model package. See LICENSE for the scope and full license texts.

Deployment Geography:

Global.

Use Case:

  • Open-vocabulary object localization and dense-scene grounding.
  • Referring-expression comprehension and visual-prompt-based localization.
  • Point-based localization, spatial reasoning, and GUI element grounding.
  • OCR with text localization and document layout understanding.
  • Perception research for robotics, embodied agents, and autonomous systems.

Release Date:

References(s):

Citation
@misc{yu2026groundanything,
  title = {GroundAnything: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed},
  author = {Qize Yu and Lianrui Fan and Bowen Ping and Xini Ding and Zetian Song and Junbo Niu and Kaixuan Wang and Tianxing Chen and Yue Chen and Minghua He and Yuran Wang and Jie Huang and Haojun Zhang and Min Chen and Hao Li and Wenxuan Song and Ruihai Wu and Xianming Liu and Shilong Liu and Shuchang Zhou and Ping Luo and Shiyu Huang},
  year = {2026},
  eprint = {2609.39600},
  archivePrefix = {arXiv},
  primaryClass = {cs.CV},
  url = {https://arxiv.org/abs/2609.39600},
}

Model Architecture:

Architecture type: a shared vision-language backbone supporting an autoregressive checkpoint and a blockwise diffusion checkpoint.

  • Vision encoder: MoonViT-V2 / Kimi-K3 vision backbone.
  • Language backbone: Qwen3-4B-Instruct-2507.
  • Multimodal projector: 2 × 2 spatial aggregation and a two-layer MLP.
  • Spatial vocabulary: 1,000 coordinate tokens shared with semantic labels and protocol markers.
  • DLM conversion: the shared decoder and vocabulary head support both causal prediction and bidirectional response-block denoising. A mask token is added for diffusion generation.

GroundAnything: vision-language architecture and autoregressive-to-diffusion conversion

Input(s):

Input types: image and text.

  • Image: one RGB image per request. The Python client accepts JPEG, PNG, and WebP files; the supplied processor handles image preparation.
  • Text: a natural-language instruction, category list, referring expression, OCR/layout query, or a prompt containing example boxes.
  • Image encoding for the HTTP API: an image_url content part containing a base64 data URI, alongside a text content part.

Use the checkpoint's own tokenizer, processor, and chat template. Multiple categories are separated by </c>. Reference boxes use the same 0–999 spatial-token vocabulary as outputs.

Output(s):

Output type: text containing semantic labels and quantized spatial coordinates.

All three released models use the GAM protocol: integer coordinate tokens <0> through <999>, object-reference delimiters, and box delimiters. A bounding box contains (x1, y1, x2, y2); a point contains (x, y). Adjacent coordinate tokens have no intervening spaces. Multiple instances of the same label are comma-separated inside one box wrapper. A missing target is represented by None.

Illustrative syntax:

<|object_ref_start|>car<|object_ref_end|><|box_start|><100><200><500><650><|box_end|>
<|object_ref_start|>car center<|object_ref_end|><|box_start|><300><425><|box_end|>
<|object_ref_start|>absent object<|object_ref_end|><|box_start|>None<|box_end|>

The client returns parsed predictions together with raw_output, finish_reason, usage, parse_error, and valid. It maps coordinates back to image pixels for visualization. A truncated or malformed response is marked invalid. For custom API clients, preserve spatial tokens with skip_special_tokens=false and avoid inserting spaces between them.

Software Integration:

Runtime engines: the repository's custom SGLang integration for the DLM and VLM services, plus a native Transformers reference route for the DLM.

Default serving environment: Linux and Python 3.12 with the source package's serving profile. Use python3 run.py setup serve to install the bundled custom engine and its pinned dependencies; installing upstream SGLang alone does not provide the same model and decoding integration.

The serving profile pins Torch 2.9.1, Transformers 5.5.4, Triton 3.5.1, sgl-kernel 0.3.20, and FlashInfer 0.5.3. Serving and evaluation use separate environments.

Tested hardware: NVIDIA B300, B200, H200, H800, and PPU. The provided SGLang serving installer is a CUDA/GPU recipe; PPU requires its matching platform runtime and is not selected by this GPU installer.

The default DLM service uses BF16, Triton attention, eager execution, one GPU, one active request, and two queued requests. Client concurrency queues requests; it does not imply a multi-request model batch. CUDA Graph and selective FP8 are discussed below as separate infrastructure experiments.

Model Version(s):

Checkpoint Generation Evaluation mode
GroundAnything Entropy-guided blockwise diffusion; optional self-speculation GAM
GroundAnything-VLM Autoregressive generation GAM

GroundAnything's main results use entropy-guided decoding; GroundAnything-VLM uses autoregressive decoding. Self-speculative decoding is an optional acceleration mode.

Evaluation

The evaluation toolkit supports 7 modes:

Mode Supported models
GAM GroundingPI, GroundAnything, GroundAnything-VLM
VLM Generic vision-language baselines
REXOMNI Rex-Omni
LOCATEANYTHING LocateAnything
GROUNDINGDINO GroundingDINO through a compatible service
DLM Legacy diffusion checkpoints using the GAM protocol
RLV2 Legacy RL checkpoints using the GAM protocol

All three released checkpoints use GAM mode.

Evaluation code and instructions: GitHub.

Quantitative Evaluation Benchmarks

GroundAnything: GroundAnything and GroundAnything-VLM benchmark overview

Inference:

Installation

Run the commands from the GroundAnything source repository root after obtaining the code package, using Linux x86_64 and Python 3.12. Use the serving profile in the supplied source bundle so that the custom model adapter, decoding implementation, and dependencies remain aligned.

python3 -m pip install -r requirements.txt huggingface_hub
python3 run.py setup serve

The installer verifies and extracts its bundled frameworks, creates .venv-serve, and records resolved packages. It prepares the custom SGLang implementation and applies the serving profile's dependency order. The profile includes the required cuDNN compatibility selection; its dependency-check report records the known Torch/cuDNN metadata exception.

Choose the checkpoint-specific recipe below. GroundAnything-VLM users need only the VLM download and launch commands; the DLM commands require the separate GroundAnything weights.

GroundAnything: Entropy-guided Service

Download the complete DLM model bundle, including its custom code and tokenizer, then launch:

hf download GroundingPI/GroundAnything --local-dir weights/dlm_bundle
python3 run.py serve --decoder denoise

The published model package is already a DLM bundle, so prepare-model is unnecessary for this download. The endpoint is http://127.0.0.1:8101/v1, with model ID groundinganything.

GroundAnything: Self-speculative Service

Stop the existing DLM service before switching its decoder:

python3 run.py serve --decoder speculative

This reuses the same DLM weights and endpoint, with diffusion proposals and greedy causal verification.

GroundAnything-VLM: Autoregressive Service

Use the separate autoregressive checkpoint and its service recipe:

hf download GroundingPI/GroundAnything-VLM --local-dir weights/vlm
python3 run.py serve --config configs/release/vlm_sglang.yaml

This service uses http://127.0.0.1:8102/v1, with model ID groundinganything-vlm. The two repositories share the spatial interface but have different loading and generation paths.

Worker (recommended)

Start the appropriate service once, then reuse the client:

from grounding_anything import GroundingAnything, visualize

client = GroundingAnything(
    base_url="http://127.0.0.1:8101/v1",
    model="groundinganything",
)
result = client.predict("example.jpg", "the red car", task="bbox")
print(result.to_dict())
if result.valid:
    visualize("example.jpg", result).save("prediction.png")

point_result = client.predict(
    "example.jpg", "the center of the red car", task="point"
)

The HTTP client retains no model weights.

Supported Tasks & Prompt Templates

Task Prompt example
Category / dense grounding Locate all the instances that match the following categories: car</c>person.
Referring boxes Locate the target referred to by the following description: the red car.
Category points Point to: car</c>person.
Referring points Point to the target referred to by the following description: the red car.
OCR OCR task detect all the text in box format.
Layout Detect all document layout elements that match the following categories: title</c>text.
GUI Point to the UI element to click for the following instruction: open the settings menu.
Visual prompting Provide reference boxes in the native spatial-token format, then request similar objects.
Given reference boxes <|box_start|><100><200><500><650><|box_end|> indicating one or more objects, find all similar objects in the image and output their bounding boxes.

The convenience client's predict() method wraps referring-box and referring-point prompts. Use an OpenAI-compatible request to /chat/completions for the other task templates; keep the image and prompt in the same user message.

Generation Modes

Entropy-guided decoding

Image/query prefill creates a causal prefix cache and a known anchor. The release recipe uses block size 32, sub-block size 4, and entropy threshold 0.8. A physical block contains the known anchor and 31 masked positions; sub-blocks are completed from left to right while the forward pass evaluates the physical block.

For each still-masked position in the active sub-block, the decoder measures entropy from the unmodified token distribution over the generatable vocabulary, excluding the mask token. Positions at or below the threshold are committed together. If no position qualifies, the lowest-entropy position is committed so decoding makes progress. Committed tokens remain fixed.

After a block is complete, a causal forward reconstructs its authoritative KV cache and supplies the next anchor. This cache-building pass does not verify or reject the generated block. A block requiring D denoising passes therefore uses D + 1 model forwards, excluding the initial prefill.

Self-speculative decoding

The model uses its own shared weights to draft tokens with bidirectional attention and verify them with causal attention. Verification accepts the longest consecutive matching prefix, stops at the first mismatch, applies the causal correction, and discards the rejected suffix cache states. The shipped speculative route uses greedy verification; it is not a general stochastic speculative sampler.

GroundAnything: linear and quadratic self-speculative schedules with shared model weights

The documented --decoder speculative service is the linear shared-weight route. Exact greedy verification is relative to the converted model's causal branch; it does not imply identical outputs to the separately trained GroundAnything-VLM checkpoint.

Inference Infrastructure

SGLang Execution

The custom SGLang integration coordinates model loading, request scheduling, attention kernels, and KV-cache ownership for blockwise generation. Denoising uses bidirectional attention inside the active block; completed history is retained as causal KV. Self-speculation additionally verifies proposals and removes rejected suffix states. These cache semantics must be preserved when optimizing execution.

The supplied recipe selects BF16 + Triton attention + eager execution. DLM recipes record the engine source, decoder, effective settings, and package versions in outputs/sglang/<decoder>/engine_runtime.json. The VLM service records its causal model/runtime configuration separately in outputs/sglang/vlm/engine_runtime.json. For higher service concurrency, use independent replicas with distinct devices, ports, and output directories; the default queue is not continuous multi-request batching.

CUDA Graph and Selective FP8

The paper evaluates a progressive infrastructure sequence:

Layer Purpose Status in the supplied default recipe
Native PyTorch eager Reference model execution Reference implementation
SGLang eager Integrated scheduling, attention, and cache execution Default serving path
CUDA Graph replay Replay compatible captured GPU work to reduce repeated launch overhead Evaluated in the paper; disabled by the default launcher
Selective FP8 Reduce arithmetic cost in eligible language-model linear operations while retaining other components in BF16 Evaluated in the paper; default serving remains BF16

CUDA Graph changes how compatible GPU work is submitted; it does not define a new token-commitment or verification rule. Captured shapes and state updates must remain compatible with the active decoding path. Selective FP8 can change logits and subsequent decoding decisions, so it is a distinct numerical configuration.

The optional graph implementation captures fixed-shape block work. Its FlashInfer path uses persistent attention masks and device-resident buffer updates for bidirectional drafting and causal verification. The Triton verifier path separates draft/verification metadata and input buffers while sharing parameters and the real KV pool. Its supported capture case is B32 with one request and no tensor/pipeline/data parallel expansion; prefill and unsupported shapes use eager execution. Optional shadow checks compare cache states, logits, and token choices against eager execution. These implementation paths are not exposed as a public run.py serve --cuda-graph switch in this release.

The paper's cumulative infrastructure speedups compare execution implementations within the same decoding mode. They are separate from the headline comparison of self-speculation against the autoregressive checkpoint. The default release commands above do not enable Graph replay or FP8, and no unsupported activation flags are implied.

Ethical Considerations:

Grounding predictions can miss small or occluded objects, repeat instances, or produce inaccurate text and coordinates. Validate localization quality on the intended task and inspect incomplete responses. GUI points describe image locations; the model does not execute interface actions. Perception outputs require task-specific validation before integration into physical systems.

Downloads last month
-
Safetensors
Model size
5B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Papers for GroundingPI/GroundAnything