How to use from the
Use from the
Transformers library
# Use a pipeline as a high-level helper
from transformers import pipeline

pipe = pipeline("image-text-to-text", model="GroundingPI/GroundAnything", trust_remote_code=True)
messages = [
    {
        "role": "user",
        "content": [
            {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"},
            {"type": "text", "text": "What animal is on the candy?"}
        ]
    },
]
pipe(text=messages)
# Load model directly
from transformers import AutoModelForCausalLM
model = AutoModelForCausalLM.from_pretrained("GroundingPI/GroundAnything", trust_remote_code=True, device_map="auto")
Quick Links

GroundAnything logo

GroundAnything: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed

Model family: GroundAnything — DLM / parallel decoding · GroundAnything-VLM — autoregressive.

This repository contains the GroundAnything DLM checkpoint. Use entropy-guided decoding for the main benchmark setting, or select optional self-speculative decoding.

GroundAnything Figure 1: broad visual grounding and parallel visual evidence extraction

🔗 Quick Links

  • 🚀 Online Demo: Coming soon — XXX.
  • 💻 GitHub Code: Coming soon — XXX.
  • 📄 Paper: arXiv:2609.39600.
  • 🧪 Evaluation data: Coming soon — XXX.

Model Overview

Description:

Autoregressive (AR) grounding models serialize spatial predictions, introducing sequential latency and imposing a causal order on output tokens. We view grounding as visual evidence extraction: objects, locations, and spatial relations are jointly constrained by the image and query, yet their dependencies do not imply an intrinsic left-to-right generation order. This distinction makes bidirectional diffusion a natural fit, allowing spatial hypotheses to emerge in parallel and be jointly refined through iterative denoising.

We introduce GroundAnything, a 4B-parameter grounding foundation model that reconciles fast parallel decoding with precise localization through blockwise denoising. Across 30 grounding benchmarks, the autoregressive variant, GroundAnything-VLM, establishes a new overall state of the art among similarly sized models at 72.42%, remaining competitive with GPT-6 Astra (71.35%). With entropy-guided decoding, GroundAnything also surpasses the prior state of the art at this scale, averaging 61.75% versus 53.32% for the fast MTP-based LocateAnything model.

An optional self-speculative mode achieves a 4.51× speedup over the AR counterpart with a 0.74 percentage-point drop in COCO F1mIoU at the paper's reported operating point. Infrastructure experiments show that progressive inference optimizations translate parallel decoding into practical speedups. These support efficient visual grounding in latency-sensitive real-world systems.

Demo Videos

Parallel decoding in action

License/Terms of Use:

This checkpoint is released under the Kimi K3 License; see the repository's LICENSE and the upstream license. The model incorporates Kimi K3-derived vision components and implementation code. The license includes additional conditions for certain commercial uses. Third-party components retain their respective licenses and copyright notices.

Deployment Geography:

Global.

Use Case:

  • Open-vocabulary object localization and dense-scene grounding.
  • Referring-expression comprehension and visual-prompt-based localization.
  • Point-based localization, spatial reasoning, and GUI element grounding.
  • OCR with text localization and document layout understanding.
  • Perception research for robotics, embodied agents, and autonomous systems.

Release Date:

References(s):

Citation
@misc{yu2026groundanything,
  title = {GroundAnything: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed},
  author = {Qize Yu and Lianrui Fan and Bowen Ping and Xini Ding and Zetian Song and Junbo Niu and Kaixuan Wang and Tianxing Chen and Yue Chen and Minghua He and Yuran Wang and Jie Huang and Haojun Zhang and Min Chen and Hao Li and Wenxuan Song and Ruihai Wu and Xianming Liu and Shilong Liu and Shuchang Zhou and Ping Luo and Shiyu Huang},
  year = {2026},
  eprint = {2609.39600},
  archivePrefix = {arXiv},
  primaryClass = {cs.CV},
  url = {https://arxiv.org/abs/2609.39600},
}

Model Architecture:

Architecture type: a shared vision-language backbone supporting an autoregressive checkpoint and a blockwise diffusion checkpoint.

  • Vision encoder: MoonViT-V2 / Kimi-K3 vision backbone.
  • Language backbone: Qwen3-4B-Instruct-2507.
  • Multimodal projector: 2 × 2 spatial aggregation and a two-layer MLP.
  • Spatial vocabulary: 1,000 coordinate tokens shared with semantic labels and protocol markers.
  • DLM conversion: the shared decoder and vocabulary head support both causal prediction and bidirectional response-block denoising. A mask token is added for diffusion generation.

GroundAnything Figure 2: vision-language architecture and autoregressive-to-diffusion conversion

Figure 2. Shared model architecture and the conversion to parallel grounding.

GroundAnything Figure 4: clean-stream causal and noisy-response block attention

Figure 4. The attention mask used during diffusion conversion. The clean stream uses causal attention. A noisy response block attends bidirectionally within its block and reads the clean conditioning context and strictly preceding clean response blocks. Conversion uses B=32; the illustration uses two-token blocks.

Input(s):

Input types: image and text.

  • Image: one RGB image per request. The Python client accepts JPEG, PNG, and WebP files; the supplied processor handles image preparation.
  • Text: a natural-language instruction, category list, referring expression, OCR/layout query, or a prompt containing example boxes.
  • Image encoding for the HTTP API: an image_url content part containing a base64 data URI, alongside a text content part.

Use the checkpoint's own tokenizer, processor, and chat template. Multiple categories are separated by </c>. Reference boxes use the same 0–999 spatial-token vocabulary as outputs.

Output(s):

Output type: text containing semantic labels and quantized spatial coordinates.

All three released models use the GAM protocol: integer coordinate tokens <0> through <999>, object-reference delimiters, and box delimiters. A bounding box contains (x1, y1, x2, y2); a point contains (x, y). Adjacent coordinate tokens have no intervening spaces. Multiple instances of the same label are comma-separated inside one box wrapper. A missing target is represented by None.

Illustrative syntax:

<|object_ref_start|>car<|object_ref_end|><|box_start|><100><200><500><650><|box_end|>
<|object_ref_start|>car center<|object_ref_end|><|box_start|><300><425><|box_end|>
<|object_ref_start|>absent object<|object_ref_end|><|box_start|>None<|box_end|>

The client returns parsed predictions together with raw_output, finish_reason, usage, parse_error, and valid. It maps coordinates back to image pixels for visualization. A truncated or malformed response is marked invalid. For custom API clients, preserve spatial tokens with skip_special_tokens=false and avoid inserting spaces between them.

Software Integration:

Runtime engines: the repository's custom SGLang integration for the DLM and VLM services, plus a native Transformers reference route for the DLM.

Default serving environment: Linux and Python 3.12 with the source package's serving profile. Use python3 run.py setup serve to install the bundled custom engine and its pinned dependencies; installing upstream SGLang alone does not provide the same model and decoding integration.

The serving profile pins Torch 2.9.1, Transformers 5.5.4, Triton 3.5.1, sgl-kernel 0.3.20, and FlashInfer 0.5.3. Serving and evaluation use separate environments.

Tested hardware: NVIDIA B300, B200, H200, H800, and PPU. The provided SGLang serving installer is a CUDA/GPU recipe; PPU requires its matching platform runtime and is not selected by this GPU installer.

The default DLM service uses BF16, Triton attention, eager execution, one GPU, one active request, and two queued requests. Client concurrency queues requests; it does not imply a multi-request model batch. CUDA Graph and selective FP8 are discussed below as separate infrastructure experiments.

Model Version(s):

Checkpoint Generation Evaluation mode
GroundAnything Entropy-guided blockwise diffusion; optional self-speculation GAM
GroundAnything-VLM Autoregressive generation GAM

The main GroundAnything benchmark results use entropy-guided diffusion without autoregressive verification. GroundAnything-VLM uses autoregressive decoding. The self-speculative speed–quality operating point is a separate experiment and should not be substituted for the main diffusion benchmark setting.

Testing and Evaluation Datasets:

Data Modality:

Image and text, with task-specific box, point, text-region, or interaction annotations.

Evaluation Dataset:

Evaluation data: XXX — download URL coming soon.

The shared evaluation toolkit provides 42 task recipes across eight task families: Grounding, Referring, Dense, OCR, Layout, GUI, Pointing, and VisualPrompt. These are executable task/split recipes, not a count of distinct datasets. Dataset paths are registered in configs/datasets.yaml; task IDs are listed in configs/eval/tasks.json.

Evaluation Modes

The evaluator accepts seven mode identifiers. A mode selects the prompt and output parser independently of the inference engine and decoding algorithm.

Mode Supported evaluation interface Output / coordinates
GAM GroundingPI, GroundAnything, and GroundAnything-VLM Native spatial tokens, 0–999
DLM Legacy compatible diffusion-checkpoint alias Same GAM spatial-token protocol
RLV2 Compatible RL-checkpoint alias Same GAM spatial-token protocol
VLM Generic VLM baselines Explicit pixel, 0–1000, or 0–1 coordinate mode
REXOMNI Rex-Omni adapter Model-specific spatial-token parser
LOCATEANYTHING LocateAnything adapter Its own output parser and generation-mode setting
GROUNDINGDINO External compatible GroundingDINO bridge JSON coordinate responses

Use mode: GAM for all three of our checkpoints. The -VLM suffix identifies an autoregressive checkpoint; it does not select the evaluator's generic VLM mode. The DLM's entropy-guided or self-speculative algorithm is selected separately through decoder.

External baseline weights and model services are separate from the evaluation toolkit. Metrics remain task-specific: box localization quality, point accuracy, OCR text-and-region matching, GUI grounding accuracy, and counting error. Full protocols and benchmark tables are provided in the paper and supplementary material.

Quantitative Evaluation Benchmarks

GroundAnything Figure 7: GroundAnything and GroundAnything-VLM benchmark overview

Figure 7. Paper-reported capability overview. Detailed task-level scores, baseline settings, and evaluation protocols are provided in the paper and supplementary material.

Inference:

Installation

Run the commands from the GroundAnything source repository root after obtaining the code package, using Linux x86_64 and Python 3.12. The code link above is reserved for the public release. Use the serving profile in the supplied source bundle so that the custom model adapter, decoding implementation, and dependencies remain aligned.

python3 -m pip install -r requirements.txt huggingface_hub
python3 run.py setup serve

The installer verifies and extracts its bundled frameworks, creates .venv-serve, and records resolved packages. It prepares the custom SGLang implementation and applies the serving profile's dependency order. The profile includes the required cuDNN compatibility selection; its dependency-check report records the known Torch/cuDNN metadata exception.

Choose the checkpoint-specific recipe below. GroundAnything-VLM users need only the VLM download and launch commands; the DLM commands require the separate GroundAnything weights.

GroundAnything: Entropy-guided Service

Download the complete DLM model bundle, including its custom code and tokenizer, then launch:

hf download GroundingPI/GroundAnything --local-dir weights/dlm_bundle
python3 run.py serve --decoder denoise

The published model package is already a DLM bundle, so prepare-model is unnecessary for this download. The endpoint is http://127.0.0.1:8101/v1, with model ID groundinganything.

GroundAnything: Self-speculative Service

Stop the existing DLM service before switching its decoder:

python3 run.py serve --decoder speculative

This reuses the same DLM weights and endpoint, with diffusion proposals and greedy causal verification.

GroundAnything-VLM: Autoregressive Service

Use the separate autoregressive checkpoint and its service recipe:

hf download GroundingPI/GroundAnything-VLM --local-dir weights/vlm
python3 run.py serve --config configs/release/vlm_sglang.yaml

This service uses http://127.0.0.1:8102/v1, with model ID groundinganything-vlm. The two repositories share the spatial interface but have different loading and generation paths.

Worker (recommended)

Start the appropriate service once, then reuse the client:

from grounding_anything import GroundingAnything, visualize

client = GroundingAnything(
    base_url="http://127.0.0.1:8101/v1",
    model="groundinganything",
)
result = client.predict("example.jpg", "the red car", task="bbox")
print(result.to_dict())
if result.valid:
    visualize("example.jpg", result).save("prediction.png")

point_result = client.predict(
    "example.jpg", "the center of the red car", task="point"
)

The HTTP client retains no model weights. A generic interactive request has no benchmark task identity; it does not automatically receive the benchmark's task-specific sampling and stopping policy. Use the evaluator below to reproduce the reported protocol.

Supported Tasks & Prompt Templates

Task Prompt example
Category / dense grounding Locate all the instances that match the following categories: car</c>person.
Referring boxes Locate the target referred to by the following description: the red car.
Category points Point to: car</c>person.
Referring points Point to the target referred to by the following description: the red car.
OCR OCR task detect all the text in box format.
Layout Detect all document layout elements that match the following categories: title</c>text.
GUI Point to the UI element to click for the following instruction: open the settings menu.
Visual prompting Provide reference boxes in the native spatial-token format, then request similar objects.
Given reference boxes <|box_start|><100><200><500><650><|box_end|> indicating one or more objects, find all similar objects in the image and output their bounding boxes.

The convenience client's predict() method wraps referring-box and referring-point prompts. Use an OpenAI-compatible request to /chat/completions for the other task templates; keep the image and prompt in the same user message.

Generation Modes

Entropy-guided decoding

Image/query prefill creates a causal prefix cache and a known anchor. The release recipe uses block size 32, sub-block size 4, and entropy threshold 0.8. A physical block contains the known anchor and 31 masked positions; sub-blocks are completed from left to right while the forward pass evaluates the physical block.

For each still-masked position in the active sub-block, the decoder measures entropy from the unmodified token distribution over the generatable vocabulary, excluding the mask token. Positions at or below the threshold are committed together. If no position qualifies, the lowest-entropy position is committed so decoding makes progress. Committed tokens remain fixed.

After a block is complete, a causal forward reconstructs its authoritative KV cache and supplies the next anchor. This cache-building pass does not verify or reject the generated block. A block requiring D denoising passes therefore uses D + 1 model forwards, excluding the initial prefill.

Self-speculative decoding

The model uses its own shared weights to draft tokens with bidirectional attention and verify them with causal attention. Verification accepts the longest consecutive matching prefix, stops at the first mismatch, applies the causal correction, and discards the rejected suffix cache states. The shipped speculative route uses greedy verification; it is not a general stochastic speculative sampler.

GroundAnything Figure 6: linear and quadratic self-speculative schedules with shared model weights

Figure 6. The paper studies two schedules: linear drafting and verification use two model forwards and 2B query tokens per round; quadratic fusion uses one forward after initialization with B(B + 1) query tokens. These counts describe queries and model calls, not total Transformer FLOPs.

The documented --decoder speculative service is the linear shared-weight route. Exact greedy verification is relative to the converted model's causal branch; it does not imply identical outputs to the separately trained GroundAnything-VLM checkpoint.

Benchmark Sampling Policy

Entropy-guided evaluation applies the five task profiles from infer/decode/configs/task_profiles.json:

Task profile Recipes Temperature Top-p Output-token budget
Strict single target 14 0 1 512
Medium, non-OCR 12 0.1 0.95 4,096
Dense, non-OCR 8 0.1 0.95 8,192
Medium OCR 4 0.3 0.95 4,096
Dense OCR 4 0 1 4,096

The evaluator sends outer sampling parameters and the nested custom_params.gam_dlm_decode contract, adds the strict-single-target policy when required, and checks the server's actual decoder before sending requests. Unknown task profiles and mismatched decoders fail early. The entropy-guided benchmark policy rejects a global max_tokens override because it would replace the task budgets.

Self-speculative evaluation uses greedy sampling (temperature=0, top_p=1, repetition penalty 1). GroundAnything-VLM retains its own autoregressive task recipes. All of these routes use GAM prompts and spatial-token parsing.

Run Evaluation

python3 run.py setup eval

# Match the command to the service that is already running.
python3 run.py eval --decoder denoise
python3 run.py eval --decoder speculative
python3 run.py eval --config configs/eval/vlm.yaml

Choose one evaluation command for the intended checkpoint/decoder:

Checkpoint / decoder Recipe Mode Endpoint model ID
GroundAnything / entropy-guided configs/eval/dlm.yaml GAM groundinganything
GroundAnything / self-speculative configs/eval/dlm_speculative.yaml GAM groundinganything
GroundAnything-VLM / autoregressive configs/eval/vlm.yaml GAM groundinganything-vlm

Use service_contract: openai and the corresponding port (8101 for DLM, 8102 for VLM). For a preflight that checks configured inputs without sending inference requests:

.venv-eval/bin/python scripts/evaluate.py configs/eval/dlm.yaml --dry-run

Configure the running service, data_root, dataset registry, task list, and a fresh run_id before launching. The default limit: 8 is a smoke test; set limit: null for full evaluation. Evaluation connects to an existing service and does not start or change the decoder.

Each run saves run.json (configuration and provenance), responses.jsonl (raw responses, finish reasons, and token usage), task logs, and summary.json (metrics and completion state). Record the checkpoint revision and task configuration when comparing results.

Inference Infrastructure

SGLang Execution

The custom SGLang integration coordinates model loading, request scheduling, attention kernels, and KV-cache ownership for blockwise generation. Denoising uses bidirectional attention inside the active block; completed history is retained as causal KV. Self-speculation additionally verifies proposals and removes rejected suffix states. These cache semantics must be preserved when optimizing execution.

The supplied recipe selects BF16 + Triton attention + eager execution. DLM recipes record the engine source, decoder, effective settings, profile digest, and package versions in outputs/sglang/<decoder>/engine_runtime.json; DLM evaluation checks decoder alignment before starting workers. The VLM service records its causal model/runtime configuration separately in outputs/sglang/vlm/engine_runtime.json. For higher service concurrency, use independent replicas with distinct devices, ports, and output directories; the default queue is not continuous multi-request batching.

CUDA Graph and Selective FP8

The paper evaluates a progressive infrastructure sequence:

Layer Purpose Status in the supplied default recipe
Native PyTorch eager Reference model execution Reference implementation
SGLang eager Integrated scheduling, attention, and cache execution Default serving path
CUDA Graph replay Replay compatible captured GPU work to reduce repeated launch overhead Evaluated in the paper; disabled by the default launcher
Selective FP8 Reduce arithmetic cost in eligible language-model linear operations while retaining other components in BF16 Evaluated in the paper; default serving remains BF16

CUDA Graph changes how compatible GPU work is submitted; it does not define a new token-commitment or verification rule. Captured shapes and state updates must remain compatible with the active decoding path. Selective FP8 can change logits and subsequent decoding decisions, so it is a distinct numerical configuration.

The optional graph implementation captures fixed-shape block work. Its FlashInfer path uses persistent attention masks and device-resident buffer updates for bidirectional drafting and causal verification. The Triton verifier path separates draft/verification metadata and input buffers while sharing parameters and the real KV pool. Its supported capture case is B32 with one request and no tensor/pipeline/data parallel expansion; prefill and unsupported shapes use eager execution. Optional shadow checks compare cache states, logits, and token choices against eager execution. These implementation paths are not exposed as a public run.py serve --cuda-graph switch in this release.

The paper's cumulative infrastructure speedups compare execution implementations within the same decoding mode. They are separate from the headline comparison of self-speculation against the autoregressive checkpoint. The default release commands above do not enable Graph replay or FP8, and no unsupported activation flags are implied.

Ethical Considerations:

Grounding predictions can miss small or occluded objects, repeat instances, or produce inaccurate text and coordinates. Validate localization quality on the intended task and inspect incomplete responses. GUI points describe image locations; the model does not execute interface actions. Perception outputs require task-specific validation before integration into physical systems.

Downloads last month
-
Safetensors
Model size
5B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Papers for GroundingPI/GroundAnything