GroundingJev

Jev-inspired Non-autoregressive Visual Grounding

Inspired by TypeSafe Jev, GroundingJev applies direct, task-specific output prediction to visual grounding.

GroundingJev replaces autoregressive coordinate decoding with continuous bounding-box regression on the Qwen3.5-0.8B multimodal backbone. A lightweight MLP head maps the last valid token's hidden state to normalized cxcywh coordinates in a single forward pass.

This repository includes the complete model weights, processor, and Python source required for inference. The custom GroundingJevModel is loaded through the included groundingjev package.

Try the online demo · Source, training, and evaluation.

Quick start

Use Python 3.12. Download the model and source, then install the package:

python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install huggingface-hub==1.5.0
python -c "from huggingface_hub import snapshot_download; snapshot_download('xyzzzh/GroundingJev', local_dir='GroundingJev')"
cd GroundingJev
python -m pip install -e .

python -m groundingjev.predict \
  --checkpoint . \
  --image /path/to/image.jpg \
  --expression 'the person wearing a red shirt' \
  --device cuda

Replace the image path and expression with your example. The JSON field bbox_xyxy contains [x1, y1, x2, y2] in original-image pixels. Add --output prediction.json to save the result. No separate base-model download is needed for inference.

Evaluation

Full test-set results.

RefCOCO and RefCOCO+ pool all testA and testB samples; RefCOCOg uses test.

mIoU ↑ (%)

Dataset Qwen3.5-0.8B GroundingJev Gain (pp)
RefCOCO 72.83 78.26 +5.43
RefCOCO+ 64.84 73.18 +8.33
RefCOCOg 71.52 74.97 +3.45

IoU@0.5

Dataset Qwen3.5-0.8B GroundingJev Gain (pp)
RefCOCO 79.74 89.11 +9.37
RefCOCO+ 70.10 82.82 +12.72
RefCOCOg 77.96 85.46 +7.50

Inference performance

Model Latency ↓ (ms) Throughput ↑ (samples/s) Speedup ↑
Qwen3.5-0.8B 1588.53 0.630 1.00×
GroundingJev 184.47 5.421 8.61×

8.61× inference speedup, with 88.39% lower mean latency.

End-to-end prediction time, excluding model loading and warmup.

Requested batch: Qwen3.5-0.8B=1, GroundingJev=1

Full split results.

Latency includes image preparation and prediction, excluding model loading and warmup. Measured on 256 RefCOCO testA samples with batch size 1 on an NVIDIA A100, after 10 warmup samples.

Training

Training uses 80,000 RefCOCO examples and the objective 5 × L1 + 2 × (1 − GIoU). The regression head is adapted for 100 steps, followed by two epochs of joint fine-tuning of the language backbone, visual merger, and head; the remaining visual encoder parameters stay frozen. Both stages use warmup and cosine learning-rate decay. Training uses ModelScope ms-swift, and evaluation uses EvalScope. See the training configuration.

Scope

The model returns one box for the described object. It does not produce segmentation masks or multiple-object detections. Performance on other languages and domains has not been evaluated.

License and acknowledgments

See LICENSE. GroundingJev builds on Qwen3.5-0.8B and is inspired by TypeSafe Jev. We thank Qwen, ModelScope ms-swift/EvalScope, SwanLab, and the COCO/RefCOCO authors. Datasets retain their respective licenses.

Downloads last month
43
Safetensors
Model size
0.9B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for xyzzzh/GroundingJev

Finetuned
(424)
this model

Space using xyzzzh/GroundingJev 1