GroundingJev
Jev-inspired Non-autoregressive Visual Grounding
Inspired by TypeSafe Jev, GroundingJev applies direct, task-specific output prediction to visual grounding.
GroundingJev replaces autoregressive coordinate decoding with continuous bounding-box regression on the Qwen3.5-0.8B multimodal backbone. A lightweight MLP head maps the last valid token's hidden state to normalized cxcywh coordinates in a single forward pass.
This repository includes the complete model weights, processor, and Python source required for inference. The custom GroundingJevModel is loaded through the included groundingjev package.
Try the online demo · Source, training, and evaluation.
Quick start
Use Python 3.12. Download the model and source, then install the package:
python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install huggingface-hub==1.5.0
python -c "from huggingface_hub import snapshot_download; snapshot_download('xyzzzh/GroundingJev', local_dir='GroundingJev')"
cd GroundingJev
python -m pip install -e .
python -m groundingjev.predict \
--checkpoint . \
--image /path/to/image.jpg \
--expression 'the person wearing a red shirt' \
--device cuda
Replace the image path and expression with your example. The JSON field bbox_xyxy contains [x1, y1, x2, y2] in original-image pixels. Add --output prediction.json to save the result. No separate base-model download is needed for inference.
Evaluation
Full test-set results.
RefCOCO and RefCOCO+ pool all testA and testB samples; RefCOCOg uses test.
mIoU ↑ (%)
| Dataset | Qwen3.5-0.8B | GroundingJev | Gain (pp) |
|---|---|---|---|
| RefCOCO | 72.83 | 78.26 | +5.43 |
| RefCOCO+ | 64.84 | 73.18 | +8.33 |
| RefCOCOg | 71.52 | 74.97 | +3.45 |
IoU@0.5
| Dataset | Qwen3.5-0.8B | GroundingJev | Gain (pp) |
|---|---|---|---|
| RefCOCO | 79.74 | 89.11 | +9.37 |
| RefCOCO+ | 70.10 | 82.82 | +12.72 |
| RefCOCOg | 77.96 | 85.46 | +7.50 |
Inference performance
| Model | Latency ↓ (ms) | Throughput ↑ (samples/s) | Speedup ↑ |
|---|---|---|---|
| Qwen3.5-0.8B | 1588.53 | 0.630 | 1.00× |
| GroundingJev | 184.47 | 5.421 | 8.61× |
8.61× inference speedup, with 88.39% lower mean latency.
End-to-end prediction time, excluding model loading and warmup.
Requested batch: Qwen3.5-0.8B=1, GroundingJev=1
Latency includes image preparation and prediction, excluding model loading and warmup. Measured on 256 RefCOCO testA samples with batch size 1 on an NVIDIA A100, after 10 warmup samples.
Training
Training uses 80,000 RefCOCO examples and the objective 5 × L1 + 2 × (1 − GIoU). The regression head is adapted for 100 steps, followed by two epochs of joint fine-tuning of the language backbone, visual merger, and head; the remaining visual encoder parameters stay frozen. Both stages use warmup and cosine learning-rate decay. Training uses ModelScope ms-swift, and evaluation uses EvalScope. See the training configuration.
Scope
The model returns one box for the described object. It does not produce segmentation masks or multiple-object detections. Performance on other languages and domains has not been evaluated.
License and acknowledgments
See LICENSE. GroundingJev builds on Qwen3.5-0.8B and is inspired by TypeSafe Jev. We thank Qwen, ModelScope ms-swift/EvalScope, SwanLab, and the COCO/RefCOCO authors. Datasets retain their respective licenses.
- Downloads last month
- 43