RF-DETR Small + text-aligned head (ONNX)
Open-vocabulary detection by adding one trained projection to a closed-set RF-DETR Small detector, served by VisionServe (Docker Hub).
The detector exports the 256-d feature f_i of each of its 300 object queries as an extra ONNX
output, query_feats. A trained matrix P maps that feature into the text-embedding space of a
CLIP or SigLIP text tower, and each class gets a cosine score against the embedding of its name:
logit_ic = a · ⟨ t̂_c , P f_i / ‖P f_i‖ ⟩ + b conf = sigmoid(logit)
The vocabulary is the prompt, for example "cup. cola can. banana.". Changing it needs no
re-export and no second detector: the text tower runs once per new vocabulary, and the result is
cached. Boxes always come from the detector.
visionserve pull rfdetr-textalign-dec1-siglip # also pulls siglip-text
visionserve run rfdetr-textalign-dec1-siglip img.png --prompt "cup. cola can. banana."
pull downloads one subfolder of this repo, writes a manifest.yaml, and pulls the text tower
the model needs as a separate dependency.
Models
visionserve pull |
folder | detector | text space (dependency) | default conf_threshold |
|---|---|---|---|---|
rfdetr-textalign-dec1 |
dec1/ |
dec1, 22 classes |
CLIP ViT-B/32, 512-d (clip-text) |
0.35 |
rfdetr-textalign-dec1-siglip |
dec1-siglip/ |
dec1 held-out, 17 classes |
SigLIP base-patch16-224, 768-d (siglip-text) |
0.35 |
rfdetr-textalign-dec1-siglip-prod |
dec1-siglip-prod/ |
dec1, 22 classes |
SigLIP base-patch16-224, 768-d (siglip-text) |
0.001 |
rfdetr-textalign-etri |
etri/ |
rfdetr-small-etri, 22 classes, frozen |
CLIP ViT-B/32, 512-d (clip-text) |
0.28 |
clip-text |
clip-text/ |
(text tower only) | CLIP ViT-B/32 text tower | — |
dec1is RF-DETR Small fine-tuned on the 22 tabletop classes with the backbone and box head frozen and only the last decoder layer trainable.dec1held-out is the same recipe trained on 17 of the 22 classes. The other 5 names were held out of every trained component, so they can only be reached through a prompt.rfdetr-small-etriis the full fine-tune published asmtbui2010/rfdetr-small-etri-ONNX, re-exported withquery_feats.
clip-text is the text half of CLIP ViT-B/32 (openai/clip-vit-base-patch32, MIT). It is
needed by rfdetr-textalign-dec1 and rfdetr-textalign-etri and had no public ONNX export, so it
is published here as its own catalog model. siglip-text comes from
mtbui2010/siglip-base-patch16-224-ONNX.
The source repository also has -probe variants of three of these models
(rfdetr-textalign-dec1-probe, -dec1-siglip-probe, -etri-probe). They are evaluation configs:
the same files with conf_threshold: 0.001, so mAP is computed over the full precision/recall
curve. They are not published separately. To get one, pull the model and lower
postprocess.conf_threshold in its manifest.
How each model scores
The method request field picks how the head is applied (default when omitted: exact).
exact: the formula above. It runs as an ONNX Runtime session (head.onnx) when the manifest names one, which every model here does.folded: drops‖P f_i‖, which makes the head a single[C, 256]matrix.gated: the detector's own class head gives the score used for thresholding and ranking, and the text-aligned head gives only the name. Changing the vocabulary then changes labels and leaves confidences unchanged.
Each manifest's conf_threshold was tuned for one method. Changing the method changes the scale
of the thresholded score.
rfdetr-textalign-etri: 0.28 was tuned forexact, the default.rfdetr-textalign-dec1and-dec1-siglip: 0.35 was tuned forgated, which is how they are served (--method gated, or-F method=gatedon/api/predict). With the defaultexact, the thresholded score is the text-aligned head's, whose scale is different, and 0.35 returns few or no detections.
visionserve run rfdetr-textalign-dec1-siglip img.png --prompt "cup. cola can." --method gated
Files
| file | contents |
|---|---|
*/detector.onnx, etri/detector-noattn.onnx |
the detector, opset 17. Input input [1,3,512,512]. Outputs dets [1,300,4] cxcywh, labels [1,300,C+1] sigmoid logits, query_feats [1,300,256], in that order |
etri/detector-qf.onnx |
the same etri detector with a fourth output, cross_attn_weights, before query_feats. Used only by /api/explain, and loaded on the first explain call |
*/proj.bin |
the trained projection: 32-byte VSTXALN1 header (d_text, d_feat, a, b, flags), then P as row-major float32 [d_text][256] |
*/head.onnx |
proj.bin exported as an ONNX graph (opset 13): inputs query_feats [batch,queries,256] and text_embeds [classes,d_text], output raw logits [batch,queries,classes] |
*/labels.txt |
the detector's own class names in head order, then N/A. Used as the vocabulary when a request has no prompt |
*/templates.txt |
the 10 prompt templates the head was trained with. Part of the head's contract |
dec1-siglip/proj-openvocab.bin |
an alternative projection for dec1-siglip (see below). Not pulled by the catalog |
dec1-siglip/vocab-all22.txt |
all 22 names, for pasting into a prompt |
clip-text/model.onnx |
CLIP ViT-B/32 text tower, opset 18: input_ids int64 [batch,77] to text_embeds [batch,512] |
clip-text/vocab.json, clip-text/merges.txt |
CLIP's BPE tokenizer files, copied unchanged from openai/clip-vit-base-patch32 |
clip-text/LICENSE |
the MIT licence of OpenAI CLIP |
dec1/detector.onnx and dec1-siglip-prod/detector.onnx are the same file (same sha256). Both
folders carry it so that each folder is complete on its own.
sha256
| file | bytes | sha256 |
|---|---|---|
dec1/detector.onnx |
114 290 612 | cd4cb2166978635de3ab2323ed0d7198cc1ae5b77dc6e21c4c90845877579125 |
dec1/head.onnx |
525 158 | ab6db90d0e921931be6303d06c2691ee777ce56fd9a7750e7741f9a8d667b3ab |
dec1/proj.bin |
524 320 | 1d81d692de4c7cd3b3345cfe317afd78f752be0347181798847709cc3a2aff2b |
dec1/labels.txt |
189 | fdd0d4e9dc1965b37e46959d1fd2f3964fb407e5b176ea84a357dbeb921d183e |
dec1/templates.txt |
370 | dd37dd428e8c0700e26b86a6c7701a9e50a26a9c32b932febdbcb9ebb45c663c |
dec1-siglip/detector.onnx |
114 280 331 | 7de8ca150390b8e5d64d4541695a6b167c76793fdc0887e60c7e1481572c5a09 |
dec1-siglip/head.onnx |
787 303 | ec724f1a1c338795e1db37dcb9892d27b8ffb6d1f69f2c47c0c2558f281cce5e |
dec1-siglip/proj.bin |
786 464 | 302640c92684e78b64e7c0fd89b4f1c2761184408c7735dbb74ab43257372c85 |
dec1-siglip/proj-openvocab.bin |
786 464 | 8e77bc51fc4700680c6e6cdcafe79adf86094c00be8161650b25b3476f20cc07 |
dec1-siglip/labels.txt |
154 | 321bf1eb6803aa638016b48b7597f4fd56e73df11c36ebcf52dacbea65daa787 |
dec1-siglip/templates.txt |
370 | dd37dd428e8c0700e26b86a6c7701a9e50a26a9c32b932febdbcb9ebb45c663c |
dec1-siglip/vocab-all22.txt |
189 | fdd0d4e9dc1965b37e46959d1fd2f3964fb407e5b176ea84a357dbeb921d183e |
dec1-siglip-prod/detector.onnx |
114 290 612 | cd4cb2166978635de3ab2323ed0d7198cc1ae5b77dc6e21c4c90845877579125 |
dec1-siglip-prod/head.onnx |
787 302 | 3f19241baa19cafb2f673a101f639aba55f445a944d055516745ab966e7fb804 |
dec1-siglip-prod/proj.bin |
786 464 | 57ceaa1539bca398f3e495353f1761422594c11b8c683064a650a4fc6dcea91c |
dec1-siglip-prod/labels.txt |
189 | fdd0d4e9dc1965b37e46959d1fd2f3964fb407e5b176ea84a357dbeb921d183e |
dec1-siglip-prod/templates.txt |
370 | dd37dd428e8c0700e26b86a6c7701a9e50a26a9c32b932febdbcb9ebb45c663c |
etri/detector-noattn.onnx |
114 399 718 | efcf3af08d5e0512095946b69866425645fca81dc5e05aaea775eff9abc11b4c |
etri/detector-qf.onnx |
114 399 892 | 5d87e22067458c9af1f679a8eeb85588569a84881a060ac0f1d8f8f252379818 |
etri/head.onnx |
525 158 | 5484c2cdd32deb74776b9b8d7b9621354a330ac417eee2ba4cf0b497333f3c38 |
etri/proj.bin |
524 320 | 314302bf7549d85ef2467375fa2d415dca2eb1dbfee6bf6e3f8efddafd89392b |
etri/labels.txt |
189 | fdd0d4e9dc1965b37e46959d1fd2f3964fb407e5b176ea84a357dbeb921d183e |
etri/templates.txt |
370 | dd37dd428e8c0700e26b86a6c7701a9e50a26a9c32b932febdbcb9ebb45c663c |
clip-text/model.onnx |
255 217 458 | a104b96e1a9ce466e24dac4e32f406ffc412eb1a459049b4040eab97b196b580 |
clip-text/vocab.json |
862 328 | 5047b556ce86ccaf6aa22b3ffccfc52d391ea4accdab9c2f2407da5b742d4363 |
clip-text/merges.txt |
524 657 | f526393189112391ce6f9795d4695f704121ce452c3aad1f5335cc41337eba85 |
clip-text/LICENSE |
1 064 | 987e63b32f6c89ff5160e429458a872ff048e6860b590a3912e938f9da8f14db |
visionserve pull checks every file it downloads against these digests, and the generated
manifest pins them again so the server re-checks them at load.
How head.onnx is made
head.onnx holds no new weights. It is proj.bin written as an ONNX graph by
models/rfdetr-textalign-dec1-siglip/export_head_onnx.py,
so that method exact runs on ONNX Runtime instead of in Go. The export is deterministic.
Re-running it on each proj.bin here, with onnx 1.21.0, gives byte-identical files:
python3 models/rfdetr-textalign-dec1-siglip/export_head_onnx.py \
--proj <folder>/proj.bin --out <folder>/head.onnx
The script also checks the graph against a NumPy reference (max |Δlogit| about 2e-6 on CPU).
The server compares one query per request against proj.bin and refuses a head.onnx that
disagrees. If you swap in another projection, for example proj-openvocab.bin renamed to
proj.bin, re-export head.onnx too. The manifests set runtime.threads: {head: 1}. The head
takes about 1 ms on one thread, and ORT's default thread pool for it slowed the detector about 3x
on CPU.
Training and data
- Detector base. RF-DETR Small, COCO-pretrained, from Roboflow (Apache-2.0).
- Fine-tuning data. tabletop-22: 247 usable images of household and tabletop objects on an
indoor work surface, 1447 boxes, 22 classes. The authors captured it with a robot arm's wrist
camera at the Korea Electronics Technology Institute (KETI). It contains no people and no
third-party imagery, and the authors release it under Apache-2.0. During development it was
called ETRI, which is where the
-etrimodel names come from. Every model uses the same split: 185 training and 62 held-out images, by image, seed 0. - Projection
P. Trained by distillation from the text tower's teacher. The local READMEs describe the recipes:etri: CLIP text anchors on the 22 class names, plus label-free CLIP region distillation over the detector's own proposals. It was trained with the(‖P f‖−1)²penalty (flags bit 0 set), witha = 13.7738andb = −5.0417.dec1: CLIP teacher,train_qf.py --n-coco 30000 --steps 6240 --lam-dec 4 --lam-norm 1, over COCO train2017 images used as an unlabelled distillation pool. No norm penalty (flags 0).dec1-siglip: SigLIP teacher, a mixed pool of in-domain tabletop crops and 2000 COCO images with proportional weighting.proj-openvocab.binuses 5000 COCO images with balanced weighting.dec1-siglip-prod: a SigLIP-space projection (d_text768) for the 22-classdec1detector. The source repository does not document its recipe or any measurement for it.
Measured results
Every number below comes from the VisionServe source READMEs named in each line. All are on the 62 held-out tabletop-22 images unless a line says otherwise, and all were measured through the VisionServe server.
rfdetr-textalign-etri (README).
RTX A6000, ONNX Runtime 1.26, CUDA EP:
| system | mAP@[.5:.95] | mAP@.5 | p50 |
|---|---|---|---|
rfdetr-small-etri (closed-set, 22 classes) |
70.0 | 95.1 | 31.8 ms |
this model, exact |
54.2 | 74.9 | 61.6 ms |
this model, folded |
54.2 | 74.9 | 37.0 ms |
grounding-dino (tiny) |
46.1 | 53.7 | 220.2 ms |
These p50 values were measured before head.onnx existed, with exact running in Go.
rfdetr-textalign-dec1 (README).
COCO-style mAP, measured through the -probe config (conf_threshold 0.001):
| system | mAP | mAP50 | AR100 | p50 |
|---|---|---|---|---|
rfdetr-small-etri (closed-set, 22 classes) |
69.97 | 95.05 | 75.72 | 38.2 ms |
this model, gated |
56.10 | 63.42 | 72.56 | 38.2 ms |
rfdetr-textalign-etri |
54.18 | 74.94 | 63.28 | 71.7 ms |
conf_threshold 0.35 is the F1 peak (0.768) for gated on the same images.
rfdetr-textalign-dec1-siglip (README).
This table is naming accuracy (top-1) on detections matched to a ground-truth box. It is not mAP.
"Unseen 5" means the 5 names that were held out of every trained component:
| head | tabletop base 17 | tabletop unseen 5 | COCO top-1 | COCO macro |
|---|---|---|---|---|
this model (proj.bin) |
81.6 | 57.3 | 85.0 | 78.3 |
proj-openvocab.bin |
78.1 | 64.0 | — | 72.2 |
| CLIP crop classifier | 72.4 | 65.3 | 57.7 | 66.4 |
| SigLIP crop classifier (the teacher) | 78.8 | 85.3 | 58.5 | 73.7 |
No end-to-end mAP is documented for this model. Its conf_threshold (0.35) was taken from the
22-class sibling and has not been re-tuned.
rfdetr-textalign-dec1-siglip-prod: no measurement is documented.
Limitations
- One domain. The detectors were fine-tuned on one indoor tabletop scene. The
etrihead does not transfer: on COCO val2017 novel categories it scores 3.7% box-conditioned top-1, against 57.0% for a CLIP crop classifier. Thedec1-sigliphead was distilled on a mixed pool to address this. Its COCO numbers are in the table above. - Near-synonyms. The head cannot separate names that the text tower barely separates, such as
coffeeandcoffee can(cosine 0.926 in CLIP text space). Four of the 22 classes score 0% with theetrihead for this reason. - Closed choice. The prompt is a closed list. Every kept box is given the nearest name in it, so a box whose real name is missing from the prompt still gets a label.
- Thresholds are per method. See How each model scores.
rfdetr-textalign-dec1-siglip-prodships withconf_threshold0.001, so it returns many low-confidence detections. Threshold them on the client, or raise the value in the manifest. - Resampling gap.
Pwas fitted on features from Pillow's bilinear resize, and the server usesdisintegration/imaging. Measured on theetrihead: median |Δconf| 0.0027, max 0.112.
License
Apache-2.0 for the detectors, projections and heads. They are fine-tunes of Roboflow RF-DETR
(Apache-2.0), plus projections trained by the authors. The text towers they depend on are also
permissive: SigLIP base-patch16-224 is Apache-2.0, and CLIP ViT-B/32 is MIT. The clip-text/
folder is redistributed under the MIT licence in clip-text/LICENSE
(Copyright (c) 2021 OpenAI). No AGPL component is used anywhere.
Model tree for mtbui2010/rfdetr-textalign-ONNX
Base model
Roboflow/rf-detr-small