RF-DETR Small + text-aligned head (ONNX)

Open-vocabulary detection by adding one trained projection to a closed-set RF-DETR Small detector, served by VisionServe (Docker Hub).

The detector exports the 256-d feature f_i of each of its 300 object queries as an extra ONNX output, query_feats. A trained matrix P maps that feature into the text-embedding space of a CLIP or SigLIP text tower, and each class gets a cosine score against the embedding of its name:

logit_ic = a · ⟨ t̂_c , P f_i / ‖P f_i‖ ⟩ + b        conf = sigmoid(logit)

The vocabulary is the prompt, for example "cup. cola can. banana.". Changing it needs no re-export and no second detector: the text tower runs once per new vocabulary, and the result is cached. Boxes always come from the detector.

visionserve pull rfdetr-textalign-dec1-siglip      # also pulls siglip-text
visionserve run rfdetr-textalign-dec1-siglip img.png --prompt "cup. cola can. banana."

pull downloads one subfolder of this repo, writes a manifest.yaml, and pulls the text tower the model needs as a separate dependency.

Models

visionserve pull folder detector text space (dependency) default conf_threshold
rfdetr-textalign-dec1 dec1/ dec1, 22 classes CLIP ViT-B/32, 512-d (clip-text) 0.35
rfdetr-textalign-dec1-siglip dec1-siglip/ dec1 held-out, 17 classes SigLIP base-patch16-224, 768-d (siglip-text) 0.35
rfdetr-textalign-dec1-siglip-prod dec1-siglip-prod/ dec1, 22 classes SigLIP base-patch16-224, 768-d (siglip-text) 0.001
rfdetr-textalign-etri etri/ rfdetr-small-etri, 22 classes, frozen CLIP ViT-B/32, 512-d (clip-text) 0.28
clip-text clip-text/ (text tower only) CLIP ViT-B/32 text tower —
  • dec1 is RF-DETR Small fine-tuned on the 22 tabletop classes with the backbone and box head frozen and only the last decoder layer trainable.
  • dec1 held-out is the same recipe trained on 17 of the 22 classes. The other 5 names were held out of every trained component, so they can only be reached through a prompt.
  • rfdetr-small-etri is the full fine-tune published as mtbui2010/rfdetr-small-etri-ONNX, re-exported with query_feats.

clip-text is the text half of CLIP ViT-B/32 (openai/clip-vit-base-patch32, MIT). It is needed by rfdetr-textalign-dec1 and rfdetr-textalign-etri and had no public ONNX export, so it is published here as its own catalog model. siglip-text comes from mtbui2010/siglip-base-patch16-224-ONNX.

The source repository also has -probe variants of three of these models (rfdetr-textalign-dec1-probe, -dec1-siglip-probe, -etri-probe). They are evaluation configs: the same files with conf_threshold: 0.001, so mAP is computed over the full precision/recall curve. They are not published separately. To get one, pull the model and lower postprocess.conf_threshold in its manifest.

How each model scores

The method request field picks how the head is applied (default when omitted: exact).

  • exact: the formula above. It runs as an ONNX Runtime session (head.onnx) when the manifest names one, which every model here does.
  • folded: drops ‖P f_i‖, which makes the head a single [C, 256] matrix.
  • gated: the detector's own class head gives the score used for thresholding and ranking, and the text-aligned head gives only the name. Changing the vocabulary then changes labels and leaves confidences unchanged.

Each manifest's conf_threshold was tuned for one method. Changing the method changes the scale of the thresholded score.

  • rfdetr-textalign-etri: 0.28 was tuned for exact, the default.
  • rfdetr-textalign-dec1 and -dec1-siglip: 0.35 was tuned for gated, which is how they are served (--method gated, or -F method=gated on /api/predict). With the default exact, the thresholded score is the text-aligned head's, whose scale is different, and 0.35 returns few or no detections.
visionserve run rfdetr-textalign-dec1-siglip img.png --prompt "cup. cola can." --method gated

Files

file contents
*/detector.onnx, etri/detector-noattn.onnx the detector, opset 17. Input input [1,3,512,512]. Outputs dets [1,300,4] cxcywh, labels [1,300,C+1] sigmoid logits, query_feats [1,300,256], in that order
etri/detector-qf.onnx the same etri detector with a fourth output, cross_attn_weights, before query_feats. Used only by /api/explain, and loaded on the first explain call
*/proj.bin the trained projection: 32-byte VSTXALN1 header (d_text, d_feat, a, b, flags), then P as row-major float32 [d_text][256]
*/head.onnx proj.bin exported as an ONNX graph (opset 13): inputs query_feats [batch,queries,256] and text_embeds [classes,d_text], output raw logits [batch,queries,classes]
*/labels.txt the detector's own class names in head order, then N/A. Used as the vocabulary when a request has no prompt
*/templates.txt the 10 prompt templates the head was trained with. Part of the head's contract
dec1-siglip/proj-openvocab.bin an alternative projection for dec1-siglip (see below). Not pulled by the catalog
dec1-siglip/vocab-all22.txt all 22 names, for pasting into a prompt
clip-text/model.onnx CLIP ViT-B/32 text tower, opset 18: input_ids int64 [batch,77] to text_embeds [batch,512]
clip-text/vocab.json, clip-text/merges.txt CLIP's BPE tokenizer files, copied unchanged from openai/clip-vit-base-patch32
clip-text/LICENSE the MIT licence of OpenAI CLIP

dec1/detector.onnx and dec1-siglip-prod/detector.onnx are the same file (same sha256). Both folders carry it so that each folder is complete on its own.

sha256

file bytes sha256
dec1/detector.onnx 114 290 612 cd4cb2166978635de3ab2323ed0d7198cc1ae5b77dc6e21c4c90845877579125
dec1/head.onnx 525 158 ab6db90d0e921931be6303d06c2691ee777ce56fd9a7750e7741f9a8d667b3ab
dec1/proj.bin 524 320 1d81d692de4c7cd3b3345cfe317afd78f752be0347181798847709cc3a2aff2b
dec1/labels.txt 189 fdd0d4e9dc1965b37e46959d1fd2f3964fb407e5b176ea84a357dbeb921d183e
dec1/templates.txt 370 dd37dd428e8c0700e26b86a6c7701a9e50a26a9c32b932febdbcb9ebb45c663c
dec1-siglip/detector.onnx 114 280 331 7de8ca150390b8e5d64d4541695a6b167c76793fdc0887e60c7e1481572c5a09
dec1-siglip/head.onnx 787 303 ec724f1a1c338795e1db37dcb9892d27b8ffb6d1f69f2c47c0c2558f281cce5e
dec1-siglip/proj.bin 786 464 302640c92684e78b64e7c0fd89b4f1c2761184408c7735dbb74ab43257372c85
dec1-siglip/proj-openvocab.bin 786 464 8e77bc51fc4700680c6e6cdcafe79adf86094c00be8161650b25b3476f20cc07
dec1-siglip/labels.txt 154 321bf1eb6803aa638016b48b7597f4fd56e73df11c36ebcf52dacbea65daa787
dec1-siglip/templates.txt 370 dd37dd428e8c0700e26b86a6c7701a9e50a26a9c32b932febdbcb9ebb45c663c
dec1-siglip/vocab-all22.txt 189 fdd0d4e9dc1965b37e46959d1fd2f3964fb407e5b176ea84a357dbeb921d183e
dec1-siglip-prod/detector.onnx 114 290 612 cd4cb2166978635de3ab2323ed0d7198cc1ae5b77dc6e21c4c90845877579125
dec1-siglip-prod/head.onnx 787 302 3f19241baa19cafb2f673a101f639aba55f445a944d055516745ab966e7fb804
dec1-siglip-prod/proj.bin 786 464 57ceaa1539bca398f3e495353f1761422594c11b8c683064a650a4fc6dcea91c
dec1-siglip-prod/labels.txt 189 fdd0d4e9dc1965b37e46959d1fd2f3964fb407e5b176ea84a357dbeb921d183e
dec1-siglip-prod/templates.txt 370 dd37dd428e8c0700e26b86a6c7701a9e50a26a9c32b932febdbcb9ebb45c663c
etri/detector-noattn.onnx 114 399 718 efcf3af08d5e0512095946b69866425645fca81dc5e05aaea775eff9abc11b4c
etri/detector-qf.onnx 114 399 892 5d87e22067458c9af1f679a8eeb85588569a84881a060ac0f1d8f8f252379818
etri/head.onnx 525 158 5484c2cdd32deb74776b9b8d7b9621354a330ac417eee2ba4cf0b497333f3c38
etri/proj.bin 524 320 314302bf7549d85ef2467375fa2d415dca2eb1dbfee6bf6e3f8efddafd89392b
etri/labels.txt 189 fdd0d4e9dc1965b37e46959d1fd2f3964fb407e5b176ea84a357dbeb921d183e
etri/templates.txt 370 dd37dd428e8c0700e26b86a6c7701a9e50a26a9c32b932febdbcb9ebb45c663c
clip-text/model.onnx 255 217 458 a104b96e1a9ce466e24dac4e32f406ffc412eb1a459049b4040eab97b196b580
clip-text/vocab.json 862 328 5047b556ce86ccaf6aa22b3ffccfc52d391ea4accdab9c2f2407da5b742d4363
clip-text/merges.txt 524 657 f526393189112391ce6f9795d4695f704121ce452c3aad1f5335cc41337eba85
clip-text/LICENSE 1 064 987e63b32f6c89ff5160e429458a872ff048e6860b590a3912e938f9da8f14db

visionserve pull checks every file it downloads against these digests, and the generated manifest pins them again so the server re-checks them at load.

How head.onnx is made

head.onnx holds no new weights. It is proj.bin written as an ONNX graph by models/rfdetr-textalign-dec1-siglip/export_head_onnx.py, so that method exact runs on ONNX Runtime instead of in Go. The export is deterministic. Re-running it on each proj.bin here, with onnx 1.21.0, gives byte-identical files:

python3 models/rfdetr-textalign-dec1-siglip/export_head_onnx.py \
    --proj <folder>/proj.bin --out <folder>/head.onnx

The script also checks the graph against a NumPy reference (max |Δlogit| about 2e-6 on CPU). The server compares one query per request against proj.bin and refuses a head.onnx that disagrees. If you swap in another projection, for example proj-openvocab.bin renamed to proj.bin, re-export head.onnx too. The manifests set runtime.threads: {head: 1}. The head takes about 1 ms on one thread, and ORT's default thread pool for it slowed the detector about 3x on CPU.

Training and data

  • Detector base. RF-DETR Small, COCO-pretrained, from Roboflow (Apache-2.0).
  • Fine-tuning data. tabletop-22: 247 usable images of household and tabletop objects on an indoor work surface, 1447 boxes, 22 classes. The authors captured it with a robot arm's wrist camera at the Korea Electronics Technology Institute (KETI). It contains no people and no third-party imagery, and the authors release it under Apache-2.0. During development it was called ETRI, which is where the -etri model names come from. Every model uses the same split: 185 training and 62 held-out images, by image, seed 0.
  • Projection P. Trained by distillation from the text tower's teacher. The local READMEs describe the recipes:
    • etri: CLIP text anchors on the 22 class names, plus label-free CLIP region distillation over the detector's own proposals. It was trained with the (‖P f‖−1)² penalty (flags bit 0 set), with a = 13.7738 and b = −5.0417.
    • dec1: CLIP teacher, train_qf.py --n-coco 30000 --steps 6240 --lam-dec 4 --lam-norm 1, over COCO train2017 images used as an unlabelled distillation pool. No norm penalty (flags 0).
    • dec1-siglip: SigLIP teacher, a mixed pool of in-domain tabletop crops and 2000 COCO images with proportional weighting. proj-openvocab.bin uses 5000 COCO images with balanced weighting.
    • dec1-siglip-prod: a SigLIP-space projection (d_text 768) for the 22-class dec1 detector. The source repository does not document its recipe or any measurement for it.

Measured results

Every number below comes from the VisionServe source READMEs named in each line. All are on the 62 held-out tabletop-22 images unless a line says otherwise, and all were measured through the VisionServe server.

rfdetr-textalign-etri (README). RTX A6000, ONNX Runtime 1.26, CUDA EP:

system mAP@[.5:.95] mAP@.5 p50
rfdetr-small-etri (closed-set, 22 classes) 70.0 95.1 31.8 ms
this model, exact 54.2 74.9 61.6 ms
this model, folded 54.2 74.9 37.0 ms
grounding-dino (tiny) 46.1 53.7 220.2 ms

These p50 values were measured before head.onnx existed, with exact running in Go.

rfdetr-textalign-dec1 (README). COCO-style mAP, measured through the -probe config (conf_threshold 0.001):

system mAP mAP50 AR100 p50
rfdetr-small-etri (closed-set, 22 classes) 69.97 95.05 75.72 38.2 ms
this model, gated 56.10 63.42 72.56 38.2 ms
rfdetr-textalign-etri 54.18 74.94 63.28 71.7 ms

conf_threshold 0.35 is the F1 peak (0.768) for gated on the same images.

rfdetr-textalign-dec1-siglip (README). This table is naming accuracy (top-1) on detections matched to a ground-truth box. It is not mAP. "Unseen 5" means the 5 names that were held out of every trained component:

head tabletop base 17 tabletop unseen 5 COCO top-1 COCO macro
this model (proj.bin) 81.6 57.3 85.0 78.3
proj-openvocab.bin 78.1 64.0 — 72.2
CLIP crop classifier 72.4 65.3 57.7 66.4
SigLIP crop classifier (the teacher) 78.8 85.3 58.5 73.7

No end-to-end mAP is documented for this model. Its conf_threshold (0.35) was taken from the 22-class sibling and has not been re-tuned.

rfdetr-textalign-dec1-siglip-prod: no measurement is documented.

Limitations

  • One domain. The detectors were fine-tuned on one indoor tabletop scene. The etri head does not transfer: on COCO val2017 novel categories it scores 3.7% box-conditioned top-1, against 57.0% for a CLIP crop classifier. The dec1-siglip head was distilled on a mixed pool to address this. Its COCO numbers are in the table above.
  • Near-synonyms. The head cannot separate names that the text tower barely separates, such as coffee and coffee can (cosine 0.926 in CLIP text space). Four of the 22 classes score 0% with the etri head for this reason.
  • Closed choice. The prompt is a closed list. Every kept box is given the nearest name in it, so a box whose real name is missing from the prompt still gets a label.
  • Thresholds are per method. See How each model scores. rfdetr-textalign-dec1-siglip-prod ships with conf_threshold 0.001, so it returns many low-confidence detections. Threshold them on the client, or raise the value in the manifest.
  • Resampling gap. P was fitted on features from Pillow's bilinear resize, and the server uses disintegration/imaging. Measured on the etri head: median |Δconf| 0.0027, max 0.112.

License

Apache-2.0 for the detectors, projections and heads. They are fine-tunes of Roboflow RF-DETR (Apache-2.0), plus projections trained by the authors. The text towers they depend on are also permissive: SigLIP base-patch16-224 is Apache-2.0, and CLIP ViT-B/32 is MIT. The clip-text/ folder is redistributed under the MIT licence in clip-text/LICENSE (Copyright (c) 2021 OpenAI). No AGPL component is used anywhere.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mtbui2010/rfdetr-textalign-ONNX

Quantized
(2)
this model