cargo-640-rfdetr

An RF-DETR-Nano fine-tuned to find two kinds of handling unit in a loading bay: steel coils and shrink-wrapped sacks. Exported to ONNX at 640Γ—640, float16, 52 MB.

It exists because the alternative did not work. Each cargo type used to get its own single-class model, fine-tuned from the last one, and that chain does not survive a site that moves two things β€” measured on the holdout sets below, the sack fine-tune cost more than half the coil accuracy and the coil model found no sacks at all. Training the classes together recovers nearly all of it in one graph.

model coil AP@50 sack AP@50
coil-640 (single-class) 0.798 0.000
sack-v3 (single-class) 0.386 0.899
cargo-640 (this model) 0.780 0.851

Classes

Class id order is the contract with everything downstream β€” the exported id2label, the serving class filter, and any stored detections document. It does not change.

0  coil
1  sack

Slot 2 exists because RF-DETR emits num_classes + 1 logits; it never wins the argmax and is named _unused_2 so a decoder cannot serve a bare index as a class name.

Intended use

Counting handling units crossing a tripwire on a fixed camera in a loading bay, at 640Γ—640, with detections consumed by a tracker. It was trained on two bays of one facility, so treat any other site as unverified until measured β€” the honest test is whether the unit is held while it travels, not whether the boxes look right in a still.

Not intended for: cargo types outside these two, moving cameras, or anything needing precise box geometry (see the limitation below).

Training

Warm-started from coil-640 rather than COCO. The set is small, so where the weights start matters more than usual, and coil-640 already knew this camera, this lighting and this failure mode β€” wrapped cargo on a pallet, in a dark bay, with a pile of wooden pallets in the background it must not fire on.

Architecture RF-DETR-Nano, resolution 640
Init coil-640 EMA checkpoint
Data 1019 train / 512 valid images; 4076 / 1000 boxes
Schedule 100 epochs, effective batch 16, seed 0; best EMA at epoch 85
Time 3 h 9 m on an RTX 4060
Holdout (framework metrics, both classes) mAP@50 0.825, mAP@50:95 0.624

The per-class AP@50 figures in the table at the top come from an independent evaluation at IoU 0.50 with class-agnostic NMS, so they are not directly comparable with the framework's own numbers above; both are reported because the first answers "is this model better than the one it replaces" and the second answers "how tight are the boxes".

Both classes carry hard negatives β€” frames of the empty bay and the background pallet stack β€” which is what stops the detector finding cargo in scenery.

Limitations

Localisation is looser than AP@50 suggests. On the training framework's own holdout metrics, mAP@50 is 0.825 while mAP@50:95 is 0.624 β€” the boxes are in the right places and their edges are approximate. That is because the sack labels are geometric projections, not hand-drawn boxes: a person fits a box to the pallet in a marking sheet and a 5-bag pinwheel model projects the individual sacks through it. That is accurate enough to count and not accurate enough to measure a sack's size or position.

Only visible sacks are labelled. The projection knows where all twenty sacks on a pallet are, including the ones buried inside it. Training on those would teach the model to report cargo it cannot see, and it would score well doing it. So a count from this model is of units in view; turning that into a load total needs the stacking pattern, which is a site property and not something the detector can supply.

It is 2.3% behind coil-640 on coil AP@50 and 5.3% behind sack-v3 on sack AP@50. That is the price of one graph instead of two. It cost nothing on the counts measured here, but a site needing the last point of coil accuracy should use the single-class model.

The sack class has a quarter of the coil class's boxes (1063 against 3013 in training), which is the most likely reason it gives up more.

Verified counts

End to end through the counting pipeline, against hand counts, with no per-clip tuning:

clip cargo expected counted
boxload1 coil 13 13 (tripwire), 13 (depletion)
boxload4, coil run coil 13 13
kpt 02 sack 22 visible 22
kpt 03 sack 10 visible 10
kpt 01 sack 25 visible 21

kpt 01 under-reads: on its first pallet a worker leans across the load for most of the crossing, and sacks that are never separately detected are never counted.

Files

model.onnx                  float16 graph, outputs pred_boxes + logits
config.json                 id2label β€” required, or a decoder serves classes as "0"
preprocessor_config.json    640Γ—640, rescale, ImageNet normalisation

do_normalize: true with the ImageNet statistics is not optional. Roboflow's exporter does not fold normalisation into the weights, so a runtime that rescales to 0–1 and stops there will silently degrade β€” measured on one clip, that mismatch cost a unit end to end.

Licence

Apache-2.0, inherited from RF-DETR.

Downloads last month
29
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support