cargo-640-rfdetr
An RF-DETR-Nano fine-tuned to find two kinds of handling unit in a loading bay: steel coils and shrink-wrapped sacks. Exported to ONNX at 640Γ640, float16, 52 MB.
It exists because the alternative did not work. Each cargo type used to get its own single-class model, fine-tuned from the last one, and that chain does not survive a site that moves two things β measured on the holdout sets below, the sack fine-tune cost more than half the coil accuracy and the coil model found no sacks at all. Training the classes together recovers nearly all of it in one graph.
| model | coil AP@50 | sack AP@50 |
|---|---|---|
coil-640 (single-class) |
0.798 | 0.000 |
sack-v3 (single-class) |
0.386 | 0.899 |
cargo-640 (this model) |
0.780 | 0.851 |
Classes
Class id order is the contract with everything downstream β the exported id2label, the
serving class filter, and any stored detections document. It does not change.
0 coil
1 sack
Slot 2 exists because RF-DETR emits num_classes + 1 logits; it never wins the argmax and
is named _unused_2 so a decoder cannot serve a bare index as a class name.
Intended use
Counting handling units crossing a tripwire on a fixed camera in a loading bay, at 640Γ640, with detections consumed by a tracker. It was trained on two bays of one facility, so treat any other site as unverified until measured β the honest test is whether the unit is held while it travels, not whether the boxes look right in a still.
Not intended for: cargo types outside these two, moving cameras, or anything needing precise box geometry (see the limitation below).
Training
Warm-started from coil-640 rather than COCO. The set is small, so where the weights start
matters more than usual, and coil-640 already knew this camera, this lighting and this
failure mode β wrapped cargo on a pallet, in a dark bay, with a pile of wooden pallets in
the background it must not fire on.
| Architecture | RF-DETR-Nano, resolution 640 |
| Init | coil-640 EMA checkpoint |
| Data | 1019 train / 512 valid images; 4076 / 1000 boxes |
| Schedule | 100 epochs, effective batch 16, seed 0; best EMA at epoch 85 |
| Time | 3 h 9 m on an RTX 4060 |
| Holdout (framework metrics, both classes) | mAP@50 0.825, mAP@50:95 0.624 |
The per-class AP@50 figures in the table at the top come from an independent evaluation at IoU 0.50 with class-agnostic NMS, so they are not directly comparable with the framework's own numbers above; both are reported because the first answers "is this model better than the one it replaces" and the second answers "how tight are the boxes".
Both classes carry hard negatives β frames of the empty bay and the background pallet stack β which is what stops the detector finding cargo in scenery.
Limitations
Localisation is looser than AP@50 suggests. On the training framework's own holdout metrics, mAP@50 is 0.825 while mAP@50:95 is 0.624 β the boxes are in the right places and their edges are approximate. That is because the sack labels are geometric projections, not hand-drawn boxes: a person fits a box to the pallet in a marking sheet and a 5-bag pinwheel model projects the individual sacks through it. That is accurate enough to count and not accurate enough to measure a sack's size or position.
Only visible sacks are labelled. The projection knows where all twenty sacks on a pallet are, including the ones buried inside it. Training on those would teach the model to report cargo it cannot see, and it would score well doing it. So a count from this model is of units in view; turning that into a load total needs the stacking pattern, which is a site property and not something the detector can supply.
It is 2.3% behind coil-640 on coil AP@50 and 5.3% behind sack-v3 on sack AP@50.
That is the price of one graph instead of two. It cost nothing on the counts measured here,
but a site needing the last point of coil accuracy should use the single-class model.
The sack class has a quarter of the coil class's boxes (1063 against 3013 in training), which is the most likely reason it gives up more.
Verified counts
End to end through the counting pipeline, against hand counts, with no per-clip tuning:
| clip | cargo | expected | counted |
|---|---|---|---|
| boxload1 | coil | 13 | 13 (tripwire), 13 (depletion) |
| boxload4, coil run | coil | 13 | 13 |
| kpt 02 | sack | 22 visible | 22 |
| kpt 03 | sack | 10 visible | 10 |
| kpt 01 | sack | 25 visible | 21 |
kpt 01 under-reads: on its first pallet a worker leans across the load for most of the crossing, and sacks that are never separately detected are never counted.
Files
model.onnx float16 graph, outputs pred_boxes + logits
config.json id2label β required, or a decoder serves classes as "0"
preprocessor_config.json 640Γ640, rescale, ImageNet normalisation
do_normalize: true with the ImageNet statistics is not optional. Roboflow's exporter
does not fold normalisation into the weights, so a runtime that rescales to 0β1 and stops
there will silently degrade β measured on one clip, that mismatch cost a unit end to end.
Licence
Apache-2.0, inherited from RF-DETR.
- Downloads last month
- 29