YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Siltframe

We measure how off-road perception breaks, find which condition the training data never covered, and fix it with the right data.

Cameras on agricultural, mining, construction and defence vehicles work in dust, at night, in fog, with rain and mud on the lens. Public off-road datasets are almost entirely daylight and dry, so a model's validation number says very little about what happens in the field. We measure the difference, per condition, per severity and per class.

siltframe.com ยท the stress test on GitHub

What is here

siltframe-stress-test โ€” public, free, CC BY-SA 4.0. 165 labelled images: 10 real off-road scenes under dust, night, fog, rain on the lens and mud on the lens at three severities each, plus 5 real rain and snow frames, and the script that turns a folder of predictions into a failure map. About ten minutes to get a number for a model you already have.

The other repositories hold training data and checkpoints and are private, because much of what we train on is licensed for research only and may not be redistributed.

What we found

Running one protocol across six architectures produced a result we did not expect and would rather publish than have a customer discover:

  • Night costs 37โ€“64 % of mIoU and no architecture escapes it. Dust costs 36โ€“49 %.
  • A model that scores badly can look like the robust one. The weakest model has the smallest relative drop under rain, because it had little left to lose. Relative drops mean nothing without the clear-weather score beside them.
  • Dataset coverage beats augmentation. Adding real frames from a dataset containing forest tracks lifted real-adverse-weather accuracy by 26โ€“35 points; weather augmentation on top of that coverage stayed within noise.

One failure traced in full: a model trained on open terrain calls overhead tree canopy "sky" โ€” 80 % of the tree pixels in one real forest-road frame. Adding synthetic bad weather made it worse, because haze removes the leaf texture that would have contradicted the shortcut. The right real frames took it to 0 %. Write-up.

That last one argues against our own product, which is the reason to trust the rest.

How we report numbers

Every figure we publish is measured, and we can say which generator and which split produced it. Evaluation deliberately uses a different generator from the one used to make training data โ€” testing on the code you trained with is how augmentation results flatter themselves. Where an independent generator exists we quote it even when our own is more flattering.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support