AI & ML interests

Robustness of off-road perception: how segmentation models fail in dust, night, fog, rain and lens mud — and which data fixes it

Recent Activity

clappoxes  updated a Space about 14 hours ago
siltframe/README
clappoxes  published a Space about 14 hours ago
siltframe/README
clappoxes  updated a model about 15 hours ago
siltframe/README
View all activity

Organization Card

Siltframe

We measure how off-road perception breaks, find which condition the training data never covered, and fix it with the right data.

Cameras on agricultural, mining, construction and defence vehicles work in dust, at night, in fog, with rain and mud on the lens. Public off-road datasets are almost entirely daylight and dry, so a model's validation number says very little about what happens in the field. We measure the difference, per condition, per severity and per class.

siltframe.com · the stress test on GitHub

What is here

siltframe-stress-test — public, free, CC BY-SA 4.0. 165 labelled images: 10 real off-road scenes under dust, night, fog, rain on the lens and mud on the lens at three severities each, plus 5 real rain and snow frames, and the script that turns a folder of predictions into a failure map. About ten minutes to get a number for a model you already have.

The other repositories hold training data and checkpoints and are private, because much of what we train on is licensed for research only and may not be redistributed.

What we found

Running one protocol across six architectures produced a result we did not expect and would rather publish than have a customer discover:

  • Night costs 37–64 % of mIoU and no architecture escapes it. Dust costs 36–49 %.
  • A model that scores badly can look like the robust one. The weakest model has the smallest relative drop under rain, because it had little left to lose. Relative drops mean nothing without the clear-weather score beside them.
  • Dataset coverage beats augmentation. Adding real frames from a dataset containing forest tracks lifted real-adverse-weather accuracy by 26–35 points; weather augmentation on top of that coverage stayed within noise.

One failure traced in full: a model trained on open terrain calls overhead tree canopy "sky" — 80 % of the tree pixels in one real forest-road frame. Adding synthetic bad weather made it worse, because haze removes the leaf texture that would have contradicted the shortcut. The right real frames took it to 0 %. Write-up.

That last one argues against our own product, which is the reason to trust the rest.

How we report numbers

Every figure we publish is measured, and we can say which generator and which split produced it. Evaluation deliberately uses a different generator from the one used to make training data — testing on the code you trained with is how augmentation results flatter themselves. Where an independent generator exists we quote it even when our own is more flattering.