Which conditions break your off-road perception model, and how badly — measured in about ten minutes.
165 labelled images: 10 real off-road scenes rendered under dust, night, fog, rain on the lens and mud on the lens at three severities each, plus 5 real rain and snow frames. Every image has a pixel-level label. A scoring script turns a folder of predictions into a failure map — mIoU per condition and severity, the drop against clear weather, traversable-ground IoU, and the classes that fail first.
This is the free sample of the Siltframe stress-test pack. It is complete and usable on its own, including commercially.
| Images | 165 JPEG, 819×400 |
| Labels | PNG, one class id per pixel, 19 classes (RELLIS-3D ontology) |
| Conditions | clean, dust, night, fog, lens rain, lens mud (×3 severities), real rain/snow |
| Licence | images and labels CC BY-SA 4.0 · evaluate.py MIT |
| Source | GOOSE validation split, Fraunhofer IOSB |
| Mirror | github.com/egeizgi/siltframe-stress-test |
clean/images, clean/labels 10 original frames and their labels
degraded/<condition>/s1..s3/ 150 rendered frames (5 conditions × 3 severities × 10)
real_rain/ 5 real rain and snow frames
baselines/ failure maps of three reference models
manifest.csv image, label, frame, source, condition, severity, seed
classes.csv id, name, colour, traversable flag
evaluate.py the scoring script (numpy + Pillow only)
from huggingface_hub import snapshot_download
path = snapshot_download("clappoxes/siltframe-stress-test", repo_type="dataset")
Write one prediction PNG per image, at the same relative path, each pixel holding a class id from classes.csv:
import csv, numpy as np
from pathlib import Path
from PIL import Image
for r in csv.DictReader(open(f"{path}/manifest.csv")):
img = np.array(Image.open(f"{path}/{r['image']}"))
pred_ids = your_model(img) # H × W array of class ids
out = Path("pred") / Path(r["image"]).with_suffix(".png")
out.parent.mkdir(parents=True, exist_ok=True)
Image.fromarray(pred_ids.astype(np.uint8)).save(out)
python evaluate.py --pred pred
Different class set? Map yours onto the 19 in classes.csv. Fewer classes is fine — anything you don't predict
scores 0, and the mean is taken over the classes present in the labels.
From the included baselines/, measured with the same evaluate.py. Mean mIoU loss against clear weather:
| clean mIoU | night | dust | rain on lens | fog | mud on lens | |
|---|---|---|---|---|---|---|
| SegFormer-B0 (RELLIS-3D only) | 14.5 | −69% | −43% | −23% | −23% | −10% |
| SegFormer-B0 (+ GOOSE) | 52.5 | −61% | −43% | −41% | −20% | −20% |
| Mask2Former Swin-T (+ GOOSE) | 58.1 | −56% | −35% | −41% | −16% | −19% |
Two things worth reading carefully, and they are why the per-condition breakdown exists:
Night is the universal failure. Every model we have measured loses more than half its mIoU. If your validation set is daytime and dry, its number is not the number you get in the field.
A model that scores badly can look like the robust one. SegFormer-B0 trained on RELLIS-3D alone has the smallest relative drop under rain and mud — because at 14.5 mIoU there was little left to lose. Relative drops mean nothing without the clear-weather score beside them.
Running this protocol across six architectures produced a result worth more than the benchmark itself: dataset coverage beats augmentation. Adding real frames from a dataset containing forest tracks lifted real-adverse-weather accuracy by 26–35 points, while weather augmentation on top of that coverage stayed within noise on real weather.
One failure traced specifically: models trained on open terrain call overhead tree canopy "sky" — 80 % of the tree pixels in one real forest-road frame. Adding the right real frames took it to 0 %. Degrading open-terrain images made it worse, because haze removes the leaf texture that would have contradicted the shortcut. Write-up.
So the useful question is not "how robust is my model" but which condition is my data missing.
Physically-modelled rather than filter-based: Koschmieder atmospheric scattering with monocular depth for dust and fog, headlight falloff with Poisson–Gaussian sensor noise for night, and droplet and splatter optics on the lens plane for rain and mud. Severity scales with resolution. Pixels behind a veil dense enough to block more than 97 % of the light are marked void and never scored, because no model can be asked to predict what the camera cannot see.
They are rendered, not captured — a cheap screen, not a substitute for a real adverse-condition validation set.
The real_rain/ frames are genuine, and are there precisely so the synthetic numbers have something to be checked
against.
evaluate.py: MIT.Frames come from the GOOSE validation split, which none of the reference models was trained on.
@misc{siltframe_stress_test_2026,
title = {Off-road segmentation stress test},
author = {Izgi, Ege},
year = {2026},
note = {Derived from the GOOSE dataset (Fraunhofer IOSB), CC BY-SA 4.0},
url = {https://siltframe.com/stress-test}
}
The full pack is 2,460 labelled images — 150 scenes under the same conditions plus 60 real rain and snow frames: siltframe.com/stress-test.
Which conditions break your off-road perception model, and how badly — measured in about ten minutes.
165 labelled images: 10 real off-road scenes rendered under dust, night, fog, rain on the lens and mud on the lens at three severities each, plus 5 real rain and snow frames. Every image has a pixel-level label. A scoring script turns a folder of predictions into a failure map — mIoU per condition and severity, the drop against clear weather, traversable-ground IoU, and the classes that fail first.
This is the free sample of the Siltframe stress-test pack. It is complete and usable on its own, including commercially.
| Images | 165 JPEG, 819×400 |
| Labels | PNG, one class id per pixel, 19 classes (RELLIS-3D ontology) |
| Conditions | clean, dust, night, fog, lens rain, lens mud (×3 severities), real rain/snow |
| Licence | images and labels CC BY-SA 4.0 · evaluate.py MIT |
| Source | GOOSE validation split, Fraunhofer IOSB |
| Mirror | github.com/egeizgi/siltframe-stress-test |
clean/images, clean/labels 10 original frames and their labels
degraded/<condition>/s1..s3/ 150 rendered frames (5 conditions × 3 severities × 10)
real_rain/ 5 real rain and snow frames
baselines/ failure maps of three reference models
manifest.csv image, label, frame, source, condition, severity, seed
classes.csv id, name, colour, traversable flag
evaluate.py the scoring script (numpy + Pillow only)
from huggingface_hub import snapshot_download
path = snapshot_download("clappoxes/siltframe-stress-test", repo_type="dataset")
Write one prediction PNG per image, at the same relative path, each pixel holding a class id from classes.csv:
import csv, numpy as np
from pathlib import Path
from PIL import Image
for r in csv.DictReader(open(f"{path}/manifest.csv")):
img = np.array(Image.open(f"{path}/{r['image']}"))
pred_ids = your_model(img) # H × W array of class ids
out = Path("pred") / Path(r["image"]).with_suffix(".png")
out.parent.mkdir(parents=True, exist_ok=True)
Image.fromarray(pred_ids.astype(np.uint8)).save(out)
python evaluate.py --pred pred
Different class set? Map yours onto the 19 in classes.csv. Fewer classes is fine — anything you don't predict
scores 0, and the mean is taken over the classes present in the labels.
From the included baselines/, measured with the same evaluate.py. Mean mIoU loss against clear weather:
| clean mIoU | night | dust | rain on lens | fog | mud on lens | |
|---|---|---|---|---|---|---|
| SegFormer-B0 (RELLIS-3D only) | 14.5 | −69% | −43% | −23% | −23% | −10% |
| SegFormer-B0 (+ GOOSE) | 52.5 | −61% | −43% | −41% | −20% | −20% |
| Mask2Former Swin-T (+ GOOSE) | 58.1 | −56% | −35% | −41% | −16% | −19% |
Two things worth reading carefully, and they are why the per-condition breakdown exists:
Night is the universal failure. Every model we have measured loses more than half its mIoU. If your validation set is daytime and dry, its number is not the number you get in the field.
A model that scores badly can look like the robust one. SegFormer-B0 trained on RELLIS-3D alone has the smallest relative drop under rain and mud — because at 14.5 mIoU there was little left to lose. Relative drops mean nothing without the clear-weather score beside them.
Running this protocol across six architectures produced a result worth more than the benchmark itself: dataset coverage beats augmentation. Adding real frames from a dataset containing forest tracks lifted real-adverse-weather accuracy by 26–35 points, while weather augmentation on top of that coverage stayed within noise on real weather.
One failure traced specifically: models trained on open terrain call overhead tree canopy "sky" — 80 % of the tree pixels in one real forest-road frame. Adding the right real frames took it to 0 %. Degrading open-terrain images made it worse, because haze removes the leaf texture that would have contradicted the shortcut. Write-up.
So the useful question is not "how robust is my model" but which condition is my data missing.
Physically-modelled rather than filter-based: Koschmieder atmospheric scattering with monocular depth for dust and fog, headlight falloff with Poisson–Gaussian sensor noise for night, and droplet and splatter optics on the lens plane for rain and mud. Severity scales with resolution. Pixels behind a veil dense enough to block more than 97 % of the light are marked void and never scored, because no model can be asked to predict what the camera cannot see.
They are rendered, not captured — a cheap screen, not a substitute for a real adverse-condition validation set.
The real_rain/ frames are genuine, and are there precisely so the synthetic numbers have something to be checked
against.
evaluate.py: MIT.Frames come from the GOOSE validation split, which none of the reference models was trained on.
@misc{siltframe_stress_test_2026,
title = {Off-road segmentation stress test},
author = {Izgi, Ege},
year = {2026},
note = {Derived from the GOOSE dataset (Fraunhofer IOSB), CC BY-SA 4.0},
url = {https://siltframe.com/stress-test}
}
The full pack is 2,460 labelled images — 150 scenes under the same conditions plus 60 real rain and snow frames: siltframe.com/stress-test.