dataset · Zenodo (CERN European Organization for Nuclear Research)
A new dashcam road segmentation dataset addresses the gap in standard urban driving benchmarks by capturing unstructured and diverse environments. Sourced from 20 driving videos, the dataset covers rural settings, mountain routes, urban roads, highways, and nighttime driving conditions. In total, it comprises 5,939 extracted frames captured at a resolution of 960 by 720 pixels. Alongside the raw imagery, the collection includes 5,939 road segmentation masks generated automatically using the Segment Anything Model ViT-H foundation model. To enable rigorous performance benchmarking, 151 frames have been manually annotated to provide ground truth labels. This resource is specifically intended to evaluate the zero-shot segmentation capabilities of foundation vision models on road conditions that are routinely missing from conventional autonomous driving datasets such as Cityscapes or CamVid.
Standard driving datasets focus heavily on predictable, well-structured city streets, leaving vision models untested on complex or rural terrain. By capturing challenging conditions such as mountain passes and nighttime routes, this benchmark helps developers understand how vision systems interpret unstructured roads. Such assessments are essential for verifying whether computer vision models can operate reliably outside strictly curated urban centres.
This dataset serves as an early-stage benchmarking resource for computer vision researchers and automated driving developers evaluating foundation models. It enables testing of perception software across non-standard environments, such as unpaved rural tracks or poorly lit mountain passes. While the dataset itself is an early testing tool rather than a finished navigation product, it supports the development of more robust vision algorithms for advanced driver assistance systems.
AI-generated from the published abstract. Always read the original work before citing.
A dashcam road segmentation dataset collected from 20 driving videos covering diverse road types including rural, mountain, urban, highway, and night conditions. Contains 5,939 extracted frames at 960×720 resolution, 5,939 SAM ViT-H generated road segmentation masks, and 151 manually annotated ground truth labels used for evaluation. Collected to benchmark zero-shot foundation model segmentation on unstructured road environments not represented in standard urban driving datasets such as Cityscapes or CamVid.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.5281/zenodo.22301900
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.