ysilabs/latam-urban-driving-sample
LATAM Urban Driving Dataset — Night / Low-Visibility / Potholes Sample (Intel RealSense D555) A free preview of a multimodal, edge-case-focused urban driving dataset captured in the State of Mexico (Edomex) with the Intel RealSense D555 (RGB + IR + metric depth), aimed at Autonomous Driving, ADAS, and Smart City perception teams. This sample: 641 contiguous night-driving frames, densely annotated with 562 pothole boxes measured in native 3D depth and vehicle classes… See the full description on the dataset page: https://huggingface.co/datasets/ysilabs/latam-urban-driving-sample.
LATAM Urban Driving Dataset — Night / Low-Visibility / Potholes Sample (Intel RealSense D555)
A free preview of a multimodal, edge-case-focused urban driving dataset captured in the State of Mexico (Edomex) with the Intel RealSense D555 (RGB + IR + metric depth), aimed at Autonomous Driving, ADAS, and Smart City perception teams.
This sample: 641 contiguous night-driving frames, densely annotated with 562 pothole boxes measured in native 3D depth and vehicle classes (car/van/taxi) disambiguated through human-in-the-loop review.
Why this dataset
Most public driving datasets are clean, well-lit, and Western. This corpus targets exactly the opposite: unstructured, high-stress, degraded-infrastructure urban driving — the scenarios where production perception models fail. Full captured session (commercial license): ~9.6 GB, 4 takes, 16,012 frames, 4 synchronized modalities per frame — all from one scenario (night, low-visibility, degraded pavement). Additional scenarios (tunnels, dense traffic, rain, jaywalking zones) require future capture sessions and are not part of this corpus yet.
What's in this sample
- Frames: 641 (contiguous)
- Modalities: RGB (
.jpg), Infrared raw +ir_alignedreprojected to the color view (.png, ~97% visually complete — measured 3D reprojection + local depth inpainting + far-field fallback for sky/background), Raw metric depthdepth_raw_mm(.png, uint16, mm), Visual depthdepth_vis(.png) - Resolution: 896×504 native
- Labels: JSONL, 1,557 boxes —
car(879),pothole(562),lane_marking(73),truck(34),van(6),taxi(3) - Video: annotated preview (
promo_video.mp4), side-by-side RGB / IR (aligned) / depth, of this exact sequence - Note: sensor hardware includes an IMU; it was not enabled for this capture, so no motion data is included.
Annotation pipeline
- Custom ensemble auto-labeling: RT-DETR (COCO classes) + Grounding DINO (open-vocabulary, queried one class at a time to avoid multi-class label fusion).
- Confidence floor 0.40 at generation time.
- Human-in-the-loop review: cross-model conflict resolution (e.g. the same vehicle detected as both
carandvan), IoU-tracking-assisted correction propagation across frames, and manual annotation of objects the automated pipeline missed entirely (a large share of the potholes in this sample were human-found). - High-confidence (≥0.75), unambiguous detections auto-approved; everything else manually decided.
License
This sample: CC BY-NC 4.0 (non-commercial, evaluation/research use). Full ~9.6 GB session: commercial license (exclusive or non-exclusive) available — contact below.
Commercial licensing / custom data collection
The full session is available at 3 data tiers:
- Tier 1 — Raw Calibrated: RGB (anonymized) + IR + IR aligned + metric depth + calibration, no labels.
- Tier 2 — Raw + Auto-Labels: Tier 1 + raw ensemble auto-labels, clearly marked as not human-reviewed.
- Tier 3 — HITL-Reviewed: Tier 1 + fully human-verified, production-ready ground truth labels.
Interested in the full session (any tier), or a custom Intel RealSense D555 capture campaign (additional scenarios) in Latin America? Contact: hola@ysilabs.com
