CoolFace
Datasetpublic

ysilabs/latam-urban-driving-sample

LATAM Urban Driving Dataset — Night / Low-Visibility / Potholes Sample (Intel RealSense D555) A free preview of a multimodal, edge-case-focused urban driving dataset captured in the State of Mexico (Edomex) with the Intel RealSense D555 (RGB + IR + metric depth), aimed at Autonomous Driving, ADAS, and Smart City perception teams. This sample: 641 contiguous night-driving frames, densely annotated with 562 pothole boxes measured in native 3D depth and vehicle classes… See the full description on the dataset page: https://huggingface.co/datasets/ysilabs/latam-urban-driving-sample.

sourceHugging Facecc-by-nc-4.0updated 5d agoView on Hugging Face
0likes47downloads
Dataset Card

LATAM Urban Driving Dataset — Night / Low-Visibility / Potholes Sample (Intel RealSense D555)

A free preview of a multimodal, edge-case-focused urban driving dataset captured in the State of Mexico (Edomex) with the Intel RealSense D555 (RGB + IR + metric depth), aimed at Autonomous Driving, ADAS, and Smart City perception teams.

This sample: 641 contiguous night-driving frames, densely annotated with 562 pothole boxes measured in native 3D depth and vehicle classes (car/van/taxi) disambiguated through human-in-the-loop review.

Why this dataset

Most public driving datasets are clean, well-lit, and Western. This corpus targets exactly the opposite: unstructured, high-stress, degraded-infrastructure urban driving — the scenarios where production perception models fail. Full captured session (commercial license): ~9.6 GB, 4 takes, 16,012 frames, 4 synchronized modalities per frame — all from one scenario (night, low-visibility, degraded pavement). Additional scenarios (tunnels, dense traffic, rain, jaywalking zones) require future capture sessions and are not part of this corpus yet.

What's in this sample

  • —Frames: 641 (contiguous)
  • —Modalities: RGB (.jpg), Infrared raw + ir_aligned reprojected to the color view (.png, ~97% visually complete — measured 3D reprojection + local depth inpainting + far-field fallback for sky/background), Raw metric depth depth_raw_mm (.png, uint16, mm), Visual depth depth_vis (.png)
  • —Resolution: 896×504 native
  • —Labels: JSONL, 1,557 boxes — car (879), pothole (562), lane_marking (73), truck (34), van (6), taxi (3)
  • —Video: annotated preview (promo_video.mp4), side-by-side RGB / IR (aligned) / depth, of this exact sequence
  • —Note: sensor hardware includes an IMU; it was not enabled for this capture, so no motion data is included.

Annotation pipeline

  1. 1.Custom ensemble auto-labeling: RT-DETR (COCO classes) + Grounding DINO (open-vocabulary, queried one class at a time to avoid multi-class label fusion).
  2. 2.Confidence floor 0.40 at generation time.
  3. 3.Human-in-the-loop review: cross-model conflict resolution (e.g. the same vehicle detected as both car and van), IoU-tracking-assisted correction propagation across frames, and manual annotation of objects the automated pipeline missed entirely (a large share of the potholes in this sample were human-found).
  4. 4.High-confidence (≥0.75), unambiguous detections auto-approved; everything else manually decided.

License

This sample: CC BY-NC 4.0 (non-commercial, evaluation/research use). Full ~9.6 GB session: commercial license (exclusive or non-exclusive) available — contact below.

Commercial licensing / custom data collection

The full session is available at 3 data tiers:

  • —Tier 1 — Raw Calibrated: RGB (anonymized) + IR + IR aligned + metric depth + calibration, no labels.
  • —Tier 2 — Raw + Auto-Labels: Tier 1 + raw ensemble auto-labels, clearly marked as not human-reviewed.
  • —Tier 3 — HITL-Reviewed: Tier 1 + fully human-verified, production-ready ground truth labels.

Interested in the full session (any tier), or a custom Intel RealSense D555 capture campaign (additional scenarios) in Latin America? Contact: hola@ysilabs.com