datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CarlaOcc
Database_structure
CarlaOcc/
├── CarlaOccV1/
│ ├── calib/
│ │ └── calib.yaml
│ ├── splits/
│ │ ├── test.txt
│ │ ├── train.txt
│ │ └── val.txt
│ ├── SceneMeshes/
│ │ ├── fg_actors/
│ │ ├── fg_actor_occ/
│ │ └── TownXX_Opt/
│ │ ├── bg_actors/
│ │ └── bg_actor_occ/
│ ├── TownXX_Opt_SeqXX/
│ │ ├── poses/
│ │ │ ├── cam_00.txt
│ │ │ └── lidar.txt
│ │ ├── rgb/
│ │ │ ├── image_00/
│ │ │ │ ├── 0000.png… See the full description on the dataset page: https://huggingface.co/datasets/fengyi233/CarlaOcc.npm3d-kitti-carlacarla-autopilot-multimodal-dataset
CARLA Autopilot Multimodal Dataset
This dataset contains synchronized multimodal driving data collected in the CARLA simulator using the autopilot feature. It provides RGB images from multiple cameras, semantic segmentation, LiDAR point clouds, 2D bounding boxes, and ego-vehicle state/control signals across varied weather, maps, and traffic densities.
The dataset is designed for research in autonomous driving, sensor fusion, imitation learning, and self-driving evaluation.… See the full description on the dataset page: https://huggingface.co/datasets/immanuelpeter/carla-autopilot-multimodal-dataset.CARLA_TLSRautonomous-driving-carla
CARLA Autonomous Driving Dataset
Custom datasets for autonomous driving in CARLA simulator
Created for CMPE 789 - Robot Perception at Rochester Institute of Technology
📊 Dataset Overview
This repository contains two custom-generated datasets from the CARLA 0.9.15 simulator for training autonomous driving perception models:
Dataset
Task
Images
Format
Size
YOLO Dataset
Object Detection
4,000
YOLOv8/v11
~1.2 GB
UFLD Dataset
Lane Detection
10,000… See the full description on the dataset page: https://huggingface.co/datasets/jkdxbns/autonomous-driving-carla.demo-tabular-benchmark-containers
📦 Carla HQ — Tabular Benchmark CuratedContainers
Centralized repository of Data Foundry CuratedContainers curated for Carla HQ, TabICLv2, and the next generation of Tabular Foundation Models (TabPFN, EXAONE, Google TabFM).
Each container directory provides:
Columnar Parquet Data (dataset.parquet): Clean, type-normalized, and validated tabular dataset binary.
Standardized Task Molds (task_metadata.predictive-ml-task-mold-v1.json): Problem definitions, target attributes, and… See the full description on the dataset page: https://huggingface.co/datasets/carlahq/demo-tabular-benchmark-containers.demo-tabular-benchmarks
📊 Carla HQ Tabular Foundation Model Benchmarks
Centralized benchmark repository of canonical tabular datasets curated for Carla HQ and TabICL (In-Context Learning foundation models for tabular data).
Each dataset is hosted as an independent subset/config with native Parquet storage, schema qualities, OpenML source links, and synchronized Google Sheets for live spreadsheet experimentation.
🚀 Quickstart & Download Options
Option 1: Using… See the full description on the dataset page: https://huggingface.co/datasets/carlahq/demo-tabular-benchmarks.CARLA_sim_data
Paired Event-Camera Collision Benchmark
The official experiment cohort contains 880 recordings / 440 matched pairs.
Each pair contains one collision and one near miss, kept in the same split.
The supplied split manifests define the fixed evaluation cohort.
Scenario
Train videos
Val videos
Test videos
Total
Head-on
66
26
26
118
Close turning
86
24
30
140
Following/braking
52
42
44
138
Pedestrian walking past
66
32
32
130
Pedestrian walking then stopping
8
12
12… See the full description on the dataset page: https://huggingface.co/datasets/Bmingg/CARLA_sim_data.carla_hdpdm_carlacarla_data_60kmini-carla-192x320-wan-2p2-vae
mini-carla-192x320-wan-2p2-vae
Wan2.2-VAE-encoded latents of a small CARLA driving dataset (192x320, native resolution,
no resize). Produced for training miniworld, a
minimal flow-matching world-model framework, by caching pixel clips through the frozen
pretrained Wan2.2 video VAE instead of a locally-trained one.
Data size
Source pixel dataset
mini_192x320_low: 320 episodes, 600 frames each (192x320, 20 fps) — ~33 GB
Clips in this cache
1,920 (6… See the full description on the dataset page: https://huggingface.co/datasets/kamwoh/mini-carla-192x320-wan-2p2-vae.carla_car_body_maskcarla_collision
CARLA Collision Video, Pose Pairs and Scene Captions
2,248 source videos and 3,764 captioned five-second windows, collected in
CARLA 0.9.16 at 16 Hz, 1280 x 704, 90-degree horizontal FOV. Each source
contains a deliberately constructed collision encounter, actual and constructed
input camera trajectories, physical vehicle controls, and collision sensor events.
Each indexed window has an English scene_static caption generated from all
81 consecutive images by Qwen/Qwen3.8-27B.… See the full description on the dataset page: https://huggingface.co/datasets/xlu11/carla_collision.carla-autopilot-images
CARLA Autopilot Images Dataset
Note: A newer, extended version of this dataset is available.🤗 CARLA Autopilot Multimodal Dataset 🤗It includes semantic segmentation, LiDAR, 2D bounding boxes, and additional environment metadata.Use it if your research requires multimodal signals beyond the RGB images and vehicle state/control data provided here.
This dataset contains autonomous driving data collected from CARLA simulator using autopilot.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/immanuelpeter/carla-autopilot-images.dontpatronizeme_pcl------------------------------------------------------ DISCLAIMER ------------------------------------------------------
The Don’t Patronize Me! dataset has been created for research purposes. Patronizing and Condescending Language (PCL) towards vulnerable communities is understood in
this dataset as a commonly used, generally unconscious and well intended writing style. We consider that the authors of the paragraphs included in this dataset do not
intend any harm towards the vulnerable… See the full description on the dataset page: https://huggingface.co/datasets/carlaperez/dontpatronizeme_pcl.Well_actually_mansplainingcarla-simlingo-raw
SimLingo CARLA Dataset (Raw, 4Hz)
Raw driving data from CARLA simulator. No transformations or derived fields - all original measurements preserved as-is.
Dataset Summary
Source: SimLingo (CVPR 2025)
Scale: 228,757 frames (23 shards)
Frame Rate: 4 FPS
Resolution: 1024x512 RGB
Routes: Complete driving episodes (routes never split across shards)
Column Schema
Core Fields
Column
Type
Description
route_id
string
Route identifier
frame_idx… See the full description on the dataset page: https://huggingface.co/datasets/TESS-Computer/carla-simlingo-raw.HumanDrive-CARLA
G29-HumanDrive-CARLA-5Runs
Dataset Description
This dataset contains 70,308 samples collected in CARLA using a Logitech G29 steering wheel.
It includes one representative run from each town:
runs_01 → Town01
runs_02 → Town03
runs_03 → Town04
runs_04 → Town05
runs_05 → Town07
Each sample contains:
RGB image embedded in JPEG
semantic segmentation image embedded in PNG
steering command
throttle command
brake command
DAgger flag
town label
run identifier
frame identifier… See the full description on the dataset page: https://huggingface.co/datasets/roboticslaburjc/HumanDrive-CARLA.Carla_dataset_town10carla-expert-racing
URJC-DeepRacer: Autonomous Driving Dataset
This dataset was generated by Sergio Robledo as part of the URJC-DeepRacer project, focused on training and validating autonomous driving agents using Deep Learning and Reinforcement Learning techniques.
Data was collected using the CARLA Simulator, featuring DeepRacer vehicle models and custom-designed racing environments.
📊 Dataset Overview
Each entry provides a synchronized capture of the front-facing RGB camera, its… See the full description on the dataset page: https://huggingface.co/datasets/urjc-deepracer/carla-expert-racing.carla_datasetcarla_text_diversity_test
Carla Text-to-Scene Generation Diveristy Test
This dataset is to test the diversity for scene generation under different scenarios.
Dataset Details
There are 8 different scenarios including:
blocking
crushing
cut-off
daily-traffic
emergency
intersection
two-wheels
weather
with prompt template like:
Please create a scene for {}
For more details, please refer to the paper.
Citation
@article{ruan2024ttsg,
title={Traffic Scene Generation from Natural Language… See the full description on the dataset page: https://huggingface.co/datasets/Justin900/carla_text_diversity_test.Carla-Road-X1
Carla-Road-X1: OpenDRIVE Road Network Generation Dataset
Text-to-xodr dataset for training language models to generate OpenDRIVE (.xodr) road network files from natural language descriptions.
Dataset Summary
Train samples: 385,929
Val samples: 42,882
Template samples: 40 (train: 36, val: 4)
Format: JSONL with Gemma chat template (system/user/assistant messages)
Data Sources
Source
Count
Description
OSM-converted xodr sub-networks
~428K… See the full description on the dataset page: https://huggingface.co/datasets/NCUT-AI/Carla-Road-X1.TedDescripción:
Dataset creado a partir de un pequeño corpus paralelo inglés-español. El corpus está formado por las transcripciones en inglés y español de tres charlas Ted, alineadas con LF Aligner.
MotoRisk-CARLA
MotoRisk-CARLA
MotoRisk-CARLA is a CARLA-based motorcycle riding dataset for accident detection and unsafe riding behavior research. It contains labeled session CSV files for normal riding, reckless/aggressive riding, zigzag/weaving, and accident scenarios collected with simulated motorcycle sensor data.
The dataset is organized into four folders: Accident, Normal_Riding, Reckless_Aggressive, and Zigzag_Weaving.
MotoRisk-CARLA
MotoRisk-CARLA
MotoRisk-CARLA is a CARLA-based motorcycle riding dataset for accident detection and unsafe riding behavior research. It contains labeled session CSV files for normal riding, reckless/aggressive riding, zigzag/weaving, and accident scenarios collected with simulated motorcycle sensor data.
The dataset is organized into four folders: Accident, Normal_Riding, Reckless_Aggressive, and Zigzag_Weaving.
CarlaSimwoman-carlaFullCarla
