datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CarlaOcc
Database_structure
CarlaOcc/
├── CarlaOccV1/
│ ├── calib/
│ │ └── calib.yaml
│ ├── splits/
│ │ ├── test.txt
│ │ ├── train.txt
│ │ └── val.txt
│ ├── SceneMeshes/
│ │ ├── fg_actors/
│ │ ├── fg_actor_occ/
│ │ └── TownXX_Opt/
│ │ ├── bg_actors/
│ │ └── bg_actor_occ/
│ ├── TownXX_Opt_SeqXX/
│ │ ├── poses/
│ │ │ ├── cam_00.txt
│ │ │ └── lidar.txt
│ │ ├── rgb/
│ │ │ ├── image_00/
│ │ │ │ ├── 0000.png… See the full description on the dataset page: https://huggingface.co/datasets/fengyi233/CarlaOcc.IPL-CARLA-dataset
IPL-CARLA-dataset
Autonomous driving semantic segmentation dataset created with CARLA (Cars Learning to Act) simulator.
Dataset information
Images are generated from two different simulated cities. They include different weather (sunny, foggy and rainy) and daytime (morning, day, sunset and night) conditions. It contains 20000 RGB-rendered images and their corresponding ground truth segmented masks. Segmentation ground truth masks have 35 different classes with colors… See the full description on the dataset page: https://huggingface.co/datasets/isp-uv-es/IPL-CARLA-dataset.carla100npm3d-kitti-carlacarla-dataset-ped2
CARLA Dataset — Pedestrian (Town05) · Extension 2
A large-scale pedestrian-following driving dataset captured from the
CARLA simulator, all in Town05. Provides synchronized RGB + depth + camera
parameters along each pedestrian trajectory. Part of the training data for the Seoul World Model.
Stored in WebDataset (.tar) format for efficient streaming.
Dataset at a glance
Metric
Value
Town
Town05 only
Actor
Pedestrian
Frames / scene
200
Scenes /… See the full description on the dataset page: https://huggingface.co/datasets/mkxdxd/carla-dataset-ped2.CARLA_OOD
Dataset Card for Carla_OOD Dataset
Dataset Summary
The Carla_OOD Dataset, created using the CARLA v0.9.13 simulator, mimics the sensor configuration of the KITTI dataset with a Velodyne HDL64 LiDAR and cameras aligned with KITTI's Camera0. It is designed to advance research in anomaly detection and segmentation in autonomous systems, addressing challenges in handling unexpected multi-modal inputs.
Supported Tasks
The Carla_OOD Dataset can be used to compare… See the full description on the dataset page: https://huggingface.co/datasets/Mona4399/CARLA_OOD.PDM_Lite_Carla_LB2
PDM-Lite Dataset for CARLA Leaderboard 2.0
Description
PDM-Lite is a state-of-the-art rule-based expert system for autonomous urban driving in CARLA Leaderboard 2.0, and the first to successfully navigate all scenarios. This dataset was used to create the QA dataset for DriveLM-Carla, a benchmark for evaluating end-to-end autonomous driving algorithms with Graph Visual Question Answering (GVQA). DriveLM introduces GVQA as a novel approach, modeling perception, prediction… See the full description on the dataset page: https://huggingface.co/datasets/autonomousvision/PDM_Lite_Carla_LB2.segmentation-carla-drivingThis dataset consists of 80 episodes of driving data collected using an autopilot agent in CARLA simulator for training imitation learning models for autonomous driving tasks.
Each frame is structured as follows:
frame_data = {
'frame': the frame index,
'hlc': an integer representing the high-level command,
'light': an integer representing current traffic light status,
'controls': an array of [throttle, steer, brake],
'measurements':… See the full description on the dataset page: https://huggingface.co/datasets/nightmare-nectarine/segmentation-carla-driving.carla-autopilot-multimodal-dataset
CARLA Autopilot Multimodal Dataset
This dataset contains synchronized multimodal driving data collected in the CARLA simulator using the autopilot feature. It provides RGB images from multiple cameras, semantic segmentation, LiDAR point clouds, 2D bounding boxes, and ego-vehicle state/control signals across varied weather, maps, and traffic densities.
The dataset is designed for research in autonomous driving, sensor fusion, imitation learning, and self-driving evaluation.… See the full description on the dataset page: https://huggingface.co/datasets/immanuelpeter/carla-autopilot-multimodal-dataset.carla-ue5-native-erp-2048x1024-5m-track-100kreal-carla-assets-anon
CARLA Modernized Assets (Unreal Engine 5.7)
A cleaned, modernized port of the CARLA simulator art
content, re-imported for Unreal Engine 5.7 (Nanite / Lumen / Substrate).
This repository contains the Unreal project's Content/ tree plus the project
and config files needed to open it in the editor.
What's included
CarlaAssets.uproject # UE 5.7 project file
Config/ # Engine/Editor/Game/Input .ini (anonymized)
Content/
├── Carla/
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/allude-occluded/real-carla-assets-anon.SidewalkPilot_carla
SidewalkPilot CARLA Synthetic Dataset
CARLA-simulator-generated steering + throttle frames used to assist the SidewalkPilot Series 1/2 models — blended with real RC-car photos and down-weighted vs real. This is synthetic data rendered in the CARLA driving simulator, not real field capture.
Series 3 does NOT use this dataset — the Series 3 line is trained on real RC-car photos only. This CARLA set is kept for the CARLA-assisted Series 1/2 history and for optional future sim2real… See the full description on the dataset page: https://huggingface.co/datasets/ram-shreyas-naik-sabavat/SidewalkPilot_carla.tcp_carla_dataCARLA_TLSRautonomous-driving-carla
CARLA Autonomous Driving Dataset
Custom datasets for autonomous driving in CARLA simulator
Created for CMPE 789 - Robot Perception at Rochester Institute of Technology
📊 Dataset Overview
This repository contains two custom-generated datasets from the CARLA 0.9.15 simulator for training autonomous driving perception models:
Dataset
Task
Images
Format
Size
YOLO Dataset
Object Detection
4,000
YOLOv8/v11
~1.2 GB
UFLD Dataset
Lane Detection
10,000… See the full description on the dataset page: https://huggingface.co/datasets/jkdxbns/autonomous-driving-carla.demo-tabular-benchmark-containers
📦 Carla HQ — Tabular Benchmark CuratedContainers
Centralized repository of Data Foundry CuratedContainers curated for Carla HQ, TabICLv2, and the next generation of Tabular Foundation Models (TabPFN, EXAONE, Google TabFM).
Each container directory provides:
Columnar Parquet Data (dataset.parquet): Clean, type-normalized, and validated tabular dataset binary.
Standardized Task Molds (task_metadata.predictive-ml-task-mold-v1.json): Problem definitions, target attributes, and… See the full description on the dataset page: https://huggingface.co/datasets/carlahq/demo-tabular-benchmark-containers.plentiful-carla-camera-rigs
Benchmark: Plentiful Carla Camera Rigs
Camera-based perception systems for autonomous driving are typically developed and evaluated using fixed sensor rigs,
while real-world vehicle fleets exhibit substantial variation in camera placement, orientation, field of view, and camera count.
This mismatch introduces a cross-rig domain gap in which only the geometric observation process changes.
To study this effect under controlled conditions, we introduce Plentiful Carla Camera… See the full description on the dataset page: https://huggingface.co/datasets/timb2001/plentiful-carla-camera-rigs.demo-tabular-benchmarks
📊 Carla HQ Tabular Foundation Model Benchmarks
Centralized benchmark repository of canonical tabular datasets curated for Carla HQ and TabICL (In-Context Learning foundation models for tabular data).
Each dataset is hosted as an independent subset/config with native Parquet storage, schema qualities, OpenML source links, and synchronized Google Sheets for live spreadsheet experimentation.
🚀 Quickstart & Download Options
Option 1: Using… See the full description on the dataset page: https://huggingface.co/datasets/carlahq/demo-tabular-benchmarks.map-carla-kashiwanoha
Kashiwanoha test map
A map of a real site, the Kashiwanoha Campus area in Kashiwa, Chiba. It holds the same roads twice: as a lanelet2 map for Autoware, and as an OpenDRIVE for CARLA. CARLA has no map of this area, so it needs the road network in a form that it can build a world from.
Contents
File
Read by
What it is
lanelet2_map.osm
Autoware
The road network in lanelet2 form
map_projector_info.yaml
Autoware
The origin of the map
kashiwanoha.xodr… See the full description on the dataset page: https://huggingface.co/datasets/AutowareFoundation/map-carla-kashiwanoha.carla-ue5-maps
CARLA UE5 Autoware maps
Autoware maps for the CARLA towns as they exist in CARLA 0.10, the release that re-authored the towns in Unreal Engine 5. One directory per CARLA world.
The point cloud maps published for Autoware alongside the CARLA towns were recorded on CARLA 0.9. The 0.10 towns are new geometry, so those maps no longer describe the world a vehicle drives in. The point clouds here were recorded from CARLA 0.10 itself.
Contents
autoware_maps/
└──… See the full description on the dataset page: https://huggingface.co/datasets/AutowareFoundation/carla-ue5-maps.mini-carla-192x320-v3
mini-carla-192x320-v3
Action- and camera-pose-conditioned driving clips rendered offline from
CARLA 0.9.16 — the v3 scale-up of
mini-carla-192x320,
built as the training corpus for
miniworld, a minimal flow-matching
world-model framework.
10.48 M frames / 145.6 hours at 192×320 (h×w), 20 Hz, across 21
training environments (7 towns × 3 camera regimes), plus a held-out-town
validation split (Town07, all three regimes).
train
val
Episodes
17,471
246
Frames
10,482,600… See the full description on the dataset page: https://huggingface.co/datasets/kamwoh/mini-carla-192x320-v3.carla-nuscenespasa-dataset
PaSa Dataset
Data for the paper: PaSa: An LLM Agent for Comprehensive Academic Paper Search
Further information: https://github.com/bytedance/pasa
PaSa: An LLM Agent for Comprehensive Academic Paper Search
Yichen He, Guanhua Huang, Peiyuan Feng, Yuan Lin, Yuchen Zhang, Hang Li, Weinan E
Paper link: https://arxiv.org/abs/2501.10120
CARLA_sim_data
Paired Event-Camera Collision Benchmark
The official experiment cohort contains 880 recordings / 440 matched pairs.
Each pair contains one collision and one near miss, kept in the same split.
The supplied split manifests define the fixed evaluation cohort.
Scenario
Train videos
Val videos
Test videos
Total
Head-on
66
26
26
118
Close turning
86
24
30
140
Following/braking
52
42
44
138
Pedestrian walking past
66
32
32
130
Pedestrian walking then stopping
8
12
12… See the full description on the dataset page: https://huggingface.co/datasets/Bmingg/CARLA_sim_data.carla_hdmini-carla-192x320
mini-carla-192x320
Action-conditioned driving clips rendered offline from CARLA 0.9.16, built as the
training set for miniworld, a minimal
flow-matching world-model framework.
192,000 frames at 192×320 (h×w), 20 Hz, across 16 environments — 8 CARLA towns
seen from two camera regimes.
Frames
192,000 (12,000 per env × 16 envs)
Episodes
320 (20 per env, 600 frames = 30 s each)
Resolution
192 × 320 (h × w), RGB uint8
Frame rate
20 Hz (fixed_delta_seconds = 0.05)… See the full description on the dataset page: https://huggingface.co/datasets/kamwoh/mini-carla-192x320.CARLA-MWRS
CARLA-MWRS
CARLA-MWRS (CARLA Multi-Weather Road Segmentation) is a fully synthetic,
paired RGB--geometry benchmark for urban-road segmentation. It was prepared
for IAF-Net: Illumination-Adaptive Fusion for Low-Light Urban Road
Segmentation. The canonical protocol deliberately holds out a town:
Town05 and Town06 are used for training, and Town04 is used only for
validation. The four weather conditions are balanced within each split.
Release contents
split… See the full description on the dataset page: https://huggingface.co/datasets/PeterNano/CARLA-MWRS.pdm_carlacarla_data_60kAD_pdm-lite_carla-lb2This is just a reproduction of https://huggingface.co/datasets/autonomousvision/PDM_Lite_Carla_LB2.
