datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lidar-warehouse-dataset
Dataset Card for LIDAR Warehouse Dayasey
This is a FiftyOne dataset with 3287 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/lidar-warehouse-dataset")
# Launch the App
session = fo.launch_app(dataset)
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/lidar-warehouse-dataset.HUGG_ARIAdual-lidar-umi-relativeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
12
],
"names": [
"umi1_x",
"umi1_y",
"umi1_z",
"umi1_rx",
"umi1_ry",
"umi1_rz",
"umi2_x"… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/dual-lidar-umi-relative.dual-lidar-umiThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
12
],
"names": [
"umi1_x",
"umi1_y",
"umi1_z",
"umi1_rx",
"umi1_ry",
"umi1_rz",
"umi2_x"… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/dual-lidar-umi.lidar-copc-san-antonioHUGG_ARIA_PINHOLE
HUGG ARIA Pinhole
This is a derived, frame-aligned pinhole version of the public
LIDAR-GT/HUGG_ARIA
dataset.
Each RGB frame is generated with the official HOT3D/Project Aria calibration
path:
Read stream 214-1 from recording.vrs at its exact TIME_CODE
timestamp.
Query the per-frame online FISHEYE624 calibration.
Query the per-frame online LINEAR calibration.
Warp the source image with projectaria_tools.core.calibration.distort_by_calibration.
This is pixel-equivalent to… See the full description on the dataset page: https://huggingface.co/datasets/LIDAR-GT/HUGG_ARIA_PINHOLE.LiDAR-Perfect-Depth-Datasetsdual-lidar-combined-filtered-long-gripper
Combined filtered dual-LiDAR UMI demonstrations
Observation-only LeRobot v3 derivative of brandonyang/dual-lidar-umi, brandonyang/dual-lidar-umi-relative. It contains 182 demonstrations (179951 frames) accepted by the continuous bimanual YAM replayability pipeline.
The 12-D observation.state contains the smoothed, trajectory-optimized YAM-achievable path in the zero-origin UMI Cartesian convention. Raw UMI gripper widths remain as separate observations. The two original UMI… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/dual-lidar-combined-filtered-long-gripper.robot-umi-lidar3d-e2eThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.images.front": {
"dtype": "video",
"shape": [
1200,
1920,
3
],
"names": [
"height",
"width",
"channels"
],
"info": {… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/robot-umi-lidar3d-e2e.dual-lidar-combined-filtered
Combined filtered dual-LiDAR UMI demonstrations
Observation-only LeRobot v3 derivative of brandonyang/dual-lidar-umi, brandonyang/dual-lidar-umi-relative. It contains 157 demonstrations (156492 frames) accepted by the continuous bimanual YAM replayability pipeline.
The 12-D observation.state contains the smoothed, trajectory-optimized YAM-achievable path in the zero-origin UMI Cartesian convention. Raw UMI gripper widths remain as separate observations. The two original UMI… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/dual-lidar-combined-filtered.dual-lidar-combined-filtered-joint-positions
Combined filtered dual-LiDAR UMI demonstrations
Observation-only LeRobot v3 derivative of brandonyang/dual-lidar-umi, brandonyang/dual-lidar-umi-relative. It contains 157 demonstrations (156492 frames) accepted by the continuous bimanual YAM replayability pipeline.
The 14-D observation.state contains left YAM joints 0–5, normalized left gripper, right YAM joints 0–5, and normalized right gripper. The two original UMI videos, timestamps, frame cadence, and task are preserved;… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/dual-lidar-combined-filtered-joint-positions.Multi-modal_dataset_for_LiDAR-Aided_CSI_Estimation
Dataset for "Synesthesia of Machines-Enhanced Wideband Multi-User CSI Learning with LiDAR Sensing"
📌 Overview
This dataset is constructed based on urban crossroad scenario from the M3SC dataset and includes CSI data between a roadside unit and three passing vehicles, as well as 64-line LiDAR point cloud data collected by the roadside devices (processed using the lightweight data processing approach proposed in the paper). The dataset contains a total of 4,500 samples… See the full description on the dataset page: https://huggingface.co/datasets/pku-pcni-lab/Multi-modal_dataset_for_LiDAR-Aided_CSI_Estimation.dual-lidar-combined-filtered-joint-positions-long-gripper
Combined filtered dual-LiDAR UMI demonstrations
Observation-only LeRobot v3 derivative of brandonyang/dual-lidar-umi, brandonyang/dual-lidar-umi-relative. It contains 182 demonstrations (179951 frames) accepted by the continuous bimanual YAM replayability pipeline.
The 14-D observation.state contains left YAM joints 0–5, normalized left gripper, right YAM joints 0–5, and normalized right gripper. The two original UMI videos, timestamps, frame cadence, and task are preserved;… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/dual-lidar-combined-filtered-joint-positions-long-gripper.ROBOMASTER-2025-LiDAR-ROSBAG
ROBOMASTER-2025 · 华北理工大学HORIZON战队 · LiDAR ROSBAG
📖 概述
数据来源: 华北理工大学 HORIZON 战队 — 雷达组依托平台: 华北理工 RM 创新实验室录制时间地点: ROBOMASTER 2025 超级对抗赛,北京理工大学(珠海)南部赛区现场实录数据用途: ROBOMASTER 场景下的点云识别、目标检测、三维建图等任务
🗂️ 数据概览
文件名
时长
大小
消息数
点云话题
RM-LiDAR-ROSBAG_01.bag
13分22秒
11.2 GB
8037
/cloudpoints
RM-LiDAR-ROSBAG_02.bag
13分59秒
12.9 GB
8399
/cloudpoints
数据格式为标准 ROS 1 .bag 文件,未压缩,采样频率约为 10 Hz。
🎥… See the full description on the dataset page: https://huggingface.co/datasets/BreCaspian/ROBOMASTER-2025-LiDAR-ROSBAG.egostation-iphone-lidar-household-v1
Zen-O Household Manipulation, iPhone LiDAR
8 first-person recordings of ordinary household work, with both hands
tracked in three dimensions and in real metres, laid out in LeRobot v2.1.
Episodes
8
Frames
128,833 at 30 fps, about 71.6 minutes
Video
observation.images.head, 1920x1440
State
7 floats, camera position and orientation
Action
20 floats, both wrists and both grippers
Coordinate frame
ROS REP 103, X forward, Y left, Z up, metric… See the full description on the dataset page: https://huggingface.co/datasets/zeno-labs/egostation-iphone-lidar-household-v1.lidar-localizationosmo360-lidar-preview
Osmo 360 Native Captures
Original DJI Osmo 360 recordings (native formats).
Files
Pair
.OSV (master)
.LRF (proxy)
CAM_20260728164614_0006_D
~16 GB
~1.2 GB
CAM_20260728170517_0007_D
~16 GB
~1.2 GB
CAM_20260728172420_0008_D
~4.4 GB
~0.3 GB
.OSV: full-resolution native Osmo 360 capture
.LRF: low-resolution proxy for quick preview
Open with DJI Mimo / DJI Studio.
robot-umi-lidar-sdk-e2eThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.images.front": {
"dtype": "video",
"shape": [
1200,
1920,
3
],
"names": [
"height",
"width",
"channels"
],
"info": {… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/robot-umi-lidar-sdk-e2e.psegs-ios-lidar-ext
PSegs iOS Lidar Extension
This project contains data captured using Lidar-equipped iPhone(s)
for use as an extension with the
PSegs project.
Structure
threeDScannerApp_data - This is test data captured
using the 3D Scanner App for iOS.
ps_external_test_fixtures - These are fixtures
created using the data in this repo and code in
PSegs. They are hosted here and
provided to power PSegs unit tests.
dual-lidar-umi-relative-filtered
Filtered dual-LiDAR UMI demonstrations
Observation-only LeRobot v3 derivative of brandonyang/dual-lidar-umi-relative. It contains 65 demonstrations (63383 frames) that pass the complete continuous YAM replayability classification.
observation.state retains the original 12-D UMI Cartesian schema. Its values are the smoothed, trajectory-optimized YAM-achievable FK path mapped back into the UMI coordinate convention. Gripper observations, videos, timestamps, and tasks are preserved… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/dual-lidar-umi-relative-filtered.HUGG_ARIA_GAUSSIANSLiDAR-LLM-Nu-Caption
Dataset Details
Dataset type:
This is the nu-Caption dataset, a QA dataset designed for training MLLM models on caption tasks in autonomous driving scenarios. It is built upon the NuScenes dataset.
Dataset keys:
"answer" is the output of the VLM models using image data. "answer_lidar" uses GPT4O-mini to filter information that cannot be obtained from the image data.
If you want to train the model like LiDAR-LLM, which only uses the LiDAR modality and does not use the vision modality… See the full description on the dataset page: https://huggingface.co/datasets/Senqiao/LiDAR-LLM-Nu-Caption.LiDARDustXdual-lidar-umi-independentThis dataset was created using LeRobot.
Dataset Description
Standalone LeRobot v3 dataset containing 296 dual-UMI orange-collection demonstrations (276,332 frames, 2.559 hours at 30 FPS). It contains synchronized observation.images.umi1 and observation.images.umi2 video observations, 14D observation.state, and 14D action; no LiDAR files or LiDAR frame features are included.
For each UMI independently, pose is expressed relative to that UMI's episode-start pose using… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/dual-lidar-umi-independent.dual-lidar-umi-testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
12
],
"names": [
"umi1_x",
"umi1_y",
"umi1_z",
"umi1_rx",
"umi1_ry",
"umi1_rz",
"umi2_x"… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/dual-lidar-umi-test.dual-lidar-combined-filtered-joint-positions-long-gripper-trainable
Long-gripper dual-UMI BiYAM joints
Published 14-D states are preserved exactly; see meta/materialization.json.
LiDAR_HD
Dataset for the repository "MoE-DSI-3D".
Motivation
To evaluate the applicability of DSI-3D and MoE-DSI-3D on a large-scale dataset, we use the LiDAR HD (LiDAR Haute Densité) dataset. LiDAR HD covers the French territory and provides an average point density of approximately 10 points/m².
The dataset was acquired through a country-wide, tile-based LiDAR campaign covering a wide variety of environments.
We selected the Paris area as our primary study region… See the full description on the dataset page: https://huggingface.co/datasets/Chahine-Nicolas/LiDAR_HD.LiDAR-LLM-Nu-Grounding
Dataset Details
Dataset type:
This is the nu-Grounding dataset, a QA dataset designed for training MLLM models for grounding in autonomous driving scenarios. This QA dataset is built upon the NuScenes dataset.
Where to send questions or comments about the dataset:
https://github.com/Yangsenqiao/LiDAR-LLM
Project Page:
https://sites.google.com/view/lidar-llm
Paper:
https://arxiv.org/abs/2312.14074
osmo360-lidar-preview
Osmo 360 Native Captures
Original DJI Osmo 360 recordings (native formats).
Files
Pair
.OSV (master)
.LRF (proxy)
CAM_20260728164614_0006_D
~16 GB
~1.2 GB
CAM_20260728170517_0007_D
~16 GB
~1.2 GB
CAM_20260728172420_0008_D
~4.4 GB
~0.3 GB
.OSV: full-resolution native Osmo 360 capture
.LRF: low-resolution proxy for quick preview
Open with DJI Mimo / DJI Studio.
lidar_degeneracy_datasets
LiDAR Degeneracy Datasets
Storage/backup for ntnu-arl/lidar_degeneracy_datasets
