monocular
monocular-geometry-evaluationProcessed versions of some open-source datasets for evaluation of monocular geometry estimation.
Dataset
Source
Publication
Num images
Storage Size
Note
NYUv2
NYU Depth Dataset V2
[1]
654
243 MB
Offical test split. Mirror, glass and window manually removed. Depth beyound 5 m truncated.
KITTI
KITTI Vision Benchmark Suite
[2, 3]
652
246 MB
Eigen's test split.
ETH3D
ETH3D SLAM & Stereo Benchmarks
[4]
454
1.3 GB
Downsized from 6202×4135 to 2048×1365
iBims-1
iBims-1 (independent… See the full description on the dataset page: https://huggingface.co/datasets/Ruicheng/monocular-geometry-evaluation.Monocular_Depth_Essentials
Monocular_Depth_Essentials
Monocular_Depth_Essentials is a lightweight, curated core dataset tailored specifically for Monocular Depth Estimation tasks.
Following the data preparation guidelines from the classic bts repository , this dataset extracts only the essential image-depth pairs from the massive raw KITTI and NYU Depth V2 datasets based strictly on the official Eigen Split train/test text lists.
If you are benchmarking or reproducing Depth Anything (V1/V2), BTS, or… See the full description on the dataset page: https://huggingface.co/datasets/Kai-Yin-UoA/Monocular_Depth_Essentials.MonocularSamples
DatraAI monocular camera-IMU sample
Episodes: 6
Raw bytes: 1351783147
Total camera-duration seconds: 537.833334
Scope
This dataset contains six monocular camera-IMU episodes. Four are complete source episodes. Episodes 000004 and 000006 contain the first 120 seconds of recording2 and recording4.
The two clips preserve the original encoded video frames. Their native VTS and IMU sidecars were shortened to the same camera interval. Sensor values were not interpolated… See the full description on the dataset page: https://huggingface.co/datasets/datraailab/MonocularSamples.APAC-Egocentric-Monocular-Labeled
APAC Egocentric Monocular (Labeled)
Twelve labeled egocentric work sequences from industrial, hospitality, logistics and retail settings, released as hand-tracking and head-tracking renders with 252 densely captioned 3-second action segments.
Preview: polishing a car in an automotive garage — hand-tracking render, downscaled to 720p.
At a glance
Samples
12
Total duration
12 min 09 s (~61 s each)
Resolution
1920×1080 @ 30 fps
Action segments… See the full description on the dataset page: https://huggingface.co/datasets/humyn-labs/APAC-Egocentric-Monocular-Labeled.CampusDepth-A-Large-Scale-Day-Night-RGB-Dataset-for-Monocular-Depth-Estimation
Advanced Driver Assistance Systems
(ADAS) require accurate and reliable perception of the
surrounding environment to ensure vehicle safety and
reduce the risk of collisions. Depth estimation plays a
crucial role in understanding object distance and spatial
relationships in traffic scenes. Traditional depth sensing
approaches, such as stereo camera systems and LiDAR,
provide accurate depth information but suffer from high
cost, increased hardware complexity, calibration… See the full description on the dataset page: https://huggingface.co/datasets/Vaibhav14/CampusDepth-A-Large-Scale-Day-Night-RGB-Dataset-for-Monocular-Depth-Estimation.KITTI-monocular
