datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CameraBench
📷 CameraBench: Towards Understanding Camera Motions in Any Video
SfMs and VLMs performance on CameraBench: Generative VLMs (evaluated with VQAScore) trail classical SfM/SLAM in pure geometry, yet they outperform discriminative VLMs that rely on CLIPScore/ITMScore and—even better—capture scene‑aware semantic cues missed by SfM
After simple supervised fine‑tuning (SFT) on ≈1,400 extra annotated clips, our 7B Qwen2.5‑VL doubles its AP, outperforming the current best… See the full description on the dataset page: https://huggingface.co/datasets/syCen/CameraBench.IDLE-OO-Camera-Traps
Dataset Card for IDLE-OO Camera Traps
IDLE-OO Camera Traps is a 5-dataset benchmark of camera trap images from the Labeled Information Library of Alexandria: Biology and Conservation (LILA BC) with a total of 2,586 images for species classification. Each of the 5 benchmarks is balanced to have the same number of images for each species within it (between 310 and 1120 images), representing between 16 and 39 species.
Supported Tasks and Leaderboards
Image… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/IDLE-OO-Camera-Traps.camera_pizza_additionalcamerabench_cutGPIC-Camera
GPIC-Camera
Per-image camera parameter annotations for the GPIC dataset
(train / test / val; train = 8,000 shards, test = 1,000,000 images, val = 200,000 images), captioned by the Puffin-World model. More captioned datasets are provided in our Puffin-16M website.
The collage above visualizes the camera maps on sample images — each
pair shows the up field (green arrows: the projected gravity-up direction)
and the latitude field (colored contours: angle above/below the horizon).… See the full description on the dataset page: https://huggingface.co/datasets/KangLiao/GPIC-Camera.multi-view-bathroom-scene-understanding-camera-relocalization
Multi-View Bathroom Scene Understanding & Camera Relocalization
Generated by datapack-import.ts
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is… See the full description on the dataset page: https://huggingface.co/datasets/physicl/multi-view-bathroom-scene-understanding-camera-relocalization.pseudo-camera-10k
pseudo-camera-10k dataset
Contents
This dataset contains 10k free images from world class photographers. The images have been resized using Lanczos antialiasing, with their smaller edge shifted to 1024px.
The aim of this dataset is a highly variable but high quality and high resolution set of images containing difficult concepts, with about half of the images being numbered group shots and family portraits with the number of subjects labeled.
No images were upsampled in… See the full description on the dataset page: https://huggingface.co/datasets/bghira/pseudo-camera-10k.camera
Dataset Card for CAMERA📷:
Table of Contents:
Dataset Card for Camera
Table of Contents
Dataset Details
Dataset Description
Dataset Sources
Uses
Direct Use
Dataset Information
Data Example
Dataset Structure
Citation
Dataset Details
Dataset Description
CAMERA (CyberAgent Multimodal Evaluation for Ad Text GeneRAtion) is the Japanese ad text generation dataset, which comprises actual data sourced from Japanese search ads and incorporates… See the full description on the dataset page: https://huggingface.co/datasets/cyberagent/camera.CameraBenchProjetson1-top-camera-083126
jetson1-top-camera-083126
Recorded dataset — captured on jetson1 — 5 episodes · 222 frames @ 20 fps (~0 min of demonstration).
Tasks
Instruction
Episodes
place the cucumber on the middle of the cutting board, aligned vertically in the image, with the back end of the cucumber close to the upper edge of the cutting board
2
Recording
Rig
jetson1 (calibration sidecar)
Recorded
2026-08-31
Operator
dorischen
Episode… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/jetson1-top-camera-083126.CameraClone-Dataset
CamCloneMaster: Enabling reference-based camera control for video generation
Paper:https://arxiv.org/abs/2506.03140
Project Page:https://camclonemaster.github.io/
Dataset:https://huggingface.co/datasets/KwaiVGI/CameraClone-Dataset
Training & Inference Code:https://github.com/KwaiVGI/CamCloneMaster
Camera Clone Dataset
1. Dataset Introduction
TL;DR: The Camera Clone Dataset, introduced in CamCloneMaster, is a large-scale synthetic dataset designed… See the full description on the dataset page: https://huggingface.co/datasets/KlingTeam/CameraClone-Dataset.CC12M-Camera
CC12M-Camera
Per-image camera parameter annotations for the CC12M (Conceptual 12M) dataset
(~10.97M images across 2,176 shards), captioned by the Puffin-World model. More captioned datasets are provided in our Puffin-16M website.
The collage above visualizes the camera maps on sample images — each
pair shows the up field (green arrows: the projected gravity-up direction)
and the latitude field (colored contours: angle above/below the horizon).
Format
One .tar per… See the full description on the dataset page: https://huggingface.co/datasets/KangLiao/CC12M-Camera.jetson1-top-camera-test082526
jetson1-top-camera-test082526
Recorded dataset — captured on jetson1 — 1 episodes · 157 frames @ 20 fps (~0 min of demonstration).
Tasks
Instruction
Episodes
place the cucumber on the middle of the cutting board, aligned vertically in the image, with the back end of the cucumber close to the upper edge of the cutting board
1
Recording
Rig
jetson1 (calibration sidecar)
Recorded
2026-08-25
Operator
dorischen
Episode… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/jetson1-top-camera-test082526.PUUM-koa-restoration-camera-trap-dataset
Dataset Card for Koa Associated Biodiversity Camera Trap Dataset
This dataset is aimed at classification of birds visiting planted Acacia koa (koa) trees in the Pu'u Maka'ala Natural Area Reserve (PUUM) on the island of Hawaii (Big Island). The dataset contains full and cropped images collected by camera trap. These images were collected from January 24th to February 25th, 2025.
Dataset Details
This dataset is aimed at classification of birds visiting planted Acacia… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/PUUM-koa-restoration-camera-trap-dataset.new_camera_sock_2_orange_color_onlyflir-camera-objects
Flir Camera Objects
This dataset is part of the Roboflow 100 benchmark, a diverse collection of 100 object detection datasets spanning 7 imagery domains.
Dataset Statistics
Split
Images
Train
9,306
Validation
2,854
Test
1,452
Total
13,612
Classes (4)
bicycle
car
dog
person
Usage
With LibreYOLO
from libreyolo import LIBREYOLO
# Load a model
model = LIBREYOLO(model_path="libreyoloXnano.pt")
# Train on this dataset… See the full description on the dataset page: https://huggingface.co/datasets/LibreYOLO/flir-camera-objects.ImageNet-1K-Camera
ImageNet-1K-Camera
Per-image camera parameter annotations for the full ImageNet-1K dataset
(1,000 training classes + the 50,000-image validation split, ~1.35M images), captioned by the Puffin-World model. More captioned datasets are provided in our Puffin-16M website.
The collage above visualizes the camera maps on sample images — each
pair shows the up field (green arrows: the projected gravity-up direction)
and the latitude field (colored contours: angle above/below the… See the full description on the dataset page: https://huggingface.co/datasets/KangLiao/ImageNet-1K-Camera.G1_Dex3_CameraPackaging_DatasetThis dataset was created using LeRobot.
Due to the inability to precisely describe spatial positions, adjust the scene to closely match the first frame of the dataset after installing the hardware as specified in Part 5 of AVP Teleoperation Documentation.
Data collection is not completed in a single session, and variations between data entries exist. Ensure these variations are accounted for during model training.
Dataset Structure
meta/info.json:
{
"codebase_version":… See the full description on the dataset page: https://huggingface.co/datasets/unitreerobotics/G1_Dex3_CameraPackaging_Dataset.lila_camera_trapsLILA Camera Traps is an aggregate data set of images taken by camera traps, which are devices that automatically (e.g. via motion detection) capture images of wild animals to help ecological research.
This data set is the first time when disparate camera trap data sets have been aggregated into a single training environment with a single taxonomy.
This data set consists of only camera trap image data sets, whereas the broader LILA website also has other data sets related to biology and conservation, intended as a resource for both machine learning (ML) researchers and those that want to harness ML for this topic.HyperSim-Absolute-Camera
HyperSim-Absolute-Camera
Per-frame camera parameter annotations for the HyperSim dataset
(448 scenes; 73,598 valid per-frame annotations),
captioned by the Puffin-World model. More captioned datasets are provided in our Puffin-16M website.
The collage above visualizes the camera maps on sample images — each
pair shows the up field (green arrows: the projected gravity-up direction)
and the latitude field (colored contours: angle above/below the horizon).
Format… See the full description on the dataset page: https://huggingface.co/datasets/KangLiao/HyperSim-Absolute-Camera.InteriorVerse-Camera
InteriorVerse-Camera
Per-image camera parameter annotations for the InteriorVerse dataset
(a large-scale photorealistic synthetic indoor dataset rendered from professionally designed scenes, shipped with per-frame albedo / depth / normal / material maps; 61,231 images), captioned by the Puffin-World model. More captioned datasets are provided in our Puffin-16M website.
The collage above visualizes the camera maps on sample images — each
pair shows the up field (green arrows:… See the full description on the dataset page: https://huggingface.co/datasets/KangLiao/InteriorVerse-Camera.lfv-raw-throw-clean-replay-current-camera
LFV raw recording: Throw_Into_Tank_Clean_Background_Replay_Current_Camera
Public raw robot-teleoperation recording archived by the LFV TUI.
Episodes: 40
Rosbag files: 40
Integrated-health BAD episodes: 1
Camera/action contract: config/lfv_camera_contract.json
Download on a local or cloud training machine:
hf auth login
hf download Onol/lfv-raw-throw-clean-replay-current-camera --repo-type dataset \
--local-dir… See the full description on the dataset page: https://huggingface.co/datasets/Onol/lfv-raw-throw-clean-replay-current-camera.new_camera_orange_color_mergedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/makermods/new_camera_orange_color_merged.plentiful-carla-camera-rigs
Benchmark: Plentiful Carla Camera Rigs
Camera-based perception systems for autonomous driving are typically developed and evaluated using fixed sensor rigs,
while real-world vehicle fleets exhibit substantial variation in camera placement, orientation, field of view, and camera count.
This mismatch introduces a cross-rig domain gap in which only the geometric observation process changes.
To study this effect under controlled conditions, we introduce Plentiful Carla Camera… See the full description on the dataset page: https://huggingface.co/datasets/timb2001/plentiful-carla-camera-rigs.dvrk_bimanual_three_camera_state20This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "franka",
"total_episodes": 74,
"total_frames": 42512,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:74"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/masondx/dvrk_bimanual_three_camera_state20.new_camera_sock_2_orange_color_only_t3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/makermods/new_camera_sock_2_orange_color_only_t3.DL3DV-Absolute-Camera
DL3DV-Absolute-Camera
Per-frame camera parameter annotations for the DL3DV dataset
(6,377 scenes across 7 buckets 1K–7K; 2,161,003 valid per-frame
annotations), captioned by the Puffin-World model. More captioned datasets are provided in our Puffin-16M website.
The collage above visualizes the camera maps on sample images — each
pair shows the up field (green arrows: the projected gravity-up direction)
and the latitude field (colored contours: angle above/below the horizon).… See the full description on the dataset page: https://huggingface.co/datasets/KangLiao/DL3DV-Absolute-Camera.lelab_3_camera_test_20260709_181658This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/makermods/lelab_3_camera_test_20260709_181658.Megalith-10M-Camera
Megalith-10M-Camera
Per-image camera parameter annotations for the Megalith-10M dataset
(the image-bearing build drawthingsai/megalith-10m; ~9.58M Flickr photos across
959 shards), captioned by the Puffin-World model. More captioned datasets are provided in our Puffin-16M website.
The collage above visualizes the camera maps on sample images — each
pair shows the up field (green arrows: the projected gravity-up direction)
and the latitude field (colored contours: angle… See the full description on the dataset page: https://huggingface.co/datasets/KangLiao/Megalith-10M-Camera.thewilds_cameratraps
Dataset Card for The Wilds Camera Trap Data
This dataset contains images and video captured from camera traps deployed at The Wilds safari park in Ohio during Summer 2025. It supports ecological monitoring, animal behavior analysis, and biodiversity studies.
Dataset Details
This dataset was created to support wildlife monitoring research using camera traps. Data were collected using various camera trap models, with each camera recording photos and videos in the… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/thewilds_cameratraps.
