datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CameraBench
📷 CameraBench: Towards Understanding Camera Motions in Any Video
SfMs and VLMs performance on CameraBench: Generative VLMs (evaluated with VQAScore) trail classical SfM/SLAM in pure geometry, yet they outperform discriminative VLMs that rely on CLIPScore/ITMScore and—even better—capture scene‑aware semantic cues missed by SfM
After simple supervised fine‑tuning (SFT) on ≈1,400 extra annotated clips, our 7B Qwen2.5‑VL doubles its AP, outperforming the current best… See the full description on the dataset page: https://huggingface.co/datasets/syCen/CameraBench.IDLE-OO-Camera-Traps
Dataset Card for IDLE-OO Camera Traps
IDLE-OO Camera Traps is a 5-dataset benchmark of camera trap images from the Labeled Information Library of Alexandria: Biology and Conservation (LILA BC) with a total of 2,586 images for species classification. Each of the 5 benchmarks is balanced to have the same number of images for each species within it (between 310 and 1120 images), representing between 16 and 39 species.
Supported Tasks and Leaderboards
Image… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/IDLE-OO-Camera-Traps.GPIC-Camera
GPIC-Camera
Per-image camera parameter annotations for the GPIC dataset
(train / test / val; train = 8,000 shards, test = 1,000,000 images, val = 200,000 images), captioned by the Puffin-World model. More captioned datasets are provided in our Puffin-16M website.
The collage above visualizes the camera maps on sample images — each
pair shows the up field (green arrows: the projected gravity-up direction)
and the latitude field (colored contours: angle above/below the horizon).… See the full description on the dataset page: https://huggingface.co/datasets/KangLiao/GPIC-Camera.multi-view-bathroom-scene-understanding-camera-relocalization
Multi-View Bathroom Scene Understanding & Camera Relocalization
Generated by datapack-import.ts
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is… See the full description on the dataset page: https://huggingface.co/datasets/physicl/multi-view-bathroom-scene-understanding-camera-relocalization.pseudo-camera-10k
pseudo-camera-10k dataset
Contents
This dataset contains 10k free images from world class photographers. The images have been resized using Lanczos antialiasing, with their smaller edge shifted to 1024px.
The aim of this dataset is a highly variable but high quality and high resolution set of images containing difficult concepts, with about half of the images being numbered group shots and family portraits with the number of subjects labeled.
No images were upsampled in… See the full description on the dataset page: https://huggingface.co/datasets/bghira/pseudo-camera-10k.camera
Dataset Card for CAMERA📷:
Table of Contents:
Dataset Card for Camera
Table of Contents
Dataset Details
Dataset Description
Dataset Sources
Uses
Direct Use
Dataset Information
Data Example
Dataset Structure
Citation
Dataset Details
Dataset Description
CAMERA (CyberAgent Multimodal Evaluation for Ad Text GeneRAtion) is the Japanese ad text generation dataset, which comprises actual data sourced from Japanese search ads and incorporates… See the full description on the dataset page: https://huggingface.co/datasets/cyberagent/camera.CC12M-Camera
CC12M-Camera
Per-image camera parameter annotations for the CC12M (Conceptual 12M) dataset
(~10.97M images across 2,176 shards), captioned by the Puffin-World model. More captioned datasets are provided in our Puffin-16M website.
The collage above visualizes the camera maps on sample images — each
pair shows the up field (green arrows: the projected gravity-up direction)
and the latitude field (colored contours: angle above/below the horizon).
Format
One .tar per… See the full description on the dataset page: https://huggingface.co/datasets/KangLiao/CC12M-Camera.PUUM-koa-restoration-camera-trap-dataset
Dataset Card for Koa Associated Biodiversity Camera Trap Dataset
This dataset is aimed at classification of birds visiting planted Acacia koa (koa) trees in the Pu'u Maka'ala Natural Area Reserve (PUUM) on the island of Hawaii (Big Island). The dataset contains full and cropped images collected by camera trap. These images were collected from January 24th to February 25th, 2025.
Dataset Details
This dataset is aimed at classification of birds visiting planted Acacia… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/PUUM-koa-restoration-camera-trap-dataset.dvrk_bimanual_three_camera_state20This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "franka",
"total_episodes": 74,
"total_frames": 42512,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:74"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/masondx/dvrk_bimanual_three_camera_state20.Megalith-10M-Camera
Megalith-10M-Camera
Per-image camera parameter annotations for the Megalith-10M dataset
(the image-bearing build drawthingsai/megalith-10m; ~9.58M Flickr photos across
959 shards), captioned by the Puffin-World model. More captioned datasets are provided in our Puffin-16M website.
The collage above visualizes the camera maps on sample images — each
pair shows the up field (green arrows: the projected gravity-up direction)
and the latitude field (colored contours: angle… See the full description on the dataset page: https://huggingface.co/datasets/KangLiao/Megalith-10M-Camera.thewilds_cameratraps
Dataset Card for The Wilds Camera Trap Data
This dataset contains images and video captured from camera traps deployed at The Wilds safari park in Ohio during Summer 2025. It supports ecological monitoring, animal behavior analysis, and biodiversity studies.
Dataset Details
This dataset was created to support wildlife monitoring research using camera traps. Data were collected using various camera trap models, with each camera recording photos and videos in the… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/thewilds_cameratraps.pseudo-camera-10k-structured-json
pseudo-camera-10k, structured JSON captions
The 9,997 training images from bghira/pseudo-camera-10k, recaptioned into the structured JSON caption schema that Ideogram 4 consumes. The images are unchanged: free photographs from world class photographers, Lanczos-resized so the shorter edge is 1024px, nothing upsampled.
The original dataset carries short CogVLM prose captions. This one replaces them with one JSON object per image describing the scene at three levels: an overall… See the full description on the dataset page: https://huggingface.co/datasets/terminusresearch/pseudo-camera-10k-structured-json.CameraBench-Pro
CameraBench-Pro
This dataset contains the testing split for the CameraBench-Pro evaluation.
camera-grandstaffcamera_settingsThis is the dataset for Generative Photography: Scene-Consistent Camera Control for Realistic Text-to-Image Synthesis, a framework for achieving realistic scene-consistent text-to-image generation with camera awareness.
Code: https://github.com/pandayuanyu/generative-photography
For more details, please check:
Project page
Multi-Camera-Multi-Vehicle-Tracking-System
Multi-Camera Multi-Vehicle Tracking System (UAV Dataset)
This dataset contains synchronized multi-UAV video footage and comprehensive tracking annotations, accompanying the paper A Topology-Aware Spatiotemporal Handover Framework for Continuous Multi-UAV Tracking.
It is designed for evaluating real-time multi-camera multi-vehicle tracking (MCMT) systems, focusing on solving trajectory fragmentation and maintaining global identity persistence across isolated UAV fields of view.… See the full description on the dataset page: https://huggingface.co/datasets/jye9/Multi-Camera-Multi-Vehicle-Tracking-System.COCO-Camera
COCO-Camera
Per-image camera parameter annotations for the COCO dataset
(detection-datasets/coco; train + val, ~122K images across 42 shards), captioned by the Puffin-World model. More captioned datasets are provided in our Puffin-16M website.
The collage above visualizes the camera maps on sample images — each
pair shows the up field (green arrows: the projected gravity-up direction)
and the latitude field (colored contours: angle above/below the horizon).
Format… See the full description on the dataset page: https://huggingface.co/datasets/KangLiao/COCO-Camera.ferma-camera-yolo
Ferma Camera YOLO
Dataset Summary
Цель датасета — обучить YOLO распознавать людей, собак, коров и хищников на камерах в ферме.
Data Sources
Данные собраны из нескольких источников:
часть изображений взята из OIDv6
часть изображений взята из датасета 8 calves
часть изображений взята с YouTube и размечена вручную
Data Structure
images/ — изображения .jpg
labels/ — разметка YOLO .txt (same stem)
Annotation Format
Формат YOLO: class_id… See the full description on the dataset page: https://huggingface.co/datasets/I77/ferma-camera-yolo.so101_camera_datasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"image": {
"dtype": "image",
"shape": [
480,
640,
3
],
"names": [
"height",
"width",
"channels"
],
"info": {
"is_depth_map": false
}… See the full description on the dataset page: https://huggingface.co/datasets/codenmood/so101_camera_dataset.vggt-camera-movecamera-movement
Rapidata Camera Movement Benchmark
Built by Rapidata.
This dataset contains 281,738 human responses, collected with the
Rapidata Python SDK, comparing how well 14 image-to-video models and world models
execute a described camera movement from a single still image. Each row is a head-to-head comparison between
two models' clips generated from the same still and the same instruction, judged by human annotators who
watched a reference animation of the requested movement.
The task… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/camera-movement.fix_cameraThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "dex-hand",
"total_episodes": 7,
"total_frames": 1400,
"total_tasks":7,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:7"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/zcyqyq/fix_camera.close_fixed_cameraweb-camera-people-behavior
Web Camera People Behavior Dataset for computer vision tasks
Dataset includes 2,300+ individuals, contributing to a total of 53,800+ videos and 9,300+ images captured via webcams. It is designed to study social interactions and behaviors in various remote meetings, including video calls, video conferencing, and online meetings.
By leveraging this dataset, developers and researchers can enhance their understanding of human behavior in digital communication settings, contributing… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/web-camera-people-behavior.so101_camera_dataset_0016This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"image": {
"dtype": "image",
"shape": [
480,
640,
3
],
"names": [
"height",
"width",
"channel"
]
},
"wrist_image": {
"dtype": "image"… See the full description on the dataset page: https://huggingface.co/datasets/codenmood/so101_camera_dataset_0016.dvrk_bimanual_three_cameraThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "franka",
"total_episodes": 74,
"total_frames": 42512,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:74"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/masondx/dvrk_bimanual_three_camera.close_2_cameras_v2run_hybrid_Camera_Control_UCPEflir-camera-objects
Dataset Card for flir-camera-objects
** The original COCO dataset is stored at dataset.tar.gz**
Dataset Summary
flir-camera-objects
Supported Tasks and Leaderboards
object-detection: The dataset can be used to train a model for Object Detection.
Languages
English
Dataset Structure
Data Instances
A data point comprises an image and its object annotations.
{
'image_id': 15,
'image': <PIL.JpegImagePlugin.JpegImageFile image… See the full description on the dataset page: https://huggingface.co/datasets/Francesco/flir-camera-objects.camera-calib-and-scene-alignment-data
