datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ProcVQA-20M-annotations
ProcVQA-20M Annotations
Project Page |
arXiv |
Code |
Model |
Media
This repository contains the text annotations for the ProcVQA-20M dataset. The full image files are hosted separately on ProcVQA-20M-media.
Overview
This dataset is constructed from over 26 embodied datasets, comprising:
20M QA pairs for training
330K original trajectories
50M annotated frames from ~5,000 hours of manipulation data
200+ different tasks
Dataset Structure
The… See the full description on the dataset page: https://huggingface.co/datasets/ce-amtic/ProcVQA-20M-annotations.4k-video-annotations
4K Video Annotations — Shot Segmentation and Camera Motion
This dataset contains 12 frame-accurate shot clips segmented from five short cinematic video sequences. Every clip is paired with a detailed, manually reviewed annotation covering visible content, subject actions, shot scale, camera angle, camera movement, movement direction, stabilization, composition, lighting, color, pacing, transitions, timecodes, and technical properties.
The footage depicts a tense nighttime… See the full description on the dataset page: https://huggingface.co/datasets/LianeMarilin/4k-video-annotations.pokemon-cards-image-and-annotations
