datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
seamless-interaction
Seamless Interaction Dataset
A large-scale multimodal dataset of 4,000+ hours of human interactions for AI research
🖼️ Blog
🌐 Website
🎮 Demo
📦 GitHub
📄 Paper
Human communication involves a complex interplay of verbal and nonverbal signals, essential for conveying meaning and achieving interpersonal goals.
The Seamless Interaction Dataset is a large-scale collection of over 4,000 hours of face-to-face interaction footage from more than 4,000 participants in… See the full description on the dataset page: https://huggingface.co/datasets/facebook/seamless-interaction.show3d-dataset
SHOW3D: Capturing Scenes of 3D Hands and Objects in the Wild
Patrick Rim, Kevin Harris, Braden Copple, Shangchen Han, Xu Xie, Ivan Shugurov, Sizhe An, He Wen,
Alex Wong, Tomas Hodan, and Kun He
CVPR 2026; https://arxiv.org/abs/2603.28760
News
September 18, 2026: Released synchronized exocentric views and camera calibrations.
SHOW3D is a large-scale multi-view dataset of hand–object interactions captured in the wild.
It is intended to advance research on… See the full description on the dataset page: https://huggingface.co/datasets/facebook/show3d-dataset.DH-FaceVid-1KS-EMBER
S-EMBER: A Large-Scale Benchmark for Streaming Egocentric Memory Retrieval
Episodic-memory video QA benchmark (face-blurred, audio-removed).
License & usage
This dataset is licensed under
CC BY-NC 4.0 and is provided
for non-commercial research use only. Access is gated: you must accept the
non-commercial terms above before downloading.
Contents
sember_mcq.jsonl — multiple-choice evaluation split.
sember_grounding.jsonl — answer-generation and… See the full description on the dataset page: https://huggingface.co/datasets/facebook/S-EMBER.IntPhys2
IntPhys 2
Dataset |
Hugging Face |
Paper |
Blog
IntPhys 2 is a video benchmark designed to evaluate the intuitive physics understanding of deep learning models. Building on the original IntPhys benchmark, IntPhys 2 focuses on four core principles related to macroscopic objects: Permanence, Immutability, Spatio-Temporal Continuity, and Solidity. These conditions are inspired by research into intuitive physical understanding emerging during early childhood. IntPhys 2… See the full description on the dataset page: https://huggingface.co/datasets/facebook/IntPhys2.wearable-ai
EgoWearBench Dataset (ECCV 2026)
Part of the Wearable AI Workshop at ECCV 2026.
A benchmark of egocentric (first-person, head-mounted wearable camera) videos paired with three complementary video question-answering tasks for evaluating wearable-AI assistants on real-world everyday activity videos.
▶ Baseline code & evaluation scripts: see starter_kit/README.md. The starter kit ships inside this repo, so git clone gives you the code and the data together.
Tasks… See the full description on the dataset page: https://huggingface.co/datasets/facebook/wearable-ai.FaceForensicsC23FaceForensics++ Dataset Overview
The dataset (downloaded using the original scripts) contains 7,010 files in total:
7,000 MP4 videos — 6,000 deepfakes and 1,000 real videos
Folder Structure:
DeepFakeDetection – 1,000 deepfake videos
Deepfakes – 1,000 deepfake videos
Face2Face – 1,000 deepfake videos
FaceShifter – 1,000 deepfake videos
FaceSwap – 1,000 deepfake videos
NeuralTextures – 1,000 deepfake videos
Real – 1,000 real videos
2M-Flores-ASL
2M-Flores
As part of the 2M-Belebele project, we have produced video recodings of ASL signing for all the dev and devtest
sentences in the original flores200 dataset.
To obtain ASL sign recordings, we provide translators of ASL and native signers with the English text version of the sentences to be recorded.
The interpreters are then asked to translate these sentences into ASL, create glosses for all sentences, and record their interpretations into ASL one sentence at a time.
The… See the full description on the dataset page: https://huggingface.co/datasets/facebook/2M-Flores-ASL.actionbench
🎬 ActionBench: Paired Video-3D Synthetic Benchmark
📖 Overview
ActionBench is a benchmark dataset of 128 paired video ↔ animated point-cloud samples for evaluating animated 3D mesh generation from video.
The dataset consists of synthetic scenes of animated objects from ObjaverseXL, rendered using Blender 3.5.1.
Each sample contains:
Video: 16 RGBA frames with alpha mask
Camera (camera.json): Camera parameters using Blender convention (X_cam = X @ R^T + T, camera looks… See the full description on the dataset page: https://huggingface.co/datasets/facebook/actionbench.face-anti-spoofing-dataset
Face Antispoofing dataset for liveness detection
Anti-Spoofing dataset: live, replay, cut, print, 3D masks - large-scale face anti spoofing
This dataset delivers a single, end-to-end resource for training and benchmarking facial liveness-detection systems. By aggregating live sessions and eleven realistic presentation-attack classes into one collection, it accelerates development toward iBeta Level 1/2 compliance and strengthens model robustness against the full spectrum of spoofing… See the full description on the dataset page: https://huggingface.co/datasets/AxonData/face-anti-spoofing-dataset.action100m-preview
Action100M: A Large-scale Video Action Dataset
Paper | GitHub
Action100M is a large-scale dataset constructed from 1.2M Internet instructional videos (14.6 years of duration), yielding ~100 million temporally localized segments with open-vocabulary action supervision and rich captions. It serves as a foundation for scalable research in video understanding and world modeling.
Load Action100M Annotations
Our data can be loaded from the 🤗 huggingface repo at… See the full description on the dataset page: https://huggingface.co/datasets/facebook/action100m-preview.PLM-VideoBench
Dataset Summary
PLM-VideoBench is a collection of human-annotated resources for evaluating Vision Language models, focused on detailed video understanding.
[📃 Tech Report]
[📂 Github]
Supported Tasks
PLM-VideoBench includes evaluation data for the following tasks:
FGQA
In this task, a model must answer a multiple-choice question (MCQ) that probes fine-grained activity understanding. Given a question and multiple options that differ in a… See the full description on the dataset page: https://huggingface.co/datasets/facebook/PLM-VideoBench.face-segmentation-image-dataset
Image Dataset of Face Segmentation for recognition tasks
Dataset comprises 87,800+ images annotated with 100+ landmarks, providing a comprehensive foundation for research in face recognition, segmentation tasks, and object recognition. It is designed to support the development of learning models, recognition algorithms, and segmentation techniques.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in facial recognition, face… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/face-segmentation-image-dataset.jepa-wms
🌍 JEPA-WMs Datasets
Robotics trajectories for world model training 🤖
Meta AI Research, FAIR
This 🤗 HuggingFace repository hosts datasets for training JEPA-WM world models.
👉 See the main repository for training code and pretrained models.
👁️ Preview Images: To view example images in the Dataset Viewer above, select a dataset configuration (e.g., metaworld, pusht) and click "Run query".
📦 Downloading Data
Use the download script… See the full description on the dataset page: https://huggingface.co/datasets/facebook/jepa-wms.face_id_v2_test_codeface-anti-spoofing
Face Antispoofing dataset for recognition systems
The dataset consists of 98,000 videos and selfies from 170 countries, providing a foundation for developing robust security systems and facial recognition algorithms.
While the dataset itself doesn't contain spoofing attacks, it's a valuable resource for testing liveness detection system, allowing researchers to simulate attacks and evaluate how effectively their systems can distinguish between real faces and various forms of… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/face-anti-spoofing.EgoAVU_data
[CVPR2026 HIGHLIGHT] EgoAVU, [ICASSP2026 Oral] Exploring Audio Hallucination in Egocentric Video Understanding
Official Implementation of EgoAVU: Egocentric Audio-Visual Understanding and Exploring Audio Hallucination in Egocentric Video Understanding
See our github for the code and setup instructions.
Check out our homepage, paper (CVPR) and paper (ICASSP) for more information.
We introduce EgoAVU, a scalable and automated data engine to enable egocentric audio–visual… See the full description on the dataset page: https://huggingface.co/datasets/facebook/EgoAVU_data.web-camera-face-liveness-detection
Web Camera Face Liveness Detection
The dataset consists of videos featuring individuals wearing various types of masks. Videos are recorded under different lighting conditions and with different attributes (glasses, masks, hats, hoods, wigs, and mustaches for men).
The dataset is created on the basis of iBeta Level 1 Dataset
In the dataset, there are 7 types of videos filmed on a web camera:
Silicone Mask - demonstration of a silicone mask attack (silicone)
2D mask with… See the full description on the dataset page: https://huggingface.co/datasets/UniqueData/web-camera-face-liveness-detection.face_segmentationAn example of a dataset that we've collected for a photo edit App.
The dataset includes 20 selfies of people (man and women)
in segmentation masks and their visualisations.face-recognition-image-dataset
Image Dataset of face images for compuer vision tasks
Dataset comprises 500,600+ images of individuals representing various races, genders, and ages, with each person having a single face image. It is designed for facial recognition and face detection research, supporting the development of advanced recognition systems.
By leveraging this dataset, researchers and developers can enhance deep learning models, improve face verification and face identification techniques, and refine… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/face-recognition-image-dataset.task03_turn_to_face_beakerThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "mobileai_robot",
"total_episodes": 117,
"total_frames": 31189,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:117"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/kiroaiseoul/task03_turn_to_face_beaker.DeepFake_Extracted_Face_Imagesface-re-identification-image-dataset
Dataset of face images with different angles and head positions
Dataset contains 23,110 individuals, each contributing 28 images featuring various angles and head positions, diverse backgrounds, and attributes, along with 1 ID photo. In total, the dataset comprises over 670,000 images in formats such as JPG and PNG. It is designed to advance face recognition and facial recognition research, focusing on person re-identification and recognition systems.
By utilizing this dataset… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/face-re-identification-image-dataset.on-device-face-liveness-detection
Mobile Face Liveness Detection
The dataset consists of videos featuring individuals wearing various types of masks. Videos are recorded under different lighting conditions and with different attributes (glasses, masks, hats, hoods, wigs, and mustaches for men).
The dataset is created on the basis of iBeta Level 1 Dataset
In the dataset, there are 4 types of videos filmed on mobile devices:
2D mask with holes for eyes - demonstration of an attack with a paper/cardboard… See the full description on the dataset page: https://huggingface.co/datasets/UniqueData/on-device-face-liveness-detection.task03_turn_to_face_beaker_gist
task03_turn_to_face_beaker_gist
GIST AI Lab이 수집한 11스테이지 실험실 프로토콜 데이터 중 stage 3 입니다.
kiroaiseoul 네이밍(task03_turn_to_face_beaker)에 맞춰 재배포한 사본이고, 원본 배포명은 task3-scan-and-move-to-beaker 입니다.
Task: "Turn in place to face the beaker"
Episodes: 2,228 / Frames: 334,996
FPS: 30 · robot_type: mobileai_robot · codebase_version: v3.0
Cameras: cam_high, cam_left_wrist, cam_right_wrist (480x640x3, h264)
action / observation.state: 16D (Mobile AI — base 2 + left 7 + right 7)… See the full description on the dataset page: https://huggingface.co/datasets/kiroaiseoul/task03_turn_to_face_beaker_gist.face_masksDataset includes 250 000 images, 4 types of mask worn on 28 000 unique faces.
All images were collected using the Toloka.ai crowdsourcing service and
validated by TrainingData.propush_cube_to_face_reward_cropped_resizedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "so100",
"total_episodes": 20,
"total_frames": 8447,
"total_tasks": 1,
"total_videos": 40,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aractingi/push_cube_to_face_reward_cropped_resized.face-anonymization-public-evaluation-v58
Origin Data Lab — Face Anonymization Public Evaluation V58
Privacy processing for real-world video datasets
Automated face anonymization combined with targeted Human QA for autonomous driving, robotics, computer vision, urban mobility, and AI data teams.
This repository presents a public engineering evaluation of Origin Data Lab's face anonymization pipeline in dense, low-light urban traffic conditions.
Evaluate Your Own Video
Need privacy processing for traffic… See the full description on the dataset page: https://huggingface.co/datasets/origindatalab/face-anonymization-public-evaluation-v58.push_cube_to_face_rewardThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "so100",
"total_episodes": 20,
"total_frames": 8447,
"total_tasks": 1,
"total_videos": 40,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aractingi/push_cube_to_face_reward.draw-smiley-face-v7
