datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
anli
Dataset Card for "anli"
Dataset Summary
The Adversarial Natural Language Inference (ANLI) is a new large-scale NLI benchmark dataset,
The dataset is collected via an iterative, adversarial human-and-model-in-the-loop procedure.
ANLI is much more difficult than its predecessors including SNLI and MNLI.
It contains three rounds. Each round has train/dev/test splits.
Supported Tasks and Leaderboards
More Information Needed
Languages
English… See the full description on the dataset page: https://huggingface.co/datasets/facebook/anli.seamless-interaction
Seamless Interaction Dataset
A large-scale multimodal dataset of 4,000+ hours of human interactions for AI research
🖼️ Blog
🌐 Website
🎮 Demo
📦 GitHub
📄 Paper
Human communication involves a complex interplay of verbal and nonverbal signals, essential for conveying meaning and achieving interpersonal goals.
The Seamless Interaction Dataset is a large-scale collection of over 4,000 hours of face-to-face interaction footage from more than 4,000 participants in… See the full description on the dataset page: https://huggingface.co/datasets/facebook/seamless-interaction.uco3dThis dataset was proposed in UnCommon Objects in 3D.
Code: https://github.com/facebookresearch/uco3d
Project page: https://uco3d.github.io/
belebele
The Belebele Benchmark for Massively Multilingual NLU Evaluation
Belebele is a multiple-choice machine reading comprehension (MRC) dataset spanning 122 language variants. This dataset enables the evaluation of mono- and multi-lingual models in high-, medium-, and low-resource languages. Each question has four multiple-choice answers and is linked to a short passage from the FLORES-200 dataset. The human annotation procedure was carefully curated to create questions that discriminate… See the full description on the dataset page: https://huggingface.co/datasets/facebook/belebele.voxpopuli
Dataset Card for Voxpopuli
Dataset Summary
VoxPopuli is a large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation.
The raw data is collected from 2009-2020 European Parliament event recordings. We acknowledge the European Parliament for creating and sharing these materials.
This implementation contains transcribed speech data for 18 languages.
It also contains 29 hours of transcribed speech data of non-native… See the full description on the dataset page: https://huggingface.co/datasets/facebook/voxpopuli.xnli
Dataset Card for "xnli"
Dataset Summary
XNLI is a subset of a few thousand examples from MNLI which has been translated
into a 14 different languages (some low-ish resource). As with MNLI, the goal is
to predict textual entailment (does sentence A imply/contradict/neither sentence
B) and is a classification task (given two sentences, predict one of three
labels).
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information… See the full description on the dataset page: https://huggingface.co/datasets/facebook/xnli.ego-1k
Ego-1K — A Large-Scale Multiview Video Dataset for Egocentric Vision
Jae Yong Lee, Daniel Scharstein, Akash Bapat, Hao Hu, Andrew Fu, Haoru Zhao, Paul Sammut, Xiang Li,
Stephen Jeapes, Anik Gupta, Lior David, Saketh Madhuvarasu, Jay Girish Joshi, and Jason Wither
CVPR 2026
arXiv:2603.13741
We present Ego-1K, a large-scale collection of time-synchronized egocentric multiview videos designed to advance neural 3D video
synthesis and dynamic scene understanding.… See the full description on the dataset page: https://huggingface.co/datasets/facebook/ego-1k.wiki_dprThis is the wikipedia split used to evaluate the Dense Passage Retrieval (DPR) model.
It contains 21M passages from wikipedia along with their DPR embeddings.
The wikipedia articles were split into multiple, disjoint text blocks of 100 words as passages.multilingual_librispeech
Dataset Card for MultiLingual LibriSpeech
Dataset Summary
This is a streamable version of the Multilingual LibriSpeech (MLS) dataset.
The data archives were restructured from the original ones from OpenSLR to make it easier to stream.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of
8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish.… See the full description on the dataset page: https://huggingface.co/datasets/facebook/multilingual_librispeech.show3d-dataset
SHOW3D: Capturing Scenes of 3D Hands and Objects in the Wild
Patrick Rim, Kevin Harris, Braden Copple, Shangchen Han, Xu Xie, Ivan Shugurov, Sizhe An, He Wen,
Alex Wong, Tomas Hodan, and Kun He
CVPR 2026; https://arxiv.org/abs/2603.28760
News
September 18, 2026: Released synchronized exocentric views and camera calibrations.
SHOW3D is a large-scale multi-view dataset of hand–object interactions captured in the wild.
It is intended to advance research on… See the full description on the dataset page: https://huggingface.co/datasets/facebook/show3d-dataset.eigen-face-dataset-256
EigenFace-256 Dataset Construction
Training a robust face embedding model requires a diverse dataset with multiple images per identity—capturing variations in angles, lighting, expressions, and age. Using real-world data for this purpose poses several challenges:
Privacy and Ethics: Real-world images can lead to legal and ethical complications.
Bias and Imbalance: Datasets based on real images may lack diversity.
Data Labeling Complexity: Annotating large datasets is… See the full description on the dataset page: https://huggingface.co/datasets/tropos-labs/eigen-face-dataset-256.DH-FaceVid-1KXVLA-Soft-Fold
Cloth-Folding Dataset for X-VLA Paper
This dataset contains 1,500 episodes of cloth folding, collected using Agilex's robotic arm. It was used in the X-VLA paper for cloth-folding tasks, showcasing a near-perfect success rate in folding accuracy.
Dataset Overview
Total Episodes: ~1,500
Task: Automated cloth folding
Robot: Agilex Aloha
Performance: Near 100% success rate in completing the folding task
Hardware setup
We observed that the camera setup of the… See the full description on the dataset page: https://huggingface.co/datasets/Facebear/XVLA-Soft-Fold.CoTracker3_Kubric
Kubric Dataset for CoTracker 3
Overview
This dataset was specifically created for training CoTracker 3, a state-of-the-art point tracking model. The dataset was generated using the Kubric engine.
Dataset Specifications
Size: ~6,000 sequences
Resolution: 512×512 pixels
Sequence Length: 120 frames per sequence
Camera Movement: Carefully rendered with subtle camera motion to simulate realistic scenarios
Format: Generated using Kubric engine
Usage
The… See the full description on the dataset page: https://huggingface.co/datasets/facebook/CoTracker3_Kubric.map-anything
MapAnything Training Metadata Dataset
Dataset Description
This dataset contains pre-computed metadata and covisibility matrices for supporting the MapAnything codebase. This metadata enables easy reproducible training for feed-forward 3D reconstruction tasks.
Please see our Data Processing README for more details.
Citation
If you use this dataset in your research, please cite our paper:
@inproceedings{keetha2026mapanything,
title={{MapAnything}: Universal… See the full description on the dataset page: https://huggingface.co/datasets/facebook/map-anything.hand_tracking_challenge_umetrackCheck out the Multiview Egocentric Hand Tracking Challenge 2024!!
To use this dataset, check out the hand_tracking_toolkit
FaceDetection
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/datnguyentien204/FaceDetection.PE-Video
PE Video Dataset (PVD)
[📃 Tech Report]
[📂 Github]
The PE Video Dataset (PVD) is a large-scale collection of 1 million diverse videos, featuring 120,000+ expertly annotated clips. The dataset was introduced in our paper "Perception Encoder".
Overview
PE Video Dataset (PVD) comprises 1M high quality and diverse videos. Among them, 120K videos are accompanied by automated and human-verified annotations. and all videos are accompanied with video description and keywords.… See the full description on the dataset page: https://huggingface.co/datasets/facebook/PE-Video.agent-task-facet-terminal-6k
Apptainer pool for hamishivi/agent-task-facet-terminal-6k
This repository hosts tmax-compatible SIF images and a unified download manifest. Training data and task archives are in hamishivi/agent-task-facet-terminal-6k. The manifest includes earlier images hosted under hamishivi and new images hosted under TMaxxx; the downloader selects the correct repository and immutable commit for each image.
Apptainer images
The pool currently contains 6,020 / 6,020 verified… See the full description on the dataset page: https://huggingface.co/datasets/TMaxxx/agent-task-facet-terminal-6k.FACETflores
Dataset Card for Flores 200
Dataset Summary
⚠️ This repository is no longer being updated ⚠️
A newer version of the FLORES dataset managed by the Open Language Data Initiative
is available at https://huggingface.co/datasets/openlanguagedata/flores_plus.
FLORES is a benchmark dataset for machine translation between English and low-resource languages.
The creation of FLORES-200 doubles the existing language coverage of FLORES-101.
Given the nature of the new… See the full description on the dataset page: https://huggingface.co/datasets/facebook/flores.facesyntheticsspigacaptioned
Dataset Card for "face_synthetics_spiga_captioned"
This is a copy of the Microsoft FaceSynthetics dataset with SPIGA-calculated landmark annotations, and additional BLIP-generated captions.
For a copy of the original FaceSynthetics dataset with no extra annotations, please refer to pcuenq/face_synthetics.
Here is the code for parsing the dataset and generating the BLIP captions:
from transformers import pipeline
dataset_name = "pcuenq/face_synthetics_spiga"
faces =… See the full description on the dataset page: https://huggingface.co/datasets/multimodalart/facesyntheticsspigacaptioned.meta-active-readingmlqa MLQA (MultiLingual Question Answering) is a benchmark dataset for evaluating cross-lingual question answering performance.
MLQA consists of over 5K extractive QA instances (12K in English) in SQuAD format in seven languages - English, Arabic,
German, Spanish, Hindi, Vietnamese and Simplified Chinese. MLQA is highly parallel, with QA instances parallel between
4 different languages on average.kilt_tasks
Dataset Card for KILT
Dataset Summary
KILT has been built from 11 datasets representing 5 types of tasks:
Fact-checking
Entity linking
Slot filling
Open domain QA
Dialog generation
All these datasets have been grounded in a single pre-processed Wikipedia dump, allowing for fairer and more consistent evaluation as well as enabling new task setups such as multitask and transfer learning with minimal effort. KILT also provides tools to analyze and understand the… See the full description on the dataset page: https://huggingface.co/datasets/facebook/kilt_tasks.S-EMBER
S-EMBER: A Large-Scale Benchmark for Streaming Egocentric Memory Retrieval
Episodic-memory video QA benchmark (face-blurred, audio-removed).
License & usage
This dataset is licensed under
CC BY-NC 4.0 and is provided
for non-commercial research use only. Access is gated: you must accept the
non-commercial terms above before downloading.
Contents
sember_mcq.jsonl — multiple-choice evaluation split.
sember_grounding.jsonl — answer-generation and… See the full description on the dataset page: https://huggingface.co/datasets/facebook/S-EMBER.omnilingual-asr-corpus
Meta Omnilingual ASR Corpus
The Omnilingual ASR Corpus is a collection of spontaneous speech recordings and their transcriptions for 348 under-served languages. The corpus was collected as part of Meta FAIR’s Omnilingual ASR project (blog, model, paper) for the purposes of training automatic speech recognition (ASR) and spoken language identification models.
Data schema
{
`language`: "lij_Latn",
`iso_639_3`: "lij",
`iso_15924`: "Latn",
`glottocode`:… See the full description on the dataset page: https://huggingface.co/datasets/facebook/omnilingual-asr-corpus.CelebA-faces-with-attributesIntPhys2
IntPhys 2
Dataset |
Hugging Face |
Paper |
Blog
IntPhys 2 is a video benchmark designed to evaluate the intuitive physics understanding of deep learning models. Building on the original IntPhys benchmark, IntPhys 2 focuses on four core principles related to macroscopic objects: Permanence, Immutability, Spatio-Temporal Continuity, and Solidity. These conditions are inspired by research into intuitive physical understanding emerging during early childhood. IntPhys 2… See the full description on the dataset page: https://huggingface.co/datasets/facebook/IntPhys2.wider_faceWIDER FACE dataset is a face detection benchmark dataset, of which images are
selected from the publicly available WIDER dataset. We choose 32,203 images and
label 393,703 faces with a high degree of variability in scale, pose and
occlusion as depicted in the sample images. WIDER FACE dataset is organized
based on 61 event classes. For each event class, we randomly select 40%/10%/50%
data as training, validation and testing sets. We adopt the same evaluation
metric employed in the PASCAL VOC dataset. Similar to MALF and Caltech datasets,
we do not release bounding box ground truth for the test images. Users are
required to submit final prediction files, which we shall proceed to evaluate.
