datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
anli
Dataset Card for "anli"
Dataset Summary
The Adversarial Natural Language Inference (ANLI) is a new large-scale NLI benchmark dataset,
The dataset is collected via an iterative, adversarial human-and-model-in-the-loop procedure.
ANLI is much more difficult than its predecessors including SNLI and MNLI.
It contains three rounds. Each round has train/dev/test splits.
Supported Tasks and Leaderboards
More Information Needed
Languages
English… See the full description on the dataset page: https://huggingface.co/datasets/facebook/anli.seamless-interaction
Seamless Interaction Dataset
A large-scale multimodal dataset of 4,000+ hours of human interactions for AI research
🖼️ Blog
🌐 Website
🎮 Demo
📦 GitHub
📄 Paper
Human communication involves a complex interplay of verbal and nonverbal signals, essential for conveying meaning and achieving interpersonal goals.
The Seamless Interaction Dataset is a large-scale collection of over 4,000 hours of face-to-face interaction footage from more than 4,000 participants in… See the full description on the dataset page: https://huggingface.co/datasets/facebook/seamless-interaction.uco3dThis dataset was proposed in UnCommon Objects in 3D.
Code: https://github.com/facebookresearch/uco3d
Project page: https://uco3d.github.io/
belebele
The Belebele Benchmark for Massively Multilingual NLU Evaluation
Belebele is a multiple-choice machine reading comprehension (MRC) dataset spanning 122 language variants. This dataset enables the evaluation of mono- and multi-lingual models in high-, medium-, and low-resource languages. Each question has four multiple-choice answers and is linked to a short passage from the FLORES-200 dataset. The human annotation procedure was carefully curated to create questions that discriminate… See the full description on the dataset page: https://huggingface.co/datasets/facebook/belebele.voxpopuli
Dataset Card for Voxpopuli
Dataset Summary
VoxPopuli is a large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation.
The raw data is collected from 2009-2020 European Parliament event recordings. We acknowledge the European Parliament for creating and sharing these materials.
This implementation contains transcribed speech data for 18 languages.
It also contains 29 hours of transcribed speech data of non-native… See the full description on the dataset page: https://huggingface.co/datasets/facebook/voxpopuli.xnli
Dataset Card for "xnli"
Dataset Summary
XNLI is a subset of a few thousand examples from MNLI which has been translated
into a 14 different languages (some low-ish resource). As with MNLI, the goal is
to predict textual entailment (does sentence A imply/contradict/neither sentence
B) and is a classification task (given two sentences, predict one of three
labels).
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information… See the full description on the dataset page: https://huggingface.co/datasets/facebook/xnli.ego-1k
Ego-1K — A Large-Scale Multiview Video Dataset for Egocentric Vision
Jae Yong Lee, Daniel Scharstein, Akash Bapat, Hao Hu, Andrew Fu, Haoru Zhao, Paul Sammut, Xiang Li,
Stephen Jeapes, Anik Gupta, Lior David, Saketh Madhuvarasu, Jay Girish Joshi, and Jason Wither
CVPR 2026
arXiv:2603.13741
We present Ego-1K, a large-scale collection of time-synchronized egocentric multiview videos designed to advance neural 3D video
synthesis and dynamic scene understanding.… See the full description on the dataset page: https://huggingface.co/datasets/facebook/ego-1k.wiki_dprThis is the wikipedia split used to evaluate the Dense Passage Retrieval (DPR) model.
It contains 21M passages from wikipedia along with their DPR embeddings.
The wikipedia articles were split into multiple, disjoint text blocks of 100 words as passages.multilingual_librispeech
Dataset Card for MultiLingual LibriSpeech
Dataset Summary
This is a streamable version of the Multilingual LibriSpeech (MLS) dataset.
The data archives were restructured from the original ones from OpenSLR to make it easier to stream.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of
8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish.… See the full description on the dataset page: https://huggingface.co/datasets/facebook/multilingual_librispeech.show3d-dataset
SHOW3D: Capturing Scenes of 3D Hands and Objects in the Wild
Patrick Rim, Kevin Harris, Braden Copple, Shangchen Han, Xu Xie, Ivan Shugurov, Sizhe An, He Wen,
Alex Wong, Tomas Hodan, and Kun He
CVPR 2026; https://arxiv.org/abs/2603.28760
News
September 18, 2026: Released synchronized exocentric views and camera calibrations.
SHOW3D is a large-scale multi-view dataset of hand–object interactions captured in the wild.
It is intended to advance research on… See the full description on the dataset page: https://huggingface.co/datasets/facebook/show3d-dataset.CoTracker3_Kubric
Kubric Dataset for CoTracker 3
Overview
This dataset was specifically created for training CoTracker 3, a state-of-the-art point tracking model. The dataset was generated using the Kubric engine.
Dataset Specifications
Size: ~6,000 sequences
Resolution: 512×512 pixels
Sequence Length: 120 frames per sequence
Camera Movement: Carefully rendered with subtle camera motion to simulate realistic scenarios
Format: Generated using Kubric engine
Usage
The… See the full description on the dataset page: https://huggingface.co/datasets/facebook/CoTracker3_Kubric.map-anything
MapAnything Training Metadata Dataset
Dataset Description
This dataset contains pre-computed metadata and covisibility matrices for supporting the MapAnything codebase. This metadata enables easy reproducible training for feed-forward 3D reconstruction tasks.
Please see our Data Processing README for more details.
Citation
If you use this dataset in your research, please cite our paper:
@inproceedings{keetha2026mapanything,
title={{MapAnything}: Universal… See the full description on the dataset page: https://huggingface.co/datasets/facebook/map-anything.hand_tracking_challenge_umetrackCheck out the Multiview Egocentric Hand Tracking Challenge 2024!!
To use this dataset, check out the hand_tracking_toolkit
PE-Video
PE Video Dataset (PVD)
[📃 Tech Report]
[📂 Github]
The PE Video Dataset (PVD) is a large-scale collection of 1 million diverse videos, featuring 120,000+ expertly annotated clips. The dataset was introduced in our paper "Perception Encoder".
Overview
PE Video Dataset (PVD) comprises 1M high quality and diverse videos. Among them, 120K videos are accompanied by automated and human-verified annotations. and all videos are accompanied with video description and keywords.… See the full description on the dataset page: https://huggingface.co/datasets/facebook/PE-Video.flores
Dataset Card for Flores 200
Dataset Summary
⚠️ This repository is no longer being updated ⚠️
A newer version of the FLORES dataset managed by the Open Language Data Initiative
is available at https://huggingface.co/datasets/openlanguagedata/flores_plus.
FLORES is a benchmark dataset for machine translation between English and low-resource languages.
The creation of FLORES-200 doubles the existing language coverage of FLORES-101.
Given the nature of the new… See the full description on the dataset page: https://huggingface.co/datasets/facebook/flores.meta-active-readingmlqa MLQA (MultiLingual Question Answering) is a benchmark dataset for evaluating cross-lingual question answering performance.
MLQA consists of over 5K extractive QA instances (12K in English) in SQuAD format in seven languages - English, Arabic,
German, Spanish, Hindi, Vietnamese and Simplified Chinese. MLQA is highly parallel, with QA instances parallel between
4 different languages on average.kilt_tasks
Dataset Card for KILT
Dataset Summary
KILT has been built from 11 datasets representing 5 types of tasks:
Fact-checking
Entity linking
Slot filling
Open domain QA
Dialog generation
All these datasets have been grounded in a single pre-processed Wikipedia dump, allowing for fairer and more consistent evaluation as well as enabling new task setups such as multitask and transfer learning with minimal effort. KILT also provides tools to analyze and understand the… See the full description on the dataset page: https://huggingface.co/datasets/facebook/kilt_tasks.S-EMBER
S-EMBER: A Large-Scale Benchmark for Streaming Egocentric Memory Retrieval
Episodic-memory video QA benchmark (face-blurred, audio-removed).
License & usage
This dataset is licensed under
CC BY-NC 4.0 and is provided
for non-commercial research use only. Access is gated: you must accept the
non-commercial terms above before downloading.
Contents
sember_mcq.jsonl — multiple-choice evaluation split.
sember_grounding.jsonl — answer-generation and… See the full description on the dataset page: https://huggingface.co/datasets/facebook/S-EMBER.omnilingual-asr-corpus
Meta Omnilingual ASR Corpus
The Omnilingual ASR Corpus is a collection of spontaneous speech recordings and their transcriptions for 348 under-served languages. The corpus was collected as part of Meta FAIR’s Omnilingual ASR project (blog, model, paper) for the purposes of training automatic speech recognition (ASR) and spoken language identification models.
Data schema
{
`language`: "lij_Latn",
`iso_639_3`: "lij",
`iso_15924`: "Latn",
`glottocode`:… See the full description on the dataset page: https://huggingface.co/datasets/facebook/omnilingual-asr-corpus.IntPhys2
IntPhys 2
Dataset |
Hugging Face |
Paper |
Blog
IntPhys 2 is a video benchmark designed to evaluate the intuitive physics understanding of deep learning models. Building on the original IntPhys benchmark, IntPhys 2 focuses on four core principles related to macroscopic objects: Permanence, Immutability, Spatio-Temporal Continuity, and Solidity. These conditions are inspired by research into intuitive physical understanding emerging during early childhood. IntPhys 2… See the full description on the dataset page: https://huggingface.co/datasets/facebook/IntPhys2.md_gender_biasMachine learning models are trained to find patterns in data.
NLP models can inadvertently learn socially undesirable patterns when training on gender biased text.
In this work, we propose a general framework that decomposes gender bias in text along several pragmatic and semantic dimensions:
bias from the gender of the person being spoken about, bias from the gender of the person being spoken to, and bias from the gender of the speaker.
Using this fine-grained framework, we automatically annotate eight large scale datasets with gender information.
In addition, we collect a novel, crowdsourced evaluation benchmark of utterance-level gender rewrites.
Distinguishing between gender bias along multiple dimensions is important, as it enables us to train finer-grained gender bias classifiers.
We show our classifiers prove valuable for a variety of important applications, such as controlling for gender bias in generative models,
detecting gender bias in arbitrary text, and shed light on offensive language in terms of genderedness.wearable-ai
EgoWearBench Dataset (ECCV 2026)
Part of the Wearable AI Workshop at ECCV 2026.
A benchmark of egocentric (first-person, head-mounted wearable camera) videos paired with three complementary video question-answering tasks for evaluating wearable-AI assistants on real-world everyday activity videos.
▶ Baseline code & evaluation scripts: see starter_kit/README.md. The starter kit ships inside this repo, so git clone gives you the code and the data together.
Tasks… See the full description on the dataset page: https://huggingface.co/datasets/facebook/wearable-ai.facebook-bot-mediagelsight-force-estimation
Dataset Details
This dataset contains paired tactile and force data, intended for use in predicting 3-axis normal and shear forces applied to the sensor's elastomer. We used three different indenter shapes to collect force-labeled data: hemisphere, sharp, and flat. To measure force ground truths, we employed the ATI nano17 force/torque sensor. The protocol consisted of applying a random normal load (up to 3N) followed by a shear load, achieved by sliding the probe 2mm on the… See the full description on the dataset page: https://huggingface.co/datasets/facebook/gelsight-force-estimation.natural_reasoningNaturalReasoning is a large-scale dataset for general reasoning tasks. It consists of high-quality challenging reasoning questions backtranslated from pretraining corpora DCLM and FineMath. The questions have been deduplicated and decontaminated from popular reasoning benchmarks including MATH, GPQA, MMLU-Pro, MMLU-STEM. For each question, we extract the reference final answer from the original document from the pretraining corpora if possible. We also provide a model-generated response from… See the full description on the dataset page: https://huggingface.co/datasets/facebook/natural_reasoning.empathetic_dialoguesPyTorch original implementation of Towards Empathetic Open-domain Conversation Models: a New Benchmark and Datasetcyberseceval3-visual-prompt-injection
Dataset Card for CyberSecEval 3 - Visual Prompt Injection Benchmark
Dataset Details
Dataset Description
This dataset provides a multimodal benchmark for visual prompt injection, with text/image inputs. It is part of CyberSecEval 3, the third edition of Meta's flagship suite of security benchmarks for LLMs to measure cybersecurity risks and capabilities across multiple domains.
Language(s): English
License: MIT
Dataset Sources
Repository: Link… See the full description on the dataset page: https://huggingface.co/datasets/facebook/cyberseceval3-visual-prompt-injection.babi_qaThe (20) QA bAbI tasks are a set of proxy tasks that evaluate reading
comprehension via question answering. Our tasks measure understanding
in several ways: whether a system is able to answer questions via chaining facts,
simple induction, deduction and many more. The tasks are designed to be prerequisites
for any system that aims to be capable of conversing with a human.
The aim is to classify these tasks into skill sets,so that researchers
can identify (and then rectify)the failings of their systems.2M-Belebele
2M-Belebele
Highly-Multilingual Speech and American Sign Language Comprehension Dataset
We introduce 2M-Belebele as the first highly multilingual speech and American Sign Language (ASL) comprehension dataset. Our dataset, which is an extension of the existing Belebele only-text dataset, covers 74 spoken languages at the intersection of Belebele and Fleurs, and one sign language (ASL).
The speech dataset is built from aligning Belebele, Flores200 and Fleurs datasets as… See the full description on the dataset page: https://huggingface.co/datasets/facebook/2M-Belebele.
