CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01facebook /anli Dataset Card for "anli" Dataset Summary The Adversarial Natural Language Inference (ANLI) is a new large-scale NLI benchmark dataset, The dataset is collected via an iterative, adversarial human-and-model-in-the-loop procedure. ANLI is much more difficult than its predecessors including SNLI and MNLI. It contains three rounds. Each round has train/dev/test splits. Supported Tasks and Leaderboards More Information Needed Languages English… See the full description on the dataset page: https://huggingface.co/datasets/facebook/anli.texttext-classification100K<n<1M53 likes179k downloads3y agoHugging Face02facebook /seamless-interaction Seamless Interaction Dataset A large-scale multimodal dataset of 4,000+ hours of human interactions for AI research 🖼️ Blog 🌐 Website 🎮 Demo 📦 GitHub 📄 Paper Human communication involves a complex interplay of verbal and nonverbal signals, essential for conveying meaning and achieving interpersonal goals. The Seamless Interaction Dataset is a large-scale collection of over 4,000 hours of face-to-face interaction footage from more than 4,000 participants in… See the full description on the dataset page: https://huggingface.co/datasets/facebook/seamless-interaction.audio197 likes100k downloads1y agoHugging Face03facebook /uco3dThis dataset was proposed in UnCommon Objects in 3D. Code: https://github.com/facebookresearch/uco3d Project page: https://uco3d.github.io/ image-to-3d10 likes98k downloads2d agoHugging Face04facebook /belebele The Belebele Benchmark for Massively Multilingual NLU Evaluation Belebele is a multiple-choice machine reading comprehension (MRC) dataset spanning 122 language variants. This dataset enables the evaluation of mono- and multi-lingual models in high-, medium-, and low-resource languages. Each question has four multiple-choice answers and is linked to a short passage from the FLORES-200 dataset. The human annotation procedure was carefully curated to create questions that discriminate… See the full description on the dataset page: https://huggingface.co/datasets/facebook/belebele.textquestion-answering100K<n<1M133 likes84k downloads2y agoHugging Face05facebook /voxpopuli Dataset Card for Voxpopuli Dataset Summary VoxPopuli is a large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation. The raw data is collected from 2009-2020 European Parliament event recordings. We acknowledge the European Parliament for creating and sharing these materials. This implementation contains transcribed speech data for 18 languages. It also contains 29 hours of transcribed speech data of non-native… See the full description on the dataset page: https://huggingface.co/datasets/facebook/voxpopuli.audioautomatic-speech-recognition1M<n<10M164 likes77k downloads8mo agoHugging Face06facebook /xnli Dataset Card for "xnli" Dataset Summary XNLI is a subset of a few thousand examples from MNLI which has been translated into a 14 different languages (some low-ish resource). As with MNLI, the goal is to predict textual entailment (does sentence A imply/contradict/neither sentence B) and is a classification task (given two sentences, predict one of three labels). Supported Tasks and Leaderboards More Information Needed Languages More Information… See the full description on the dataset page: https://huggingface.co/datasets/facebook/xnli.text1M<n<10M73 likes68k downloads3y agoHugging Face07facebook /ego-1k Ego-1K — A Large-Scale Multiview Video Dataset for Egocentric Vision Jae Yong Lee, Daniel Scharstein, Akash Bapat, Hao Hu, Andrew Fu, Haoru Zhao, Paul Sammut, Xiang Li, Stephen Jeapes, Anik Gupta, Lior David, Saketh Madhuvarasu, Jay Girish Joshi, and Jason Wither CVPR 2026 &nbsp;&nbsp;&nbsp; arXiv:2603.13741 We present Ego-1K, a large-scale collection of time-synchronized egocentric multiview videos designed to advance neural 3D video synthesis and dynamic scene understanding.… See the full description on the dataset page: https://huggingface.co/datasets/facebook/ego-1k.tabulardepth-estimation100K<n<1M17 likes53k downloads3mo agoHugging Face08facebook /wiki_dprThis is the wikipedia split used to evaluate the Dense Passage Retrieval (DPR) model. It contains 21M passages from wikipedia along with their DPR embeddings. The wikipedia articles were split into multiple, disjoint text blocks of 100 words as passages.fill-mask10M<n<100M45 likes36k downloads3y agoHugging Face09facebook /multilingual_librispeech Dataset Card for MultiLingual LibriSpeech Dataset Summary This is a streamable version of the Multilingual LibriSpeech (MLS) dataset. The data archives were restructured from the original ones from OpenSLR to make it easier to stream. MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish.… See the full description on the dataset page: https://huggingface.co/datasets/facebook/multilingual_librispeech.audioautomatic-speech-recognition1M<n<10M190 likes34k downloads2y agoHugging Face10facebook /show3d-dataset SHOW3D: Capturing Scenes of 3D Hands and Objects in the Wild Patrick Rim, Kevin Harris, Braden Copple, Shangchen Han, Xu Xie, Ivan Shugurov, Sizhe An, He Wen, Alex Wong, Tomas Hodan, and Kun He CVPR 2026; https://arxiv.org/abs/2603.28760 News September 18, 2026: Released synchronized exocentric views and camera calibrations. SHOW3D is a large-scale multi-view dataset of hand–object interactions captured in the wild. It is intended to advance research on… See the full description on the dataset page: https://huggingface.co/datasets/facebook/show3d-dataset.tabularother1K<n<10K6 likes20k downloads2d agoHugging Face11facebook /CoTracker3_Kubric Kubric Dataset for CoTracker 3 Overview This dataset was specifically created for training CoTracker 3, a state-of-the-art point tracking model. The dataset was generated using the Kubric engine. Dataset Specifications Size: ~6,000 sequences Resolution: 512×512 pixels Sequence Length: 120 frames per sequence Camera Movement: Carefully rendered with subtle camera motion to simulate realistic scenarios Format: Generated using Kubric engine Usage The… See the full description on the dataset page: https://huggingface.co/datasets/facebook/CoTracker3_Kubric.10 likes12k downloads2y agoHugging Face12facebook /map-anything MapAnything Training Metadata Dataset Dataset Description This dataset contains pre-computed metadata and covisibility matrices for supporting the MapAnything codebase. This metadata enables easy reproducible training for feed-forward 3D reconstruction tasks. Please see our Data Processing README for more details. Citation If you use this dataset in your research, please cite our paper: @inproceedings{keetha2026mapanything, title={{MapAnything}: Universal… See the full description on the dataset page: https://huggingface.co/datasets/facebook/map-anything.image-to-3d100B<n<1T9 likes8.2k downloads8mo agoHugging Face13facebook /hand_tracking_challenge_umetrackCheck out the Multiview Egocentric Hand Tracking Challenge 2024!! To use this dataset, check out the hand_tracking_toolkit image100K<n<1M1 likes7.8k downloads2y agoHugging Face14facebook /PE-Video PE Video Dataset (PVD) [📃 Tech Report] [📂 Github] The PE Video Dataset (PVD) is a large-scale collection of 1 million diverse videos, featuring 120,000+ expertly annotated clips. The dataset was introduced in our paper "Perception Encoder". Overview PE Video Dataset (PVD) comprises 1M high quality and diverse videos. Among them, 120K videos are accompanied by automated and human-verified annotations. and all videos are accompanied with video description and keywords.… See the full description on the dataset page: https://huggingface.co/datasets/facebook/PE-Video.text100K<n<1M51 likes6k downloads1y agoHugging Face15facebook /floresgated Dataset Card for Flores 200 Dataset Summary ⚠️ This repository is no longer being updated ⚠️ A newer version of the FLORES dataset managed by the Open Language Data Initiative is available at https://huggingface.co/datasets/openlanguagedata/flores_plus. FLORES is a benchmark dataset for machine translation between English and low-resource languages. The creation of FLORES-200 doubles the existing language coverage of FLORES-101. Given the nature of the new… See the full description on the dataset page: https://huggingface.co/datasets/facebook/flores.tabulartext-generation1M<n<10M120 likes5.2k downloads4mo agoHugging Face16facebook /meta-active-readingtext1B<n<10B37 likes4.6k downloads1y agoHugging Face17facebook /mlqa MLQA (MultiLingual Question Answering) is a benchmark dataset for evaluating cross-lingual question answering performance. MLQA consists of over 5K extractive QA instances (12K in English) in SQuAD format in seven languages - English, Arabic, German, Spanish, Hindi, Vietnamese and Simplified Chinese. MLQA is highly parallel, with QA instances parallel between 4 different languages on average.question-answering10K<n<100K44 likes4.5k downloads3y agoHugging Face18facebook /kilt_tasks Dataset Card for KILT Dataset Summary KILT has been built from 11 datasets representing 5 types of tasks: Fact-checking Entity linking Slot filling Open domain QA Dialog generation All these datasets have been grounded in a single pre-processed Wikipedia dump, allowing for fairer and more consistent evaluation as well as enabling new task setups such as multitask and transfer learning with minimal effort. KILT also provides tools to analyze and understand the… See the full description on the dataset page: https://huggingface.co/datasets/facebook/kilt_tasks.textfill-mask1M<n<10M68 likes4.5k downloads3y agoHugging Face19facebook /S-EMBERgated S-EMBER: A Large-Scale Benchmark for Streaming Egocentric Memory Retrieval Episodic-memory video QA benchmark (face-blurred, audio-removed). License & usage This dataset is licensed under CC BY-NC 4.0 and is provided for non-commercial research use only. Access is gated: you must accept the non-commercial terms above before downloading. Contents sember_mcq.jsonl — multiple-choice evaluation split. sember_grounding.jsonl — answer-generation and… See the full description on the dataset page: https://huggingface.co/datasets/facebook/S-EMBER.tabular10K<n<100K6 likes4.4k downloads3mo agoHugging Face20facebook /omnilingual-asr-corpus Meta Omnilingual ASR Corpus The Omnilingual ASR Corpus is a collection of spontaneous speech recordings and their transcriptions for 348 under-served languages. The corpus was collected as part of Meta FAIR’s Omnilingual ASR project (blog, model, paper) for the purposes of training automatic speech recognition (ASR) and spoken language identification models. Data schema { `language`: "lij_Latn", `iso_639_3`: "lij", `iso_15924`: "Latn", `glottocode`:… See the full description on the dataset page: https://huggingface.co/datasets/facebook/omnilingual-asr-corpus.audioautomatic-speech-recognition100K<n<1M210 likes4.1k downloads10mo agoHugging Face21facebook /IntPhys2 IntPhys 2 Dataset   |   Hugging Face   |   Paper   |   Blog IntPhys 2 is a video benchmark designed to evaluate the intuitive physics understanding of deep learning models. Building on the original IntPhys benchmark, IntPhys 2 focuses on four core principles related to macroscopic objects: Permanence, Immutability, Spatio-Temporal Continuity, and Solidity. These conditions are inspired by research into intuitive physical understanding emerging during early childhood. IntPhys 2… See the full description on the dataset page: https://huggingface.co/datasets/facebook/IntPhys2.text1K<n<10K14 likes4k downloads1y agoHugging Face22facebook /md_gender_biasMachine learning models are trained to find patterns in data. NLP models can inadvertently learn socially undesirable patterns when training on gender biased text. In this work, we propose a general framework that decomposes gender bias in text along several pragmatic and semantic dimensions: bias from the gender of the person being spoken about, bias from the gender of the person being spoken to, and bias from the gender of the speaker. Using this fine-grained framework, we automatically annotate eight large scale datasets with gender information. In addition, we collect a novel, crowdsourced evaluation benchmark of utterance-level gender rewrites. Distinguishing between gender bias along multiple dimensions is important, as it enables us to train finer-grained gender bias classifiers. We show our classifiers prove valuable for a variety of important applications, such as controlling for gender bias in generative models, detecting gender bias in arbitrary text, and shed light on offensive language in terms of genderedness.text-classification100K<n<1M19 likes3k downloads3y agoHugging Face23facebook /wearable-aigated EgoWearBench Dataset (ECCV 2026) Part of the Wearable AI Workshop at ECCV 2026. A benchmark of egocentric (first-person, head-mounted wearable camera) videos paired with three complementary video question-answering tasks for evaluating wearable-AI assistants on real-world everyday activity videos. ▶ Baseline code & evaluation scripts: see starter_kit/README.md. The starter kit ships inside this repo, so git clone gives you the code and the data together. Tasks… See the full description on the dataset page: https://huggingface.co/datasets/facebook/wearable-ai.textvisual-question-answering1K<n<10K24 likes2.9k downloads13d agoHugging Face24Sami96ads61 /facebook-bot-mediaimagen<1K1 likes2.8k downloads7h agoHugging Face25facebook /gelsight-force-estimation Dataset Details This dataset contains paired tactile and force data, intended for use in predicting 3-axis normal and shear forces applied to the sensor's elastomer. We used three different indenter shapes to collect force-labeled data: hemisphere, sharp, and flat. To measure force ground truths, we employed the ATI nano17 force/torque sensor. The protocol consisted of applying a random normal load (up to 3N) followed by a shear load, achieved by sliding the probe 2mm on the… See the full description on the dataset page: https://huggingface.co/datasets/facebook/gelsight-force-estimation.imagen<1K2 likes2.8k downloads2y agoHugging Face26facebook /natural_reasoningNaturalReasoning is a large-scale dataset for general reasoning tasks. It consists of high-quality challenging reasoning questions backtranslated from pretraining corpora DCLM and FineMath. The questions have been deduplicated and decontaminated from popular reasoning benchmarks including MATH, GPQA, MMLU-Pro, MMLU-STEM. For each question, we extract the reference final answer from the original document from the pretraining corpora if possible. We also provide a model-generated response from… See the full description on the dataset page: https://huggingface.co/datasets/facebook/natural_reasoning.texttext-generation1M<n<10M585 likes2.8k downloads2y agoHugging Face27facebook /empathetic_dialoguesPyTorch original implementation of Towards Empathetic Open-domain Conversation Models: a New Benchmark and Datasetquestion-answering10K<n<100K134 likes2.7k downloads3y agoHugging Face28facebook /cyberseceval3-visual-prompt-injection Dataset Card for CyberSecEval 3 - Visual Prompt Injection Benchmark Dataset Details Dataset Description This dataset provides a multimodal benchmark for visual prompt injection, with text/image inputs. It is part of CyberSecEval 3, the third edition of Meta's flagship suite of security benchmarks for LLMs to measure cybersecurity risks and capabilities across multiple domains. Language(s): English License: MIT Dataset Sources Repository: Link… See the full description on the dataset page: https://huggingface.co/datasets/facebook/cyberseceval3-visual-prompt-injection.imagetext-generation1K<n<10K10 likes2.7k downloads2y agoHugging Face29facebook /babi_qaThe (20) QA bAbI tasks are a set of proxy tasks that evaluate reading comprehension via question answering. Our tasks measure understanding in several ways: whether a system is able to answer questions via chaining facts, simple induction, deduction and many more. The tasks are designed to be prerequisites for any system that aims to be capable of conversing with a human. The aim is to classify these tasks into skill sets,so that researchers can identify (and then rectify)the failings of their systems.textquestion-answering10K<n<100K13 likes2.2k downloads4y agoHugging Face30facebook /2M-Belebele 2M-Belebele Highly-Multilingual Speech and American Sign Language Comprehension Dataset We introduce 2M-Belebele as the first highly multilingual speech and American Sign Language (ASL) comprehension dataset. Our dataset, which is an extension of the existing Belebele only-text dataset, covers 74 spoken languages at the intersection of Belebele and Fleurs, and one sign language (ASL). The speech dataset is built from aligning Belebele, Flores200 and Fleurs datasets as… See the full description on the dataset page: https://huggingface.co/datasets/facebook/2M-Belebele.tabularquestion-answering10K<n<100K13 likes2.2k downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.