datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
seamless-interaction
Seamless Interaction Dataset
A large-scale multimodal dataset of 4,000+ hours of human interactions for AI research
🖼️ Blog
🌐 Website
🎮 Demo
📦 GitHub
📄 Paper
Human communication involves a complex interplay of verbal and nonverbal signals, essential for conveying meaning and achieving interpersonal goals.
The Seamless Interaction Dataset is a large-scale collection of over 4,000 hours of face-to-face interaction footage from more than 4,000 participants in… See the full description on the dataset page: https://huggingface.co/datasets/facebook/seamless-interaction.voxpopuli
Dataset Card for Voxpopuli
Dataset Summary
VoxPopuli is a large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation.
The raw data is collected from 2009-2020 European Parliament event recordings. We acknowledge the European Parliament for creating and sharing these materials.
This implementation contains transcribed speech data for 18 languages.
It also contains 29 hours of transcribed speech data of non-native… See the full description on the dataset page: https://huggingface.co/datasets/facebook/voxpopuli.multilingual_librispeech
Dataset Card for MultiLingual LibriSpeech
Dataset Summary
This is a streamable version of the Multilingual LibriSpeech (MLS) dataset.
The data archives were restructured from the original ones from OpenSLR to make it easier to stream.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of
8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish.… See the full description on the dataset page: https://huggingface.co/datasets/facebook/multilingual_librispeech.omnilingual-asr-corpus
Meta Omnilingual ASR Corpus
The Omnilingual ASR Corpus is a collection of spontaneous speech recordings and their transcriptions for 348 under-served languages. The corpus was collected as part of Meta FAIR’s Omnilingual ASR project (blog, model, paper) for the purposes of training automatic speech recognition (ASR) and spoken language identification models.
Data schema
{
`language`: "lij_Latn",
`iso_639_3`: "lij",
`iso_15924`: "Latn",
`glottocode`:… See the full description on the dataset page: https://huggingface.co/datasets/facebook/omnilingual-asr-corpus.2M-Belebele
2M-Belebele
Highly-Multilingual Speech and American Sign Language Comprehension Dataset
We introduce 2M-Belebele as the first highly multilingual speech and American Sign Language (ASL) comprehension dataset. Our dataset, which is an extension of the existing Belebele only-text dataset, covers 74 spoken languages at the intersection of Belebele and Fleurs, and one sign language (ASL).
The speech dataset is built from aligning Belebele, Flores200 and Fleurs datasets as… See the full description on the dataset page: https://huggingface.co/datasets/facebook/2M-Belebele.EgoAVU_data
[CVPR2026 HIGHLIGHT] EgoAVU, [ICASSP2026 Oral] Exploring Audio Hallucination in Egocentric Video Understanding
Official Implementation of EgoAVU: Egocentric Audio-Visual Understanding and Exploring Audio Hallucination in Egocentric Video Understanding
See our github for the code and setup instructions.
Check out our homepage, paper (CVPR) and paper (ICASSP) for more information.
We introduce EgoAVU, a scalable and automated data engine to enable egocentric audio–visual… See the full description on the dataset page: https://huggingface.co/datasets/facebook/EgoAVU_data.faceit_top_demos_883_voice_split_cut_20251126facebook_multilingual_librispeech
This dataset only contains test data, which is integrated into UltraEval-Audio(https://github.com/OpenBMB/UltraEval-Audio) framework.
python audio_evals/main.py --dataset mls_dutch --model gpt4o_audio
python audio_evals/main.py --dataset mls_french --model gpt4o_audio
python audio_evals/main.py --dataset mls_german --model gpt4o_audio
python audio_evals/main.py --dataset mls_italian --model gpt4o_audio
python audio_evals/main.py --dataset mls_polish --model gpt4o_audio
python… See the full description on the dataset page: https://huggingface.co/datasets/TwinkStart/facebook_multilingual_librispeech.mmione_voice_FACEBOOK_PARQUET
Artificial Omnivoice Hungarian Speaker Dataset
Ez egy teljesen szintetikus magyar nyelvű beszédadatbázis, amely kiváló minőségű szövegfelolvasó (TTS) és beszédfelismerő (ASR) modellek tanításához és finomhangolásához készült.
Adatforrás és Referencia Hang
A dataset alapjául szolgáló referencia hang (speaker identity) egy 20 másodperces részlet az alábbi YouTube videóból:
Forrás: Hogyan legyél tökéletes magyar várvédő tutorial
Licenc: A videó CC (Creative Commons)… See the full description on the dataset page: https://huggingface.co/datasets/fablevi/one_voice_FACEBOOK_PARQUET.female_facebook_datafacebook_mms-tts-eng_GPU-CPUfacebook_voxpopulik_16k_Whisper_Compatibleenhanced_facebook_voxpopulik_16k_Whisper_Compatiblegolos-annotation-length-basedcommon-voice-speaker-parameters-basedfaceocultanoise-augmented-russian-librispeechfacebook_m4t_v2_syndata
