datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Speech-MASSIVE
Speech-MASSIVE
Dataset Description
Speech-MASSIVE is a multilingual Spoken Language Understanding (SLU) dataset comprising the speech counterpart for a portion of the MASSIVE textual corpus. Speech-MASSIVE covers 12 languages (Arabic, German, Spanish, French, Hungarian, Korean, Dutch, Polish, European Portuguese, Russian, Turkish, and Vietnamese) from different families and inherits from MASSIVE the annotations for the intent prediction and slot-filling tasks. MASSIVE… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/Speech-MASSIVE.mosel
Dataset Description, Collection, and Source
The MOSEL corpus is a multilingual dataset collection including up to 950K hours of open-source speech recordings covering the 24 official languages of the European Union. We collect data by surveying labeled and unlabeled speech corpora under open-source compliant licenses.
In particular, MOSEL includes the automatic transcripts of 441k hours of unlabeled speech from VoxPopuli and LibriLight. The data is transcribed using Whisper large… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/mosel.mtedx
Multilingual TEDx (mTEDx) — SLR100
mTEDx is a multilingual speech recognition and translation corpus built from
TEDx Talks.Original resource: https://www.openslr.org/100/
The corpus provides audio recordings and VTT transcripts for 8 languages
(Spanish, French, Portuguese, Italian, Russian, Greek, Arabic, German) with
aligned translations into up to 5 languages (English, Spanish, French,
Portuguese, Italian).
License: CC BY-NC-ND 4.0Contact: Elizabeth Salesky (esalesky@jhu.edu)… See the full description on the dataset page: https://huggingface.co/datasets/deepdml/mtedx.MCIF
Dataset Description, Collection, and Source
MCIF (Multimodal Crosslingual Instruction Following) is a multilingual human-annotated benchmark
based on scientific talks that is designed to evaluate instruction-following in crosslingual,
multimodal settings over both short- and long-form inputs.
MCIF spans three core modalities -- speech, vision, and text -- and four diverse languages (English, German, Italian, and Chinese),
enabling a comprehensive evaluation of MLLMs'… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/MCIF.mTEDx-ptbr
Multilingual TEDx (Portuguese speech and transcripts)
NOTE: This dataset contains only the Portuguese portion of the mTEDx dataset, already processed and segmented into parts.
Multilingual TEDx (mTEDx) is a multilingual speech recognition and translation corpus to facilitate the training of ASR and SLT models in additional languages.
The corpus comprises audio recordings and transcripts from TEDx Talks in 8 languages (Spanish, French, Portuguese, Italian, Russian, Greek, Arabic… See the full description on the dataset page: https://huggingface.co/datasets/dominguesm/mTEDx-ptbr.MTR-DuplexBench
MTR-DuplexBench: Towards a Comprehensive Evaluation of Multi-Round Conversations for Full-Duplex Speech Language Models
🎉🎉 MTR-DuplexBench has been accepted by ACL 2026 Findings!
📄 Paper | 🤗 HuggingFace Dataset
Affiliations: Tsinghua University, The Chinese University of Hong Kong, Huawei
Dataset Description
This dataset is constructed to evaluate multi-modal audio models across four critical dimensions: Conversational Features, Instruction Following, Safety… See the full description on the dataset page: https://huggingface.co/datasets/Jeff0918/MTR-DuplexBench.svq
Simple Voice Questions
Simple Voice Questions (SVQ) is a set of short audio questions recorded in 26 locales across 17 languages under multiple audio conditions.
Data Collection
Speakers were presented with recording instructions specifying the recording environment and text query to be recorded.
They recorded using their own phones or tablets under four conditions:
clean: Record in quiet environment
background speech noise: Record while audio from sources like podcasts… See the full description on the dataset page: https://huggingface.co/datasets/mteb/svq.fama-data
Dataset Description, Collection, and Source
The FAMA training data is the collection of English and Italian datasets for automatic speech recognition (ASR) and speech translation (ST)
used to train the FAMA models family.
The ASR section of FAMA is derived from the MOSEL data collection, including the automatic
transcripts obtained with Whisper and available in the HuggingFace MOSEL Dataset.
The ASR is further augmented with automatically transcribed speech from the… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/fama-data.MCIF-ST
MCIF-ST: Context-aware Speech Recognition and Speech Translation from MCIF
MCIF-ST provides both long-form and short-form ready-to-use
Automatic Speech Recogniton (ASR) and Speech Translation (ST) data derived
from MCIF (Multimodal
Crosslingual Instruction Following), a multilingual benchmark based on
scientific talks. While the original MCIF release packages its content as
instruction-following rows (multimodal context + prompt + expected
answer, for… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/MCIF-ST.Speech-MASSIVE-test
Speech-MASSIVE Test Split
This dataset repository is only for test split of Speech-MASSIVE.
train and dev splits are available in the separate dataset repository. https://huggingface.co/datasets/FBK-MT/Speech-MASSIVE
Dataset Description
Speech-MASSIVE is a multilingual Spoken Language Understanding (SLU) dataset comprising the speech counterpart for a portion of the MASSIVE textual corpus. Speech-MASSIVE covers 12 languages (Arabic, German, Spanish, French, Hungarian… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/Speech-MASSIVE-test.mTEDx-ptbr
Multilingual TEDx (Portuguese speech and transcripts)
NOTE: This dataset contains only the Portuguese portion of the mTEDx dataset, already processed and segmented into parts.
Multilingual TEDx (mTEDx) is a multilingual speech recognition and translation corpus to facilitate the training of ASR and SLT models in additional languages.
The corpus comprises audio recordings and transcripts from TEDx Talks in 8 languages (Spanish, French, Portuguese, Italian, Russian, Greek, Arabic… See the full description on the dataset page: https://huggingface.co/datasets/FERNAN89/mTEDx-ptbr.SpokenWords-GA-EN-MTed
Dataset Card for Dataset Name
This is the Irish portion of the Spoken Words dataset (available at MLCommons/ml_spoken_words),
with merged splits “train”, “validation”, and “test”, augmented with machine translation.
The Irish sentences are automatically translated into English using Google Translation API.
The dataset includes approximately 3 hours and 2 minutes of audio (03:02:02), spoken by multiple narrators.
Dataset Structure
Dataset({
features: ['keyword'… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/SpokenWords-GA-EN-MTed.mtedx-v-eval
mTEDx-V Eval: Long-Form Multilingual X->En Speech Translation with Recoverable Video Evidence
Talk-level (long-form) evaluation manifests for visual-context-aware simultaneous
speech translation, built from the Multilingual TEDx (mTEDx)
corpus. Each record is one full TEDx talk whose talk_id is the real YouTube video ID,
so the original talk video can be obtained from its official source and frames can be
aligned to the sentence-level segment timestamps below (timestamps are on… See the full description on the dataset page: https://huggingface.co/datasets/gavinlaw/mtedx-v-eval.banc-trawsgrifiadau-bangor_mt
Dataset Card for banc-trawsgrifiadau-bangor_mt
Source Dataset Provenance
This dataset was derived from the source dataset at commit SHA: da64c66e467dac267ee3aeaa2f520ef82722379e
(short: da64c66e)
