datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MCIF
Dataset Description, Collection, and Source
MCIF (Multimodal Crosslingual Instruction Following) is a multilingual human-annotated benchmark
based on scientific talks that is designed to evaluate instruction-following in crosslingual,
multimodal settings over both short- and long-form inputs.
MCIF spans three core modalities -- speech, vision, and text -- and four diverse languages (English, German, Italian, and Chinese),
enabling a comprehensive evaluation of MLLMs'… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/MCIF.speech-translation-and-summarization
English-Centric Multilingual Audio Dataset
This dataset contains generated article and summary audio for English-centric multilingual directions.
Each direction folder contains metadata JSONL files and corresponding audio files for few_shot and test splits.
Included directions
amharic_english / english_amharic
arabic_english / english_arabic
bengali_english / english_bengali
chinese_simplified_english / english_chinese_simplified
english_english
french_english /… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/speech-translation-and-summarization.the-mc-speech-datasetThis is public domain speech dataset consisting of 24018 short audio clips of a single speaker reading sentences in Polish. A transcription is provided for each clip. Clips have total length of more than 22 hours.
Texts are in public domain. The audio was recorded in 2021-22 as a part of my master's thesis and is in public domain.
If you use this dataset, please cite:
@masterthesis{mcspeech,
title={Analiza porównawcza korpusów nagrań mowy dla celów syntezy mowy w języku polskim}… See the full description on the dataset page: https://huggingface.co/datasets/czyzi0/the-mc-speech-dataset.MCIF-ST
MCIF-ST: Context-aware Speech Recognition and Speech Translation from MCIF
MCIF-ST provides both long-form and short-form ready-to-use
Automatic Speech Recogniton (ASR) and Speech Translation (ST) data derived
from MCIF (Multimodal
Crosslingual Instruction Following), a multilingual benchmark based on
scientific talks. While the original MCIF release packages its content as
instruction-following rows (multimodal context + prompt + expected
answer, for… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/MCIF-ST.vtl-speech-landmarks
VTL Speech Landmarks Dataset
Articulatory speech synthesis dataset with acoustic landmarks, generated using VocalTractLab (VTL).
Dataset Description
This dataset contains synthesized speech for 117,497 English words from the CMU Pronouncing Dictionary, generated with two speakers (male and female). Each word includes:
Audio: 48kHz WAV files
Landmarks: Acoustic-phonetic event markers (JSON)
Articulatory data: Full vocal tract trajectories from VTL (JSON)
Speakers… See the full description on the dataset page: https://huggingface.co/datasets/mcamara/vtl-speech-landmarks.MCGA
MCGA: A Multi-task Classical Chinese Literary Genre Audio Corpus
MCGA (Multi-task Classical Chinese Literary Genre Audio Corpus) is the first large-scale, open-source, and fully copyrighted audio corpus dedicated to Classical Chinese Studies, comprising 119 hours (22,000 samples) of standard Mandarin recordings by native speakers that span five major literary genres (Fu, Shi, Wen, Ci, and Qu) across 11 historical periods, specifically constructed to support six core… See the full description on the dataset page: https://huggingface.co/datasets/yxdu/MCGA.MCIF
Dataset Description, Collection, and Source
MCIF (Multimodal Crosslingual Instruction Following) is a multilingual human-annotated benchmark
based on scientific talks that is designed to evaluate instruction-following in crosslingual,
multimodal settings over both short- and long-form inputs.
MCIF spans three core modalities -- speech, vision, and text -- and four diverse languages (English, German, Italian, and Chinese),
enabling a comprehensive evaluation of MLLMs'… See the full description on the dataset page: https://huggingface.co/datasets/Rendra86318/MCIF.MCIF
Dataset Description, Collection, and Source
MCIF (Multimodal Crosslingual Instruction Following) is a multilingual human-annotated benchmark
based on scientific talks that is designed to evaluate instruction-following in crosslingual,
multimodal settings over both short- and long-form inputs.
MCIF spans three core modalities -- speech, vision, and text -- and four diverse languages (English, German, Italian, and Chinese),
enabling a comprehensive evaluation of MLLMs'… See the full description on the dataset page: https://huggingface.co/datasets/vaishnavikedar4/MCIF.voxpopolo_2A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation.
