CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01FBK-MT /MCIF Dataset Description, Collection, and Source MCIF (Multimodal Crosslingual Instruction Following) is a multilingual human-annotated benchmark based on scientific talks that is designed to evaluate instruction-following in crosslingual, multimodal settings over both short- and long-form inputs. MCIF spans three core modalities -- speech, vision, and text -- and four diverse languages (English, German, Italian, and Chinese), enabling a comprehensive evaluation of MLLMs'… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/MCIF.audioautomatic-speech-recognition1K<n<10K70 likes1.5k downloads2mo agoHugging Face02McGill-NLP /speech-translation-and-summarization English-Centric Multilingual Audio Dataset This dataset contains generated article and summary audio for English-centric multilingual directions. Each direction folder contains metadata JSONL files and corresponding audio files for few_shot and test splits. Included directions amharic_english / english_amharic arabic_english / english_arabic bengali_english / english_bengali chinese_simplified_english / english_chinese_simplified english_english french_english /… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/speech-translation-and-summarization.audioautomatic-speech-recognition10K<n<100K6 likes760 downloads1mo agoHugging Face03czyzi0 /the-mc-speech-datasetThis is public domain speech dataset consisting of 24018 short audio clips of a single speaker reading sentences in Polish. A transcription is provided for each clip. Clips have total length of more than 22 hours. Texts are in public domain. The audio was recorded in 2021-22 as a part of my master's thesis and is in public domain. If you use this dataset, please cite: @masterthesis{mcspeech, title={Analiza porównawcza korpusów nagrań mowy dla celów syntezy mowy w języku polskim}… See the full description on the dataset page: https://huggingface.co/datasets/czyzi0/the-mc-speech-dataset.audiotext-to-speech10K<n<100K8 likes327 downloads3y agoHugging Face04FBK-MT /MCIF-ST MCIF-ST: Context-aware Speech Recognition and Speech Translation from MCIF MCIF-ST provides both long-form and short-form ready-to-use Automatic Speech Recogniton (ASR) and Speech Translation (ST) data derived from MCIF (Multimodal Crosslingual Instruction Following), a multilingual benchmark based on scientific talks. While the original MCIF release packages its content as instruction-following rows (multimodal context + prompt + expected answer, for… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/MCIF-ST.audioautomatic-speech-recognition1K<n<10K0 likes162 downloads2mo agoHugging Face05mcamara /vtl-speech-landmarks VTL Speech Landmarks Dataset Articulatory speech synthesis dataset with acoustic landmarks, generated using VocalTractLab (VTL). Dataset Description This dataset contains synthesized speech for 117,497 English words from the CMU Pronouncing Dictionary, generated with two speakers (male and female). Each word includes: Audio: 48kHz WAV files Landmarks: Acoustic-phonetic event markers (JSON) Articulatory data: Full vocal tract trajectories from VTL (JSON) Speakers… See the full description on the dataset page: https://huggingface.co/datasets/mcamara/vtl-speech-landmarks.automatic-speech-recognition100K<n<1M0 likes155 downloads8mo agoHugging Face06yxdu /MCGA MCGA: A Multi-task Classical Chinese Literary Genre Audio Corpus MCGA (Multi-task Classical Chinese Literary Genre Audio Corpus) is the first large-scale, open-source, and fully copyrighted audio corpus dedicated to Classical Chinese Studies, comprising 119 hours (22,000 samples) of standard Mandarin recordings by native speakers that span five major literary genres (Fu, Shi, Wen, Ci, and Qu) across 11 historical periods, specifically constructed to support six core… See the full description on the dataset page: https://huggingface.co/datasets/yxdu/MCGA.automatic-speech-recognition1 likes134 downloads2mo agoHugging Face07Rendra86318 /MCIF Dataset Description, Collection, and Source MCIF (Multimodal Crosslingual Instruction Following) is a multilingual human-annotated benchmark based on scientific talks that is designed to evaluate instruction-following in crosslingual, multimodal settings over both short- and long-form inputs. MCIF spans three core modalities -- speech, vision, and text -- and four diverse languages (English, German, Italian, and Chinese), enabling a comprehensive evaluation of MLLMs'… See the full description on the dataset page: https://huggingface.co/datasets/Rendra86318/MCIF.audioautomatic-speech-recognition1K<n<10K0 likes72 downloads9mo agoHugging Face08vaishnavikedar4 /MCIF Dataset Description, Collection, and Source MCIF (Multimodal Crosslingual Instruction Following) is a multilingual human-annotated benchmark based on scientific talks that is designed to evaluate instruction-following in crosslingual, multimodal settings over both short- and long-form inputs. MCIF spans three core modalities -- speech, vision, and text -- and four diverse languages (English, German, Italian, and Chinese), enabling a comprehensive evaluation of MLLMs'… See the full description on the dataset page: https://huggingface.co/datasets/vaishnavikedar4/MCIF.audioautomatic-speech-recognition1K<n<10K0 likes39 downloads9mo agoHugging Face09mcapozi /voxpopolo_2A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation.automatic-speech-recognition0 likes12 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.