CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01FBK-MT /MCIF Dataset Description, Collection, and Source MCIF (Multimodal Crosslingual Instruction Following) is a multilingual human-annotated benchmark based on scientific talks that is designed to evaluate instruction-following in crosslingual, multimodal settings over both short- and long-form inputs. MCIF spans three core modalities -- speech, vision, and text -- and four diverse languages (English, German, Italian, and Chinese), enabling a comprehensive evaluation of MLLMs'… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/MCIF.audioautomatic-speech-recognition1K<n<10K70 likes1.5k downloads3mo agoHugging Face02McGill-NLP /speech-translation-and-summarization English-Centric Multilingual Audio Dataset This dataset contains generated article and summary audio for English-centric multilingual directions. Each direction folder contains metadata JSONL files and corresponding audio files for few_shot and test splits. Included directions amharic_english / english_amharic arabic_english / english_arabic bengali_english / english_bengali chinese_simplified_english / english_chinese_simplified english_english french_english /… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/speech-translation-and-summarization.audioautomatic-speech-recognition10K<n<100K6 likes768 downloads2mo agoHugging Face03czyzi0 /the-mc-speech-datasetThis is public domain speech dataset consisting of 24018 short audio clips of a single speaker reading sentences in Polish. A transcription is provided for each clip. Clips have total length of more than 22 hours. Texts are in public domain. The audio was recorded in 2021-22 as a part of my master's thesis and is in public domain. If you use this dataset, please cite: @masterthesis{mcspeech, title={Analiza porównawcza korpusów nagrań mowy dla celów syntezy mowy w języku polskim}… See the full description on the dataset page: https://huggingface.co/datasets/czyzi0/the-mc-speech-dataset.audiotext-to-speech10K<n<100K8 likes324 downloads3y agoHugging Face04FBK-MT /MCIF-ST MCIF-ST: Context-aware Speech Recognition and Speech Translation from MCIF MCIF-ST provides both long-form and short-form ready-to-use Automatic Speech Recogniton (ASR) and Speech Translation (ST) data derived from MCIF (Multimodal Crosslingual Instruction Following), a multilingual benchmark based on scientific talks. While the original MCIF release packages its content as instruction-following rows (multimodal context + prompt + expected answer, for… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/MCIF-ST.audioautomatic-speech-recognition1K<n<10K0 likes169 downloads2mo agoHugging Face05Rendra86318 /MCIF Dataset Description, Collection, and Source MCIF (Multimodal Crosslingual Instruction Following) is a multilingual human-annotated benchmark based on scientific talks that is designed to evaluate instruction-following in crosslingual, multimodal settings over both short- and long-form inputs. MCIF spans three core modalities -- speech, vision, and text -- and four diverse languages (English, German, Italian, and Chinese), enabling a comprehensive evaluation of MLLMs'… See the full description on the dataset page: https://huggingface.co/datasets/Rendra86318/MCIF.audioautomatic-speech-recognition1K<n<10K0 likes74 downloads9mo agoHugging Face06MCAA1-MSU /twb_luo_kegated Utavu Foundation Dholuo Speech Dataset This is the dataset card for the Utavu Foundation and CLEAR Global TWB Voice pilot. Dataset summary This dataset contains spontaneous Dholuo speech about agriculture. Speakers answered open questions in Dholuo and spoke freely in response. Each recording has a Dholuo transcription. The data was collected through the TWB Voice platform and is intended for speech technology research and language technology development for… See the full description on the dataset page: https://huggingface.co/datasets/MCAA1-MSU/twb_luo_ke.audioautomatic-speech-recognition1K<n<10K0 likes74 downloads17h agoHugging Face07vaishnavikedar4 /MCIF Dataset Description, Collection, and Source MCIF (Multimodal Crosslingual Instruction Following) is a multilingual human-annotated benchmark based on scientific talks that is designed to evaluate instruction-following in crosslingual, multimodal settings over both short- and long-form inputs. MCIF spans three core modalities -- speech, vision, and text -- and four diverse languages (English, German, Italian, and Chinese), enabling a comprehensive evaluation of MLLMs'… See the full description on the dataset page: https://huggingface.co/datasets/vaishnavikedar4/MCIF.audioautomatic-speech-recognition1K<n<10K0 likes41 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.