CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01espnet /yodas-granary Dataset Card for YODAS-Granary Repository: NeMo-speech-data-processor: Granary Paper: Granary: Speech Recognition and Translation Dataset in 25 European Languages Shared by: ESPnet Dataset Description YODAS-Granary is a curated subset of the larger nvidia/Granary dataset, focusing on high-quality pseudo-labeled speech data for Automatic Speech Recognition (ASR) and Automatic Speech Translation (AST) across 23 European languages. Overview… See the full description on the dataset page: https://huggingface.co/datasets/espnet/yodas-granary.audioautomatic-speech-recognition10M<n<100M33 likes75k downloads1y agoHugging Face02espnet /yodasUpdates 2024/07/09: we also uploaded a new version of YODAS as YODAS2, it provides unsegmented audios and higher sampling rate (24k) README This is the YODAS manual/automatic subset from our YODAS dataset, it has 369,510 hours of speech. This dataset contains audio utterances and corresponding captions (manual or automatic) from YouTube. Note that manual caption only indicates that it is uploaded by users, but not necessarily transcribed by a human For more details about YODAS… See the full description on the dataset page: https://huggingface.co/datasets/espnet/yodas.155 likes55k downloads2y agoHugging Face03espnet /yodas2YODAS2 is the long-form dataset from YODAS dataset. It provides the same dataset as espnet/yodas but YODAS2 has the following new features: formatted in the long-form (video-level) where audios are not segmented. audios are encoded using higher sampling rates (i.e. 24k) For detailed information about YODAS dataset, please refer to our paper and the espnet/yodas repo. Usage: Each data point corresponds to an entire video on YouTube, it contains the following fields: video_id:… See the full description on the dataset page: https://huggingface.co/datasets/espnet/yodas2.56 likes38k downloads1y agoHugging Face04espnet /Bagpiper_SFT_Data Bagpiper SFT Data Release status: the validated Parquet release is being uploaded. The homepage and metadata may appear before every large shard is committed. Bagpiper SFT Data is the supervised fine-tuning corpus for Bagpiper, an open-ended audio language model that understands and generates speech, music, environmental sound, and their mixtures through rich textual captions and planning. The public release has exactly two configurations: Configuration Direction… See the full description on the dataset page: https://huggingface.co/datasets/espnet/Bagpiper_SFT_Data.audioaudio-classification1M<n<10M1 likes6.5k downloads2mo agoHugging Face05espnet /floras FLORAS FLORAS is a 50-language benchmark For LOng-form Recognition And Summarization of spoken language. The goal of FLORAS is to create a more realistic benchmarking environment for speech recognition, translation, and summarization models. Unlike typical academic benchmarks like LibriSpeech and FLEURS that uses pre-segmented single-speaker read-speech, FLORAS tests the capabilities of models on raw long-form conversational audio, which can have one or many speakers. To… See the full description on the dataset page: https://huggingface.co/datasets/espnet/floras.audioautomatic-speech-recognition10K<n<100K15 likes3.5k downloads2mo agoHugging Face06espnet /yodas_owsmv4🏆 News: Our OWSM v4 paper won the Best Student Paper Award at INTERSPEECH 2025! Dataset Card for YODAS_OWSMv4 Paper: OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning (Best Student Paper at INTERSPEECH 2025) Authors: Yifan Peng, Muhammad Shakeel, Yui Sudo, William Chen, Jinchuan Tian, Chyi-Jiunn Lin, Shinji Watanabe Data Cleaning Scripts: ESPnet Model Demo: Gradio Dataset Description Open Whisper-style Speech Model (OWSM)is the first… See the full description on the dataset page: https://huggingface.co/datasets/espnet/yodas_owsmv4.imageautomatic-speech-recognitionn<1K18 likes3.5k downloads1y agoHugging Face07espnet /Bagpiper_PreTrain_Data Bagpiper Pretraining Data Bagpiper Pretraining Data is the public rich-captioned audio snapshot associated with Bagpiper, an open-ended audio language model that learns bidirectional mappings between audio and comprehensive text descriptions across speech, music, environmental sound, and mixtures. The en metadata describes the primary rich-caption language. Source audio can contain speech or singing in other languages; it is not an English-only audio guarantee. The repository… See the full description on the dataset page: https://huggingface.co/datasets/espnet/Bagpiper_PreTrain_Data.tabularautomatic-speech-recognition10K<n<100K0 likes2.9k downloads2mo agoHugging Face08espnet /Bagpiper_TTS_SFT_Data Bagpiper-TTS SFT Data Release status: the validated Parquet release is being uploaded. The homepage and metadata may appear before every large shard is committed. Bagpiper-TTS SFT Data supports Bagpiper-TTS, a universal speech-synthesis model that interprets free-form natural-language requests, plans the requested delivery, produces a rich textual caption, and synthesizes the target audio. The release is organized into the six applications used by the paper:… See the full description on the dataset page: https://huggingface.co/datasets/espnet/Bagpiper_TTS_SFT_Data.audiotext-to-speech100K<n<1M0 likes2.7k downloads2mo agoHugging Face09Stanwang1210 /raw_tts_emilia_ESPnet_espnet_mls-multi_soundstream_16k0 likes1.6k downloads2y agoHugging Face10espnet /ace-opencpop-segments Citation Information @misc{shi2024singingvoicedatascalingup, title={Singing Voice Data Scaling-up: An Introduction to ACE-Opencpop and ACE-KiSing}, author={Jiatong Shi and Yueqian Lin and Xinyi Bai and Keyi Zhang and Yuning Wu and Yuxun Tang and Yifeng Yu and Qin Jin and Shinji Watanabe}, year={2024}, eprint={2401.17619}, archivePrefix={arXiv}, primaryClass={cs.SD}, url={https://arxiv.org/abs/2401.17619}, } audiotext-to-audio100K<n<1M8 likes1.1k downloads2y agoHugging Face11Stanwang1210 /raw_tts_esc_ESPnet_espnet_mls-audioset_soundstream_16ktextn<1K0 likes932 downloads2y agoHugging Face12espnet /mms_ulab_v2MMS ulab v2 is a a massively multilingual speech dataset that contains 8900 hours of unlabeled speech across 4023 languages. In total, it contains 189 language families. It can be used for language identification, spoken language modelling, or speech representation learning. MMS ulab v2 is a reproduced and extended version of the MMS ulab dataset originally proposed in Scaling Speech Technology to 1000+ Languages, covering more languages and containing more data. This dataset includes the raw… See the full description on the dataset page: https://huggingface.co/datasets/espnet/mms_ulab_v2.audioaudio-to-audio10K<n<100K27 likes918 downloads2y agoHugging Face13eligrayy /OE-LoL-Esports-Dataset0 likes866 downloads3y agoHugging Face14Stanwang1210 /speechlm_bpe_tts_mls_mix_train_valle_espnet_mls-multi_soundstream_16k0 likes768 downloads2y agoHugging Face15espressocheese /LichessGamestabular1B<n<10B1 likes735 downloads4mo agoHugging Face16Stanwang1210 /raw_tts_esc_ESPnet_espnet_mls-multi_soundstream_16ktextn<1K0 likes697 downloads2y agoHugging Face17gptilt /lol-esports-matches GPTilt: League of Legends Esports Matches This dataset is part of the GPTilt open-source initiative, aimed at democratizing access to high-quality LoL data for research and analysis, fostering public exploration, and advancing the community's understanding of League of Legends through data science and AI. It provides a clean, canonical record of the competitive matches and games of professional League of Legends. By using this dataset, users accept full responsibility for any… See the full description on the dataset page: https://huggingface.co/datasets/gptilt/lol-esports-matches.text100K<n<1M0 likes617 downloads9d agoHugging Face18espejelomar /code_search_net_python_10000_examplestext10K<n<100K14 likes466 downloads5y agoHugging Face19ESpeech /ESpeech-webinars2 Webinar Audio Dataset Dataset Description This dataset contains 850 hours processed webinar audio segments with corresponding metadata. Each audio file represents a segment extracted from webinar recordings, processed at 44.1kHz sample rate. Dataset Summary Language: Russian Task: TTS, ASR, Quality Asessment Audio format: MP3, 44.1kHz sample rate Structure: Segmented audio files with JSON metadata Dataset Structure Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/ESpeech/ESpeech-webinars2.audiotext-to-speech100K<n<1M8 likes413 downloads1y agoHugging Face20Stanwang1210 /raw_bpe_tts_mls_ESPnet_espnet_mls-english_soundstream_16k0 likes403 downloads2y agoHugging Face21espnet /ace-kising-segments Citation Information @misc{shi2024singingvoicedatascalingup, title={Singing Voice Data Scaling-up: An Introduction to ACE-Opencpop and ACE-KiSing}, author={Jiatong Shi and Yueqian Lin and Xinyi Bai and Keyi Zhang and Yuning Wu and Yuxun Tang and Yifeng Yu and Qin Jin and Shinji Watanabe}, year={2024}, eprint={2401.17619}, archivePrefix={arXiv}, primaryClass={cs.SD}, url={https://arxiv.org/abs/2401.17619}, } audiotext-to-audio10K<n<100K7 likes361 downloads2y agoHugging Face22Stanwang1210 /speechlm_bpe_tts_mls_mix_train_valle_espnet_mls-audioset_soundstream_16k0 likes361 downloads2y agoHugging Face23Stanwang1210 /speechlm_bpe_tts_mls_mix_train_valle_espnet_mls-english_soundstream_16k0 likes343 downloads2y agoHugging Face24EsportsBench /EsportsBench EsportsBench: A Collection of Datasets for Benchmarking Rating Systems in Esports EsportsBench is a collection of 20 esports competition datasets. Each row of each dataset represents a match played between either two players or two teams in a professional video game tournament. The goal of the datasets is to provide a resource for comparison and development of rating systems used to predict the results of esports matches based on past results. Date is complete up to 2026-03-31.… See the full description on the dataset page: https://huggingface.co/datasets/EsportsBench/EsportsBench.text1M<n<10M4 likes338 downloads1mo agoHugging Face25Stanwang1210 /raw_tts_esc_ESPnet_espnet_mls-english_soundstream_16ktextn<1K0 likes298 downloads2y agoHugging Face26espnet /DSUChallenge2024 The Interspeech 2024 Challenge on Speech Processing Using Discrete Units Paper: https://www.isca-archive.org/interspeech_2024/chang24b_interspeech.html Arxiv: https://arxiv.org/abs/2406.07725 Challenge details: https://www.wavlab.org/activities/2024/Interspeech2024-Discrete-Speech-Unit-Challenge/ To cite: @inproceedings{chang24b_interspeech, title = {The Interspeech 2024 Challenge on Speech Processing Using Discrete Units}, author = {Xuankai Chang and Jiatong Shi and… See the full description on the dataset page: https://huggingface.co/datasets/espnet/DSUChallenge2024.audio100K<n<1M1 likes273 downloads2y agoHugging Face27gptilt /lol-esports-entities GPTilt: League of Legends Esports Directory This dataset is part of the GPTilt open-source initiative, aimed at democratizing access to high-quality LoL data for research and analysis, fostering public exploration, and advancing the community's understanding of League of Legends through data science and AI. It provides a clean, canonical reference for the people and organizations of competitive League of Legends. By using this dataset, users accept full responsibility for any… See the full description on the dataset page: https://huggingface.co/datasets/gptilt/lol-esports-entities.text100K<n<1M1 likes272 downloads9d agoHugging Face28instinct-org /espeech_podcasts_chunked_tts_traingated espeech_podcasts_chunked_tts_train This is a gated Russian TTS training dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: ru (Russian) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Prepared for TTS training… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/espeech_podcasts_chunked_tts_train.audiotext-to-speech0 likes253 downloads4mo agoHugging Face29espnet /wikitonguesThe WikiTongues speech corpus is a collection of conversational audio across 700+ languages. It can be used for spoken language modelling or speech representation learning. This dataset includes the raw unsegmented audio in a 16kHz single channel format. Each clip is usually 2-10 minutes long, and contains one or more speakers conversing in their language(s). Sometimes, a speaker may switch languages within a single clip. The total dataset size is around 70 hours. The current version of the… See the full description on the dataset page: https://huggingface.co/datasets/espnet/wikitongues.audioaudio-to-audion<1K4 likes238 downloads2y agoHugging Face30espnet /ml_superb_hfaudio100K<n<1M7 likes231 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.