CoolFace
19 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01espnet /yodas-granary Dataset Card for YODAS-Granary Repository: NeMo-speech-data-processor: Granary Paper: Granary: Speech Recognition and Translation Dataset in 25 European Languages Shared by: ESPnet Dataset Description YODAS-Granary is a curated subset of the larger nvidia/Granary dataset, focusing on high-quality pseudo-labeled speech data for Automatic Speech Recognition (ASR) and Automatic Speech Translation (AST) across 23 European languages. Overview… See the full description on the dataset page: https://huggingface.co/datasets/espnet/yodas-granary.audioautomatic-speech-recognition10M<n<100M33 likes82k downloads1y agoHugging Face02espnet /Bagpiper_SFT_Data Bagpiper SFT Data Release status: the validated Parquet release is being uploaded. The homepage and metadata may appear before every large shard is committed. Bagpiper SFT Data is the supervised fine-tuning corpus for Bagpiper, an open-ended audio language model that understands and generates speech, music, environmental sound, and their mixtures through rich textual captions and planning. The public release has exactly two configurations: Configuration Direction… See the full description on the dataset page: https://huggingface.co/datasets/espnet/Bagpiper_SFT_Data.audioaudio-classification1M<n<10M1 likes6.5k downloads2mo agoHugging Face03espnet /floras FLORAS FLORAS is a 50-language benchmark For LOng-form Recognition And Summarization of spoken language. The goal of FLORAS is to create a more realistic benchmarking environment for speech recognition, translation, and summarization models. Unlike typical academic benchmarks like LibriSpeech and FLEURS that uses pre-segmented single-speaker read-speech, FLORAS tests the capabilities of models on raw long-form conversational audio, which can have one or many speakers. To… See the full description on the dataset page: https://huggingface.co/datasets/espnet/floras.audioautomatic-speech-recognition10K<n<100K15 likes3.7k downloads2mo agoHugging Face04espnet /yodas_owsmv4🏆 News: Our OWSM v4 paper won the Best Student Paper Award at INTERSPEECH 2025! Dataset Card for YODAS_OWSMv4 Paper: OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning (Best Student Paper at INTERSPEECH 2025) Authors: Yifan Peng, Muhammad Shakeel, Yui Sudo, William Chen, Jinchuan Tian, Chyi-Jiunn Lin, Shinji Watanabe Data Cleaning Scripts: ESPnet Model Demo: Gradio Dataset Description Open Whisper-style Speech Model (OWSM)is the first… See the full description on the dataset page: https://huggingface.co/datasets/espnet/yodas_owsmv4.imageautomatic-speech-recognitionn<1K18 likes3.5k downloads1y agoHugging Face05espnet /Bagpiper_PreTrain_Data Bagpiper Pretraining Data Bagpiper Pretraining Data is the public rich-captioned audio snapshot associated with Bagpiper, an open-ended audio language model that learns bidirectional mappings between audio and comprehensive text descriptions across speech, music, environmental sound, and mixtures. The en metadata describes the primary rich-caption language. Source audio can contain speech or singing in other languages; it is not an English-only audio guarantee. The repository… See the full description on the dataset page: https://huggingface.co/datasets/espnet/Bagpiper_PreTrain_Data.tabularautomatic-speech-recognition10K<n<100K0 likes2.4k downloads2mo agoHugging Face06espnet /ace-opencpop-segments Citation Information @misc{shi2024singingvoicedatascalingup, title={Singing Voice Data Scaling-up: An Introduction to ACE-Opencpop and ACE-KiSing}, author={Jiatong Shi and Yueqian Lin and Xinyi Bai and Keyi Zhang and Yuning Wu and Yuxun Tang and Yifeng Yu and Qin Jin and Shinji Watanabe}, year={2024}, eprint={2401.17619}, archivePrefix={arXiv}, primaryClass={cs.SD}, url={https://arxiv.org/abs/2401.17619}, } audiotext-to-audio100K<n<1M8 likes1.2k downloads2y agoHugging Face07ESpeech /ESpeech-webinars2 Webinar Audio Dataset Dataset Description This dataset contains 850 hours processed webinar audio segments with corresponding metadata. Each audio file represents a segment extracted from webinar recordings, processed at 44.1kHz sample rate. Dataset Summary Language: Russian Task: TTS, ASR, Quality Asessment Audio format: MP3, 44.1kHz sample rate Structure: Segmented audio files with JSON metadata Dataset Structure Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/ESpeech/ESpeech-webinars2.audiotext-to-speech100K<n<1M8 likes419 downloads1y agoHugging Face08espnet /ace-kising-segments Citation Information @misc{shi2024singingvoicedatascalingup, title={Singing Voice Data Scaling-up: An Introduction to ACE-Opencpop and ACE-KiSing}, author={Jiatong Shi and Yueqian Lin and Xinyi Bai and Keyi Zhang and Yuning Wu and Yuxun Tang and Yifeng Yu and Qin Jin and Shinji Watanabe}, year={2024}, eprint={2401.17619}, archivePrefix={arXiv}, primaryClass={cs.SD}, url={https://arxiv.org/abs/2401.17619}, } audiotext-to-audio10K<n<100K7 likes396 downloads2y agoHugging Face09espnet /yodas3 YODAS v3 Paper YODAS v3 is a large web-crawled dataset containing over 1.1 million hours of audio that were originally released under a CC-BY-3.0 license. The dataset contains audio in over 100 languages. YODAS v3 can be used for a variety of multi-modal tasks, including Automatic Speech Recognition, Text-to-Speech, and Audio Representation Learning. We crawl a distinct set of videos from the v1 and v2 versions of YODAS, to guarantee that there are no overlaps in the data. For… See the full description on the dataset page: https://huggingface.co/datasets/espnet/yodas3.tabularaudio-to-audio1M<n<10M2 likes361 downloads18h agoHugging Face10lab260 /espeech_balalaika ESpeech datasets (w/o podcasts) Annotated by Balalaika [!IMPORTANT] Official dataset for our INTERSPEECH 2026 paper "A Data-Centric Framework for Addressing Phonetic and Prosodic Challenges in Russian Speech Generative Models" (arXiv:2507.13563). Part of the Balalaika Russian speech data-processing pipeline — code: https://github.com/lab260ru/balalaika. If you use this resource, please cite it. A curated Russian speech dataset for advanced speech generative tasks.… See the full description on the dataset page: https://huggingface.co/datasets/lab260/espeech_balalaika.tabulartext-to-speech100K<n<1M4 likes189 downloads3mo agoHugging Face11ESpeech /ESpeech-podcastsgated Podcasts Audio Dataset Dataset Description This dataset contains 3200 hours of processed audio segments extracted from various podcasts with corresponding metadata. Each audio file represents a segment from podcast episodes, processed at 44.1kHz sample rate. Dataset Summary Language: Russian Total Duration: 3200 hours of speech Task: TTS, ASR, Quality Assessment Audio format: MP3, 44.1kHz sample rate Structure: Segmented audio files with JSON… See the full description on the dataset page: https://huggingface.co/datasets/ESpeech/ESpeech-podcasts.audiotext-to-speech10M<n<100M14 likes115 downloads10mo agoHugging Face12ESpeech /ESpeech-tuchniyzhab Tuchniy Zhab YouTube Audio Dataset Dataset Description This dataset contains 306 hours of processed audio segments extracted from the "Tuchniy Zhab" YouTube channel with corresponding metadata. Each audio file represents a segment from the channel's videos and content, processed at 44.1kHz sample rate. Dataset Summary Language: Russian Task: TTS, ASR, Quality Assessment Audio format: MP3, 44.1kHz sample rate Structure: Segmented audio files with JSON metadata… See the full description on the dataset page: https://huggingface.co/datasets/ESpeech/ESpeech-tuchniyzhab.audiotext-to-speech100K<n<1M2 likes113 downloads1y agoHugging Face13ESpeech /ESpeech-upvote Upvote YouTube Audio Dataset Dataset Description This dataset contains 296 hours of processed audio segments extracted from the "Upvote" YouTube channel with corresponding metadata. Each audio file represents a segment from the channel's videos and content, processed at 44.1kHz sample rate. Dataset Summary Language: Russian Task: TTS, ASR, Quality Assessment Audio format: MP3, 44.1kHz sample rate Structure: Segmented audio files with JSON metadata Source:… See the full description on the dataset page: https://huggingface.co/datasets/ESpeech/ESpeech-upvote.audiotext-to-speech100K<n<1M0 likes107 downloads1y agoHugging Face14ESpeech /ESpeech-buldjat Buldjat YouTube Audio Dataset Dataset Description This dataset contains 54 hours of processed audio segments extracted from the "Buldjat" YouTube channel with corresponding metadata. Each audio file represents a segment from the channel's videos and content, processed at 44.1kHz sample rate. Dataset Summary Language: Russian Task: TTS, ASR, Quality Assessment Audio format: MP3, 44.1kHz sample rate Structure: Segmented audio files with JSON metadata Source:… See the full description on the dataset page: https://huggingface.co/datasets/ESpeech/ESpeech-buldjat.audiotext-to-speech10K<n<100K3 likes90 downloads1y agoHugging Face15ia-espirita /pinga-fogo-chico-xavier 🎙️ Pinga-Fogo com Chico Xavier — TV Tupi, 1971 As duas entrevistas históricas do médium Chico Xavier, transmitidas ao vivo pela TV Tupi em 1971, transcritas e estruturadas em turnos de fala com timestamp. 345 turnos (115 deles respostas do próprio Chico Xavier), a partir de 6 horas de áudio — o registro mais extenso do médium falando de improviso, sem edição, diante de um painel de jornalistas. Arquivos Arquivo Programa Turnos Respostas do Chico… See the full description on the dataset page: https://huggingface.co/datasets/ia-espirita/pinga-fogo-chico-xavier.tabularquestion-answeringn<1K1 likes72 downloads1mo agoHugging Face16instinct-org /espeech_podcasts_chunked_speech_restorised_nfa_alignedgated Forced-Aligned STT Dataset Source dataset: instinct-org/espeech_podcasts_chunked_speech_restorised Aligned dataset: instinct-org/espeech_podcasts_chunked_speech_restorised_nfa_aligned Rows: 2467471 successfully aligned rows This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality control. Source rows that did not produce usable CTM alignments are excluded from the published data. Added columns: nfa_token_alignments: NeMo token/subword CTM spans… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/espeech_podcasts_chunked_speech_restorised_nfa_aligned.tabularautomatic-speech-recognition1M<n<10M0 likes3 downloads4mo agoHugging Face17instinct-org /espeech_podcasts_chunkedgated espeech_podcasts_chunked This is a gated Russian chunked speech dataset from instinct-org. It contains podcast speech segments and transcript text for speech-to-text training, evaluation, and data preparation workflows. Current Audio Availability This repository currently has a mixed audio-byte state: data/train-00000-of-00130.parquet through data/train-00010-of-00130.parquet contain embedded audio bytes. data/train-00011-of-00130.parquet through… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/espeech_podcasts_chunked.audioautomatic-speech-recognition1M<n<10M0 likes2 downloads4mo agoHugging Face18instinct-org /espeech_podcasts_chunked_speech_restorisedgated espeech_podcasts_chunked_speech_restorised This is a gated Russian speech-restorised chunked speech dataset from instinct-org. It contains byte-backed speech audio and transcripts for speech-to-text training, evaluation, forced alignment, and data preparation workflows. STT Alignment Source Use this repository as the current repaired source candidate for Russian podcast STT alignment. The matching base repository preserves the original transcript and segment… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/espeech_podcasts_chunked_speech_restorised.audioautomatic-speech-recognition1M<n<10M0 likes2 downloads4mo agoHugging Face19instinct-org /espeech_podcasts_chunked_nfa_alignedgated Forced-Aligned STT Dataset Source dataset: instinct-org/espeech_podcasts_chunked Aligned dataset: instinct-org/espeech_podcasts_chunked_nfa_aligned Rows: 231121 successfully aligned rows This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality control. Source rows that did not produce usable CTM alignments are excluded from the published data. Added columns: nfa_token_alignments: NeMo token/subword CTM spans nfa_word_alignments: word-level CTM spans… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/espeech_podcasts_chunked_nfa_aligned.tabularautomatic-speech-recognition100K<n<1M0 likes1 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.