CoolFace
16 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01espnet /Bagpiper_SFT_Data Bagpiper SFT Data Release status: the validated Parquet release is being uploaded. The homepage and metadata may appear before every large shard is committed. Bagpiper SFT Data is the supervised fine-tuning corpus for Bagpiper, an open-ended audio language model that understands and generates speech, music, environmental sound, and their mixtures through rich textual captions and planning. The public release has exactly two configurations: Configuration Direction… See the full description on the dataset page: https://huggingface.co/datasets/espnet/Bagpiper_SFT_Data.audioaudio-classification1M<n<10M1 likes6.3k downloads2mo agoHugging Face02espnet /floras FLORAS FLORAS is a 50-language benchmark For LOng-form Recognition And Summarization of spoken language. The goal of FLORAS is to create a more realistic benchmarking environment for speech recognition, translation, and summarization models. Unlike typical academic benchmarks like LibriSpeech and FLEURS that uses pre-segmented single-speaker read-speech, FLORAS tests the capabilities of models on raw long-form conversational audio, which can have one or many speakers. To… See the full description on the dataset page: https://huggingface.co/datasets/espnet/floras.audioautomatic-speech-recognition10K<n<100K15 likes3.8k downloads2mo agoHugging Face03espnet /Bagpiper_TTS_SFT_Data Bagpiper-TTS SFT Data Release status: the validated Parquet release is being uploaded. The homepage and metadata may appear before every large shard is committed. Bagpiper-TTS SFT Data supports Bagpiper-TTS, a universal speech-synthesis model that interprets free-form natural-language requests, plans the requested delivery, produces a rich textual caption, and synthesizes the target audio. The release is organized into the six applications used by the paper:… See the full description on the dataset page: https://huggingface.co/datasets/espnet/Bagpiper_TTS_SFT_Data.audiotext-to-speech100K<n<1M0 likes3k downloads2mo agoHugging Face04espnet /Bagpiper_PreTrain_Data Bagpiper Pretraining Data Bagpiper Pretraining Data is the public rich-captioned audio snapshot associated with Bagpiper, an open-ended audio language model that learns bidirectional mappings between audio and comprehensive text descriptions across speech, music, environmental sound, and mixtures. The en metadata describes the primary rich-caption language. Source audio can contain speech or singing in other languages; it is not an English-only audio guarantee. The repository… See the full description on the dataset page: https://huggingface.co/datasets/espnet/Bagpiper_PreTrain_Data.tabularautomatic-speech-recognition10K<n<100K0 likes2.6k downloads2mo agoHugging Face05espnet /ace-opencpop-segments Citation Information @misc{shi2024singingvoicedatascalingup, title={Singing Voice Data Scaling-up: An Introduction to ACE-Opencpop and ACE-KiSing}, author={Jiatong Shi and Yueqian Lin and Xinyi Bai and Keyi Zhang and Yuning Wu and Yuxun Tang and Yifeng Yu and Qin Jin and Shinji Watanabe}, year={2024}, eprint={2401.17619}, archivePrefix={arXiv}, primaryClass={cs.SD}, url={https://arxiv.org/abs/2401.17619}, } audiotext-to-audio100K<n<1M8 likes1k downloads2y agoHugging Face06espnet /mms_ulab_v2MMS ulab v2 is a a massively multilingual speech dataset that contains 8900 hours of unlabeled speech across 4023 languages. In total, it contains 189 language families. It can be used for language identification, spoken language modelling, or speech representation learning. MMS ulab v2 is a reproduced and extended version of the MMS ulab dataset originally proposed in Scaling Speech Technology to 1000+ Languages, covering more languages and containing more data. This dataset includes the raw… See the full description on the dataset page: https://huggingface.co/datasets/espnet/mms_ulab_v2.audioaudio-to-audio10K<n<100K27 likes992 downloads2y agoHugging Face07espnet /ace-kising-segments Citation Information @misc{shi2024singingvoicedatascalingup, title={Singing Voice Data Scaling-up: An Introduction to ACE-Opencpop and ACE-KiSing}, author={Jiatong Shi and Yueqian Lin and Xinyi Bai and Keyi Zhang and Yuning Wu and Yuxun Tang and Yifeng Yu and Qin Jin and Shinji Watanabe}, year={2024}, eprint={2401.17619}, archivePrefix={arXiv}, primaryClass={cs.SD}, url={https://arxiv.org/abs/2401.17619}, } audiotext-to-audio10K<n<100K7 likes421 downloads2y agoHugging Face08espnet /DSUChallenge2024 The Interspeech 2024 Challenge on Speech Processing Using Discrete Units Paper: https://www.isca-archive.org/interspeech_2024/chang24b_interspeech.html Arxiv: https://arxiv.org/abs/2406.07725 Challenge details: https://www.wavlab.org/activities/2024/Interspeech2024-Discrete-Speech-Unit-Challenge/ To cite: @inproceedings{chang24b_interspeech, title = {The Interspeech 2024 Challenge on Speech Processing Using Discrete Units}, author = {Xuankai Chang and Jiatong Shi and… See the full description on the dataset page: https://huggingface.co/datasets/espnet/DSUChallenge2024.audio100K<n<1M1 likes326 downloads2y agoHugging Face09espnet /wikitonguesThe WikiTongues speech corpus is a collection of conversational audio across 700+ languages. It can be used for spoken language modelling or speech representation learning. This dataset includes the raw unsegmented audio in a 16kHz single channel format. Each clip is usually 2-10 minutes long, and contains one or more speakers conversing in their language(s). Sometimes, a speaker may switch languages within a single clip. The total dataset size is around 70 hours. The current version of the… See the full description on the dataset page: https://huggingface.co/datasets/espnet/wikitongues.audioaudio-to-audion<1K4 likes258 downloads2y agoHugging Face10espnet /ml_superb_hfaudio100K<n<1M7 likes236 downloads2y agoHugging Face11espnet /jesus_dramasJesus Dramas is a collection of religious audio dramas across 430 languages. In total, there is around 640 hours of audio. It can be used for language identification, spoken language modelling, or speech representation learning. This dataset includes the raw unsegmented audio in a 16kHz single channel format. Each audio drama can have multiple speakers, for both male and female voices. It can be segmented into utterances with a voice activity detection (VAD) model such as this one. The… See the full description on the dataset page: https://huggingface.co/datasets/espnet/jesus_dramas.audioaudio-to-audion<1K4 likes138 downloads2y agoHugging Face12espnet /long-yodas-unsegmentedaudio1K<n<10K0 likes120 downloads2y agoHugging Face13espnet /kising_score_segmentstabularn<1K0 likes29 downloads1y agoHugging Face14espnet /long-yodas-segmentedaudio100K<n<1M2 likes12 downloads2y agoHugging Face15chiyuanhsiao /espnet_prob_swb_16k_faketextn<1K0 likes10 downloads9mo agoHugging Face16chiyuanhsiao /espnet_prob_swb_16ktext1K<n<10K0 likes8 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.