CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SparkAudio /voxbox VoxBox This dataset is a curated collection of bilingual speech corpora annotated clean transcriptions and rich metadata incluing age, gender, and emotion. Dataset Structure . ├── audios/ │ └── aishell-3/ # Audio files (organised by sub-corpus) │ └── ... └── metadata/ ├── aishell-3.jsonl ├── casia.jsonl ├── commonvoice_cn.jsonl ├── ... └── wenetspeech4tts.jsonl # JSONL metadata files Each JSONL file corresponds to a… See the full description on the dataset page: https://huggingface.co/datasets/SparkAudio/voxbox.audiotext-to-speech10M<n<100M76 likes40k downloads1y agoHugging Face02LGB666 /CosyVoice2-SparkTTSaudion<1K0 likes1.5k downloads1y agoHugging Face03sirekist98 /tokenise_spanish_datasetDataset tokenizado para TTS en español. Incluye audio, transcripción, emoción, y códigos SNAC. audio100K<n<1M0 likes1.4k downloads1y agoHugging Face04sirekist98 /spanish_Audiosaudio100K<n<1M3 likes1.2k downloads1y agoHugging Face05SpeechAntiSpoofingBenchmarks /J-SPAW_LA J-SPAW (LA track, eval) ⚠️ NON-COMMERCIAL USE ONLY The upstream J-SPAW dataset is released "For non-commercial use only" (see the J-SPAW repository). This packaging inherits that restriction: do not use it for any commercial purpose. It is provided solely for non-commercial academic research and benchmarking. The upstream terms are sparse and do not spell out redistribution; contact the original authors for any use beyond non-commercial research. Benchmark-ready… See the full description on the dataset page: https://huggingface.co/datasets/SpeechAntiSpoofingBenchmarks/J-SPAW_LA.audioaudio-classification1K<n<10K0 likes1k downloads3mo agoHugging Face06zhisheng01 /SpatialAudio SpatialAudio This repo hosts the dataset and models of "BAT: Learning to Reason about Spatial Sounds with Large Language Models" [ICML 2024 bib]. Spatial Audio Dataset (Mono/Binaural/Ambisonics) AudioSet (Anechoic Audio Source) We provide Balanced train and Evaluation set for your convenience. You can download from SpatialAudio. For the Unbalanced train set, please refer to Official AudioSet. Metadata can be downloaded from metadata. AudioSet ├──… See the full description on the dataset page: https://huggingface.co/datasets/zhisheng01/SpatialAudio.audio12 likes914 downloads2y agoHugging Face07shraavb /spanish-slang-stt-data Spanish Regional Speech-to-Text Dataset A multilingual Spanish speech recognition dataset covering 4 regional dialects for fine-tuning Whisper and other ASR models. Dataset Description This dataset contains ~39,000 audio samples with transcriptions across 4 Spanish-speaking regions: Region Samples Description Mexico 17,725 Mexican Spanish including CIEMPIESS corpus Spain 11,360 Castilian Spanish from TEDx and Common Voice Argentina 5,839 Rioplatense Spanish… See the full description on the dataset page: https://huggingface.co/datasets/shraavb/spanish-slang-stt-data.audioautomatic-speech-recognition10K<n<100K0 likes895 downloads8mo agoHugging Face08anuj-inavlabs /kupe-spark-asr-270m-data kupe-spark-asr-270m — data Multilingual ASR corpus for kupe-spark-asr-270m (Gemma-3-270m + Mimi codec). Languages: en (English), hi (Hindi), gu (Gujarati), bn (Bengali), ur (Urdu), mr (Marathi) Configs audio — raw speech resampled to 24 kHz mono (audio/data/shard_*.parquet). mimi — Mimi codebook-0 tokens (12.5 tok/s) + transcripts (mimi/*.parquet). Used for training. Shards are uploaded one-by-one as they are fetched. Resume state lives in… See the full description on the dataset page: https://huggingface.co/datasets/anuj-inavlabs/kupe-spark-asr-270m-data.audio0 likes678 downloads19d agoHugging Face09Spaiche /SPC Dataset Card for "SPC-v2" More Information needed audio10K<n<100K0 likes609 downloads4y agoHugging Face10freds0 /cml_tts_dataset_spanishaudio100K<n<1M3 likes459 downloads2y agoHugging Face11lrauch /spassLicense: Creative Commons Attribution 4.0 International Source Rhoddy Viveros-Muñoz, Pablo Huijse, Victor Poblete, Victor Vargas, Diego Espejo, Matthieu Vernier, Diego Vergara, Jorge Arenas, & Enrique Suárez. (2022). SPASS dataset: A synthetic polyphonic dataset with spatiotemporal labels of sound sources (1.0). Zenodo. https://doi.org/10.5281/zenodo.7484370 audio10K<n<100K0 likes377 downloads1y agoHugging Face12SPARCO-project /benchmark_DEMAND_noise benchmark_DEMAND_noise This dataset is a segmented subset derived from DEMAND: Diverse Environments Multichannel Acoustic Noise Database. It is prepared for the SPARCO noise ablation benchmark. The intended use is to provide fixed 4-second environmental noise segments for: AUROC-based SAE noise-related feature selection binary noise-presence scorer training scorer threshold calibration final held-out benchmark evaluation Source Original source: DEMAND: Diverse… See the full description on the dataset page: https://huggingface.co/datasets/SPARCO-project/benchmark_DEMAND_noise.audioaudio-classification1K<n<10K0 likes367 downloads4mo agoHugging Face13ebellob /voxforge_spanish_enhanced VoxForge Spanish Enhanced (CleanUNet + FlashSR) Dataset Summary This dataset is a processed and enhanced version of the Spanish subset of: VoxForge. Furthermore, as this is a personal project, we give no guarantees that the audio is completely clean from any artifacts or noise the CleanUNet model could not remove. However, we have personally tested the corpus via the fine-tuning of some SOTA speech models and the results have been satisfactory. It has been created to… See the full description on the dataset page: https://huggingface.co/datasets/ebellob/voxforge_spanish_enhanced.audio10K<n<100K1 likes357 downloads5mo agoHugging Face14Yujin6 /evaluate_spatialaudio1K<n<10K0 likes350 downloads4mo agoHugging Face15GianDiego /latam-spanish-speech-orpheus-tts-24khz LATAM Spanish High-Quality Speech Dataset (24kHz - Orpheus TTS Ready) Dataset Description This dataset contains approximately 24 hours of high-quality speech audio in Latin American Spanish, specifically prepared for Text-to-Speech (TTS) applications like OrpheusTTS, which require a 24kHz sampling rate. The audio files are derived from the Crowdsourced high-quality speech datasets made by Google and were obtained via OpenSLR. The original recordings were high-quality… See the full description on the dataset page: https://huggingface.co/datasets/GianDiego/latam-spanish-speech-orpheus-tts-24khz.audiotext-to-speech10K<n<100K16 likes328 downloads1y agoHugging Face16ylacombe /google-chilean-spanish Dataset Card for Tamil Speech Dataset Summary This dataset consists of 7 hours of transcribed high-quality audio of Chilean Spanish sentences recorded by 31 volunteers. The dataset is intended for speech technologies. The data archives were restructured from the original ones from OpenSLR to make it easier to stream. Supported Tasks text-to-speech, text-to-audio: The dataset can be used to train a model for Text-To-Speech (TTS). automatic-speech-recognition… See the full description on the dataset page: https://huggingface.co/datasets/ylacombe/google-chilean-spanish.audiotext-to-speech1K<n<10K24 likes306 downloads3y agoHugging Face17kwatcharasupat /dnr-v3-spaaudio0 likes280 downloads11mo agoHugging Face18ylacombe /google-colombian-spanish Dataset Card for "google-colombian-spanish" More Information needed audio1K<n<10K14 likes224 downloads3y agoHugging Face19AdoCleanCode /spanish_voxpopuli_alignedaudio10K<n<100K0 likes198 downloads8mo agoHugging Face20ciempiess /librivox_spanish Dataset Card for librivox_spanish Dataset Summary Librivox is a non-commercial, non-profit and ad-free project that is dedicated to make all books in the public domain available, for free, in audio format on the internet. According to this, we downloaded 300 titles in Spanish to create the LIBRIVOX SPANISH CORPUS. The LIBRIVOX SPANISH CORPUS has a duration of 73 hours and it is constituted by audio files between 3 and 10 seconds long, manually segmented. Transcription are… See the full description on the dataset page: https://huggingface.co/datasets/ciempiess/librivox_spanish.audioautomatic-speech-recognition10K<n<100K5 likes183 downloads2y agoHugging Face21ylacombe /google-argentinian-spanish Dataset Card for "google-argentinian-spanish" More Information needed audio1K<n<10K19 likes178 downloads3y agoHugging Face22mimbres /spansynth-edit-gallery SpanSynth-Edit Gallery Finished MIDI-guided music edits made with SpanSynth-Edit. Open a work's shared link to compare the original and edited audio, explore its score, and make an editable copy. Each work has its own folder: input.wav and output.wav: the original clip and saved generated audio, at 48 kHz mono. original.mid and edited.mid: the source and edited score, aligned to the clip. context.wav: the normalized audio used by the model, including the history before the… See the full description on the dataset page: https://huggingface.co/datasets/mimbres/spansynth-edit-gallery.audioaudio-to-audion<1K0 likes165 downloads14h agoHugging Face23einarolafsson /spacr-tutorials spaCR tutorial media Narration and 4K video for the spaCR interactive tutorial library, served directly to https://einarolafsson.github.io/spacr/tutorials/. spaCR is a toolkit for microscopy and single-cell analysis of pooled CRISPR screens. This repository holds the media its 40-lesson tutorial player streams; it is not a training dataset. Why it lives here GitHub Pages caps a published site at 1 GB. The full narration set is 2,662 MiB across 54 voices, so the… See the full description on the dataset page: https://huggingface.co/datasets/einarolafsson/spacr-tutorials.audion<1K0 likes146 downloads12d agoHugging Face24Thermostatic /CommonVoice-17.0-Spanishaudio1M<n<10M1 likes140 downloads1y agoHugging Face25BrunoHays /Bangor-Miami-Spanish-English-Corpus Bangor Miami Spanish-English Corpus The Bangor Miami Corpus is a naturalistic Spanish-English code-switching speech dataset collected by Jon Russell Herring at Bangor University. It captures spontaneous bilingual conversations recorded in Miami, Florida, involving proficient Spanish-English bilinguals across multiple speaker groups. Dataset description Total recordings 56 Total duration ~32 h Languages English (en), Spanish (es) Format MP3 audio +… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/Bangor-Miami-Spanish-English-Corpus.audioautomatic-speech-recognitionn<1K0 likes140 downloads4mo agoHugging Face26UniDataPro /spanish-speech-recognition-dataset Spanish Speech Dataset for recognition task Dataset comprises 10 hours of telephone dialogues in Spanish, collected from 10 native speakers across various topics and domains. It is a valuable resource for advancing speech recognition technology. By utilizing this dataset, researchers and developers can advance their understanding and capabilities in automatic speech recognition (ASR) systems, transcribing audio, and natural language processing (NLP). - Get the data The dataset… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/spanish-speech-recognition-dataset.audioautomatic-speech-recognitionn<1K2 likes135 downloads1mo agoHugging Face27rjnieto /spanish-dialectsgated Dataset Card For "spanish_dialects" Dataset Summary This dataset contains 10+ hours of high-quality Spanish audio covering numerous speakers across Spain and Latin America. The following dialects are present in the dataset: Spain, Mexico, Chile, Argentina, Dominican Republic audio1K<n<10K3 likes133 downloads9mo agoHugging Face28ciempiess /voxforge_spanish Dataset Card for voxforge_spanish Dataset Summary VoxForge was set up to collect transcribed speech for use with Free and Open Source Speech Recognition Engines (on Linux, Windows and Mac). They promise they will make available all submitted audio files under the GPL license, and then 'compile' them into acoustic models for use with Open Source speech recognition engines such as CMU Sphinx, ISIP, Julius and HTK. According to this, we downloaded the Spanish recordings of… See the full description on the dataset page: https://huggingface.co/datasets/ciempiess/voxforge_spanish.audioautomatic-speech-recognition10K<n<100K4 likes132 downloads2y agoHugging Face29ittailup /spanish_tokenizedaudio10K<n<100K1 likes121 downloads2y agoHugging Face30Cnam-LMSSC /multilingual_librispeech_spanish_phoneme Multilingual LibriSpeech Spanish Phoneme Dataset Summary This dataset is a curated version of the Spanish subset of Multilingual LibriSpeech (MLS), enriched with a phonetic transcription column (phoneme). The Laboratoire de Mécanique des Structures et des Systèmes Couplés (Cnam-LMSSC) created this version to facilitate research into Spanish acoustic modeling, phoneme recognition, and speech synthesis. It builds upon the high-quality audio derived from LibriVox audiobooks… See the full description on the dataset page: https://huggingface.co/datasets/Cnam-LMSSC/multilingual_librispeech_spanish_phoneme.audioautomatic-speech-recognition100K<n<1M1 likes111 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.