CoolFace
28 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01XRXRX /X-Voice-Dataset-Train X-Voice Training Dataset Overview The X-Voice training dataset is a large-scale multilingual speech corpus curated for high-performance speech models. It provides a robust foundation for cross-lingual phonetic and prosodic modeling. Also the train set of X-Voice Model. Core Statistics Total Speech Duration: 420K hours 30 languages European: bg (Bulgarian), cs (Czech), da (Danish), de (German), el (Greek), en (English), es (Spanish), et (Estonian), fi… See the full description on the dataset page: https://huggingface.co/datasets/XRXRX/X-Voice-Dataset-Train.audiotext-to-speech10M<n<100M11 likes4.5k downloads5mo agoHugging Face02Yehor /audiobooks-xxlaudioautomatic-speech-recognition10M<n<100M1 likes470 downloads11mo agoHugging Face03xmodar /commonvoice-12.0-arabic-voice-converted Dataset Card for Voice Converted Arabic Common Voice 12.0 This dataset is derived from the Common Voice Arabic Corpus 12.0 and includes automatically diacritized transcriptions and phoneme representations for the original augmented audio data. The recordings feature Arabic text read aloud by users, where the text was initially undiacritized, allowing for potential reading errors. The diacritization and phonemes were generated automatically, resulting in a dataset that is valuable… See the full description on the dataset page: https://huggingface.co/datasets/xmodar/commonvoice-12.0-arabic-voice-converted.audioautomatic-speech-recognition100K<n<1M8 likes360 downloads2y agoHugging Face04SALT-Research /DeepDialogue-xtts DeepDialogue-xtts DeepDialogue-xtts is a large-scale multimodal dataset containing 40,150 high-quality multi-turn dialogues spanning 41 domains and incorporating 20 distinct emotions with coherent emotional progressions. This repository contains the XTTS-v2 variant of the dataset, where speech is generated using XTTS-v2 with explicit emotional conditioning. 🚨 Important This dataset is large (~180GB) due to the inclusion of high-quality audio files. When cloning the… See the full description on the dataset page: https://huggingface.co/datasets/SALT-Research/DeepDialogue-xtts.audioaudio-classification100K<n<1M8 likes250 downloads1y agoHugging Face05CLEAR-Global /Hausa-Synthetic-ASR-Dataset-XTTSgatedSynthetic Hausa ASR dataset generated using a fine-tuned version of the XTTS-v2 model. Sample rate: 24kHz. Total duration: 574 hours. audioautomatic-speech-recognition100K<n<1M1 likes149 downloads1y agoHugging Face06ilyes25 /wjbmattingly_xhosa_merged_audio Xhosa Merged Audio This dataset was cultivated from Beijuka/xhosa_parakeet_50hr. This dataset orginally came from NCHLT isiXhosa Speech Corpus (see below). The original corpus contained audio and transcription in 3-5 word segments. This meant that the majority of the dataset was ~5 seconds long. Whisper can receive an input of 30 seconds. This meant that the dataset required substantial padding. To reduce the amount of padding, the audio segments were merged together sequentially… See the full description on the dataset page: https://huggingface.co/datasets/ilyes25/wjbmattingly_xhosa_merged_audio.audioautomatic-speech-recognition1K<n<10K0 likes134 downloads1y agoHugging Face07syzym /xbmu_amdo31 Dataset Card for [XBMU-AMDO31] Dataset Summary XBMU-AMDO31 dataset is a speech recognition corpus of Amdo Tibetan dialect. The open source corpus contains 31 hours of speech data and resources related to build speech recognition systems, including transcribed texts and a Tibetan pronunciation dictionary. Supported Tasks and Leaderboards automatic-speech-recognition: The dataset can be used to train a model for Amdo Tibetan Automatic Speech Recognition (ASR). It… See the full description on the dataset page: https://huggingface.co/datasets/syzym/xbmu_amdo31.textautomatic-speech-recognition10K<n<100K4 likes130 downloads4y agoHugging Face08DOVEXAI /XIMA-2122 DOVEXAI Nigerian Languages Speech & Translation Dataset (XIMA-2122) Curated by DOVEXAI LTD (Nigeria) · dovexai.africa · dovexai.io A provenance-tracked dataset of Nigerian-language voice recordings paired with human transcriptions and translations, built for evaluation and supervised fine-tuning (SFT) of speech and translation models. Every recording is cryptographically receipted, fully anonymized, and quality-gated before release. Dataset snapshot Feature… See the full description on the dataset page: https://huggingface.co/datasets/DOVEXAI/XIMA-2122.audioautomatic-speech-recognitionn<1K1 likes114 downloads14d agoHugging Face09BrunoHays /english-x-code-switching Synthetic English Code-Switching Evaluation Set This dataset contains synthetic long-form English code-switching audio samples built from ML-SUPERB hybrid data. Each mixed sample combines English with exactly one additional language. Durations are randomly drawn between 5 and 15 minutes, and each sample contains one or two code switches. The random seed is stored per row. Each selected utterance chunk is RMS-normalized to -20.0 dBFS before concatenation, with peak limiting at 0.99.… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/english-x-code-switching.audioautomatic-speech-recognitionn<1K0 likes111 downloads5mo agoHugging Face10BrunoHays /english-en-x-code-switching-main-lang English EN-X Code-Switching Main-Language This dataset contains synthetic English-plus-one-language code-switching samples built from FLEURS. Source data is google/fleurs at revision refs/convert/parquet, split test, resampled to 16000 Hz. The generator uses seed 42 and creates 50 mixed samples. Each mixed sample contains English and exactly one of Spanish, Portuguese, French, German, or Italian, sampled uniformly. Each selected utterance is RMS-normalized to -20.0 dBFS before… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/english-en-x-code-switching-main-lang.audioautomatic-speech-recognitionn<1K0 likes96 downloads5mo agoHugging Face11xacer /librivox-tracksA dataset of all audio files uploaded to LibriVox before 26th September 2023. Forked from https://huggingface.co/datasets/pykeio/librivox-tracks Changes: Used archive.org metadata API to annotate rows with "duration" column tabulartext-to-speech100K<n<1M2 likes95 downloads2y agoHugging Face12wjbmattingly /xhosa_merged_audio Xhosa Merged Audio This dataset was cultivated from Beijuka/xhosa_parakeet_50hr. This dataset orginally came from NCHLT isiXhosa Speech Corpus (see below). The original corpus contained audio and transcription in 3-5 word segments. This meant that the majority of the dataset was ~5 seconds long. Whisper can receive an input of 30 seconds. This meant that the dataset required substantial padding. To reduce the amount of padding, the audio segments were merged together sequentially… See the full description on the dataset page: https://huggingface.co/datasets/wjbmattingly/xhosa_merged_audio.audioautomatic-speech-recognition1K<n<10K2 likes93 downloads2y agoHugging Face13simpra /fleurs_xho FLEURS -- isiXhosa (xh_za) Re-mirrored from google/fleurs, config xh_za. n-way parallel read speech built on FLORES-101 text -- an evaluation-sized corpus (~19 h), not training scale, but the de facto African-language ASR/TTS benchmark (used in Whisper, MMS, SeamlessM4T, USM papers). Licence CC BY 4.0 -- inherited unchanged from the source. What changed from the source audio peak-normalized per clip (see below); sample rate and encoding otherwise… See the full description on the dataset page: https://huggingface.co/datasets/simpra/fleurs_xho.textautomatic-speech-recognition1K<n<10K0 likes80 downloads11d agoHugging Face14XedriX /malagasy-nwt-bible Malagasy NWT Dataset Dataset from the Malagasy New World Translation Bible (JW.org 2021). Format audio — 16 kHz mono WAV text — clean Malagasy transcript, digits converted to Malagasy words Speaker Information This dataset contains multiple speakers — several readers who each narrate different books or chapters of the Bible. Automatic speaker diarization and clustering was attempted using pyannote/speaker-diarization-3.1, but reliable speaker identity… See the full description on the dataset page: https://huggingface.co/datasets/XedriX/malagasy-nwt-bible.audioautomatic-speech-recognition10K<n<100K0 likes69 downloads6mo agoHugging Face15ia-espirita /pinga-fogo-chico-xavier 🎙️ Pinga-Fogo com Chico Xavier — TV Tupi, 1971 As duas entrevistas históricas do médium Chico Xavier, transmitidas ao vivo pela TV Tupi em 1971, transcritas e estruturadas em turnos de fala com timestamp. 345 turnos (115 deles respostas do próprio Chico Xavier), a partir de 6 horas de áudio — o registro mais extenso do médium falando de improviso, sem edição, diante de um painel de jornalistas. Arquivos Arquivo Programa Turnos Respostas do Chico… See the full description on the dataset page: https://huggingface.co/datasets/ia-espirita/pinga-fogo-chico-xavier.tabularquestion-answeringn<1K1 likes63 downloads1mo agoHugging Face16XRXRX /X-Voice-TestsetX-Voice Multilingual Test Set High-Fidelity Test Set for Multilingual Text-to-Speech across 30 Languages This test set is built as part of the research: X-Voice: One Speaker, 30+ Languages with Zero-Shot Voice Cloning, serving as the evaluation benchmark for our model. Dataset Summary 30 languages European: bg (Bulgarian), cs (Czech), da (Danish), de (German), el (Greek), en (English), es (Spanish), et (Estonian), fi (Finnish), fr (French), hr (Croatian), hu (Hungarian), it… See the full description on the dataset page: https://huggingface.co/datasets/XRXRX/X-Voice-Testset.audiotext-to-speech10K<n<100K4 likes60 downloads5mo agoHugging Face17xezpeleta /mozilla-common-voice-20-eugated Mozilla Common Voice Basque Dataset v20.0 This is the Basque portion of the Mozilla Common Voice dataset version 20.0. audioautomatic-speech-recognition100K<n<1M0 likes52 downloads2y agoHugging Face18Itbanque /ScreenTalk-XS 🎬 ScreenTalk-XS: Sample Speech Dataset from Screen Content 🖥️ 📢 What is ScreenTalk-XS? ScreenTalk-XS is a high-quality transcribed speech dataset containing 10k speech samples from diverse screen content.It is designed for automatic speech recognition (ASR), natural language processing (NLP), and conversational AI research. ✅ This dataset is freely available for research and educational use.🔹 If you need a larger dataset with more diverse speech samples… See the full description on the dataset page: https://huggingface.co/datasets/Itbanque/ScreenTalk-XS.audioautomatic-speech-recognition10K<n<100K2 likes46 downloads1y agoHugging Face19danielshaps /nchlt_speech_xho NCHLT Speech Corpus -- isiXhosa This is the isiXhosa language part of the NCHLT Speech Corpus of the South African languages. Language code (ISO 639): xho URI: https://hdl.handle.net/20.500.12185/279 Licence: Creative Commons Attribution 3.0 Unported License (CC BY 3.0): http://creativecommons.org/licenses/by/3.0/legalcode Attribution: The Department of Arts and Culture of the government of the Republic of South Africa (DAC), Council for Scientific and… See the full description on the dataset page: https://huggingface.co/datasets/danielshaps/nchlt_speech_xho.audioautomatic-speech-recognition10K<n<100K0 likes42 downloads2y agoHugging Face20Max5ive /xhosa_merged_audio Xhosa Merged Audio This dataset was cultivated from Beijuka/xhosa_parakeet_50hr. This dataset orginally came from NCHLT isiXhosa Speech Corpus (see below). The original corpus contained audio and transcription in 3-5 word segments. This meant that the majority of the dataset was ~5 seconds long. Whisper can receive an input of 30 seconds. This meant that the dataset required substantial padding. To reduce the amount of padding, the audio segments were merged together sequentially… See the full description on the dataset page: https://huggingface.co/datasets/Max5ive/xhosa_merged_audio.audioautomatic-speech-recognition1K<n<10K0 likes42 downloads6mo agoHugging Face21BrunoHays /english-x-code-switching-samples Synthetic English Code-Switching Evaluation Set Samples This dataset contains the individual normalized utterance chunks used to build the paired mixed dataset. Each mixed sample combines English with exactly one additional language. Durations are randomly drawn between 5 and 15 minutes, and each sample contains one or two code switches. The random seed is stored per row. Each selected utterance chunk is RMS-normalized to -20.0 dBFS before concatenation, with peak limiting at 0.99.… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/english-x-code-switching-samples.audioautomatic-speech-recognition10K<n<100K0 likes41 downloads5mo agoHugging Face22xiaofff /omnievalkit-data-test OmniEvalKit Evaluation Datasets Evaluation datasets for OmniEvalKit, a comprehensive evaluation framework for omni-modal (audio + video + image + text) models. Overview Total subsets: 89 Total samples: 353,610 Total size: 352.3 GB (Parquet with embedded audio/image, no video) Subsets requiring video download: 42 Note: Video files are NOT embedded in the Parquet files due to size constraints. Usage from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/xiaofff/omnievalkit-data-test.audioaudio-classification10K<n<100K0 likes36 downloads6mo agoHugging Face23BrunoHays /english-en-x-code-switching-main-lang-samples-merged English EN-X Code-Switching Main-Language Merged Samples This dataset contains contiguous same-language segments from the paired mixed dataset. Source data is google/fleurs at revision refs/convert/parquet, split test, resampled to 16000 Hz. The generator uses seed 42 and creates 50 mixed samples. Each mixed sample contains English and exactly one of Spanish, Portuguese, French, German, or Italian, sampled uniformly. Each selected utterance is RMS-normalized to -20.0 dBFS before… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/english-en-x-code-switching-main-lang-samples-merged.audioautomatic-speech-recognitionn<1K0 likes33 downloads5mo agoHugging Face24Max5ive /nchlt_speech_xhosa NCHLT Speech Corpus -- isiXhosa This is the isiXhosa language part of the NCHLT Speech Corpus of the South African languages. Language code (ISO 639): xho URI: https://hdl.handle.net/20.500.12185/279 Licence: Creative Commons Attribution 3.0 Unported License (CC BY 3.0): http://creativecommons.org/licenses/by/3.0/legalcode Attribution: The Department of Arts and Culture of the government of the Republic of South Africa (DAC), Council for Scientific and… See the full description on the dataset page: https://huggingface.co/datasets/Max5ive/nchlt_speech_xhosa.audioautomatic-speech-recognition10K<n<100K0 likes18 downloads6mo agoHugging Face25BrunoHays /english-en-x-code-switching-main-lang-samples English EN-X Code-Switching Main-Language Samples This dataset contains the individual full FLEURS utterance chunks used to build the paired mixed dataset. Source data is google/fleurs at revision refs/convert/parquet, split test, resampled to 16000 Hz. The generator uses seed 42 and creates 50 mixed samples. Each mixed sample contains English and exactly one of Spanish, Portuguese, French, German, or Italian, sampled uniformly. Each selected utterance is RMS-normalized to -20.0… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/english-en-x-code-switching-main-lang-samples.audioautomatic-speech-recognition1K<n<10K0 likes11 downloads5mo agoHugging Face26Max5ive /nchlt_speech_xitsonga NCHLT Speech Corpus -- Xitsonga This is the Xitsonga language part of the NCHLT Speech Corpus of the South African languages. Language code (ISO 639): tso URI: https://hdl.handle.net/20.500.12185/277 Licence: Creative Commons Attribution 3.0 Unported License (CC BY 3.0): http://creativecommons.org/licenses/by/3.0/legalcode Attribution: The Department of Arts and Culture of the government of the Republic of South Africa (DAC), Council for Scientific and… See the full description on the dataset page: https://huggingface.co/datasets/Max5ive/nchlt_speech_xitsonga.audioautomatic-speech-recognition10K<n<100K0 likes8 downloads6mo agoHugging Face27jssaluja /youtube_watch_XDw2cw6Fh8ggatedaudioautomatic-speech-recognitionn<1K0 likes6 downloads2mo agoHugging Face28CAS-SIAT-XinHai /AudiPsy-Synthesisgated AudiPsy-Synthesis: Multilingual Emotional Counseling Dialogue Dataset AudiPsy is a multilingual emotional counseling dialogue dataset containing paired speech audio and text transcripts for psychological counseling conversations. The dataset includes synthetic speech generated using modern TTS systems and annotated emotional information. It is designed to support research in: emotional speech understanding mental health dialogue modeling speech-text multimodal learning counseling… See the full description on the dataset page: https://huggingface.co/datasets/CAS-SIAT-XinHai/AudiPsy-Synthesis.audioautomatic-speech-recognition10K<n<100K0 likes4 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.