CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01projecte-aina /synthetic_dem Dataset Card for synthetic_dem Dataset Summary The Synthetic DEM Corpus is the result of the first phase of a collaboration between El Colegio de México (COLMEX) and the Barcelona Supercomputing Center (BSC). It all began when COLMEX was looking for a way to have its Diccionario del Español de México (DEM), which can be accessed online, include the option to play each of its words with a Mexican accent through synthetic speech files. On the other hand, BSC is always on… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/synthetic_dem.audioautomatic-speech-recognition100K<n<1M2 likes1.5k downloads1y agoHugging Face02TigreGotico /barranquenho-ipa-dict-synthetic Barranquenho IPA Pronunciation Dictionary The first and only IPA pronunciation dictionary of Barranquenho — the Ibero-Romance contact variety spoken in Barrancos (Baixo Alentejo, Portugal), a mixed system born of centuries of Portuguese–Spanish (Extremaduran / Andalusian) contact on the raia. Every headword is written in the Convenção Ortográfica do Barranquenho (2025) orthography and paired with a broad-phonemic IPA transcription plus Portuguese and Spanish glosses. This… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/barranquenho-ipa-dict-synthetic.texttext-to-speech1K<n<10K0 likes510 downloads2mo agoHugging Face03AfriSpeech /multivoice-synthetic-speech Synthetic Voice Samples · Africa Synthetic speech. No human speaker was recorded for any clip in this dataset. Generated with afrispeech-synth: text from africa-corpus, normalised to a universal orthography with africa-g2p, spoken by Google Gemini's Live API. 17,010 clips · 38.8 hours · 566 languages · 30 voices Every clip is a distinct sentence — no sentence is repeated Each language is read by up to 30 different voices, one sentence per voice ~1.29 hours per voice… See the full description on the dataset page: https://huggingface.co/datasets/AfriSpeech/multivoice-synthetic-speech.audiotext-to-speech10K<n<100K1 likes303 downloads6d agoHugging Face04yuriyvnv /synthetic_transcript_pt Portuguese Speech Dataset with Multiple Training Configurations A comprehensive Portuguese speech dataset offering three distinct training configurations for speech recognition research, each designed for different experimental scenarios and training paradigms. 🎯 Dataset Configurations Overview This dataset provides three carefully curated subsets to enable comprehensive speech recognition research: Configuration Training Data Validation Test Total Samples Use Case… See the full description on the dataset page: https://huggingface.co/datasets/yuriyvnv/synthetic_transcript_pt.audioautomatic-speech-recognition100K<n<1M0 likes277 downloads5mo agoHugging Face05Hani89 /Synthetic-Medical-Speech-Dataset Synthetic Medical Speech Dataset Overview Synthetic Medical Speech Dataset is a synthetic dataset of audio–text pairs designed for developing and evaluating automatic speech recognition (ASR) models in the medical domain.The corpus contains thousands of short audio clips generated from medically relevant text using a text-to-speech (TTS) system.Each clip is paired with its corresponding transcript.Because all content is synthetically produced, the dataset does not contain… See the full description on the dataset page: https://huggingface.co/datasets/Hani89/Synthetic-Medical-Speech-Dataset.audioautomatic-speech-recognition10K<n<100K4 likes248 downloads11mo agoHugging Face06MLRS /masri_synthetic Dataset Card for masri_synthetic Dataset Summary The MASRI-SYNTHETIC is a corpus made out of synthesized speech in Maltese. The text-to-speech (TTS) system utilized to produce the utterances was developed by the Research & Development Department of Crimsonwing p.l.c. The sentences used to create the corpus were extracted from the MLRS Corpus, which is a corpus of written or transcribed Maltese divided into different genres, including: culture, news, academic, religion… See the full description on the dataset page: https://huggingface.co/datasets/MLRS/masri_synthetic.audioautomatic-speech-recognition10K<n<100K2 likes246 downloads2y agoHugging Face07yuriyvnv /synthetic_transcript_nl Dutch Synthetic Speech Transcripts This dataset contains 34,898 synthetic Dutch speech samples generated using GPT-4o-mini for transcript creation and OpenAI's TTS-1 model for speech synthesis. It was designed to augment Automatic Speech Recognition (ASR) training for low-resource scenarios, matching the linguistic distribution of Common Voice 17.0 Dutch. Dataset Description Purpose This dataset addresses the challenge of limited labeled speech data for Dutch… See the full description on the dataset page: https://huggingface.co/datasets/yuriyvnv/synthetic_transcript_nl.audioautomatic-speech-recognition10K<n<100K0 likes172 downloads10mo agoHugging Face08routsourav1729 /synthetic-tts Munni — Hindi/English Synthetic TTS Voice Dataset Single-speaker, code-mixed Hindi (Devanagari) + English speech dataset built for XTTS-v2 fine-tuning. Domain is call-center / customer-support style dialogue (warranty, billing, appointment scheduling), with scripted placeholder phone numbers spoken digit-by-digit. 8,796 clips / 18.16 hours, 24 kHz mono 16-bit PCM Split: train 8,621 / validation 175 (98/2, seed 42) Duration per clip: mean 7.43s, min 3.18s, max 11.00s Single… See the full description on the dataset page: https://huggingface.co/datasets/routsourav1729/synthetic-tts.texttext-to-speech1K<n<10K0 likes168 downloads20d agoHugging Face09abdo1819 /arabic-english-code-switching-synthetic-asr Synthetic Arabic-English Code-Switched Speech for ASR This dataset contains synthetic speech generated for Egyptian Arabic-English code-switched automatic speech recognition. It is published separately from the human review annotations so the human audio remains in its upstream Hugging Face repository. Configurations Configuration Train Test Publication status synthetic 8,655 962 Contains 5,673 ArE-CSTD-derived texts; noncommercial/share-alike terms… See the full description on the dataset page: https://huggingface.co/datasets/abdo1819/arabic-english-code-switching-synthetic-asr.audioautomatic-speech-recognition1K<n<10K0 likes164 downloads2mo agoHugging Face10CLEAR-Global /Hausa-Synthetic-ASR-Dataset-XTTSgatedSynthetic Hausa ASR dataset generated using a fine-tuned version of the XTTS-v2 model. Sample rate: 24kHz. Total duration: 574 hours. audioautomatic-speech-recognition100K<n<1M1 likes146 downloads1y agoHugging Face11Wi-Fi /korean-full-duplex-synthetic-dataset-preview Korean Full-Duplex Synthetic Dataset Preview Overview Public preview of a Korean full-duplex synthetic speech dataset. This repository contains 100 conversations sampled from a corpus of 89,273 conversations (2,000.5 hours); it does not publish the full corpus audio. Preview contents 100 conversation WAV files data/representative.jsonl 24 kHz, mono, 16-bit PCM Events: normal, barge_in, backchannel, cutoff_by_user Annotation format… See the full description on the dataset page: https://huggingface.co/datasets/Wi-Fi/korean-full-duplex-synthetic-dataset-preview.audioautomatic-speech-recognitionn<1K1 likes138 downloads1mo agoHugging Face12Anilosan15 /Synthetic_Turkish_TTS_Data Synthetic Turkish TTS Data This dataset was created by generating synthetic Turkish text across multiple speech scenarios. The text was produced in the following domains: finance_master, cs_master, parcel_delivery, ecommerce, telecom, isp_support, technical_support, subscription, insurance, health_appointments, public_services, education_registration, and daily_speech. These synthetic texts were then synthesized with a high-quality Turkish TTS model. The dataset is intended to be… See the full description on the dataset page: https://huggingface.co/datasets/Anilosan15/Synthetic_Turkish_TTS_Data.audiotext-to-speech10K<n<100K6 likes116 downloads5mo agoHugging Face13milanakdj /nepali-tts-synthetic-v2gated Nepali TTS Synthetic v2 383,298 synthetic Nepali (ne) speech/text pairs, 24 kHz mono 16-bit WAV embedded as-is (no re-encode, no resampling). Generated by the synthetic_pipeline in milanakdj/TTS_training: Edge TTS synthesis → optional voice conversion against a pool of 600 real multi-speaker reference clips → ASR-based QC gate on character error rate. Read this before training on it Only 48% of rows are voice-converted. Each row carries a kept field recording… See the full description on the dataset page: https://huggingface.co/datasets/milanakdj/nepali-tts-synthetic-v2.audiotext-to-speech100K<n<1M0 likes116 downloads21d agoHugging Face14uncleMehrzad /synthetic-speaker-diarization-dataset-fa-large-3000audioaudio-classification1K<n<10K3 likes108 downloads1y agoHugging Face15DatarrX /burmese-synthetic-speech-corpus Burmese Synthetic Speech Corpus (DatarrX/burmese-synthetic-speech-corpus) Overview The Burmese Synthetic Speech Corpus is a high-fidelity, manually curated audio dataset specifically designed to advance Text-to-Speech (TTS) systems, speech recognition, and other audio-driven Machine Learning tasks for the Burmese (Myanmar) language. Created by DatarrX, this dataset bridges the gap in low-resource speech technologies by providing highly natural, native-sounding… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/burmese-synthetic-speech-corpus.audiotext-to-speech1K<n<10K7 likes107 downloads4mo agoHugging Face16tsdocode /open-vi-dialog-synthetic-100h OpenDialog Vietnamese Synthetic Dialogue 100h Synthetic Vietnamese two-speaker dialogue for ZipVoice-Dialog experiments. 12,000 chunks 30 seconds per chunk 100.0 hours total Each item contains S1/S2 speaker labels, turn timings, target text, relationship, pronouns, environment, topic, mood, and source reference IDs. Audio renderer: vLLM-Omni VoxCPM2 Audio format: mono WAV, 48 kHz, 30 seconds per chunk This is a research dataset. Review the source/reference licensing and the… See the full description on the dataset page: https://huggingface.co/datasets/tsdocode/open-vi-dialog-synthetic-100h.audiotext-to-speech10K<n<100K0 likes97 downloads1mo agoHugging Face17ivkond /synthetic-speech-diarization-ru synthetic-speech-diarization-ru Synthetic speech diarization dataset in Parquet format. Dataset Details Number of tracks: 2000 Sampling rate: 16000 Hz Audio format: Embedded in Parquet files (Audio feature compatible) Storage: Parquet format for efficient loading Dataset Structure The dataset contains audio tracks with speaker diarization annotations, stored directly in Parquet format. Features audio: Audio waveform (Audio feature with array and… See the full description on the dataset page: https://huggingface.co/datasets/ivkond/synthetic-speech-diarization-ru.tabularautomatic-speech-recognition1K<n<10K0 likes93 downloads10mo agoHugging Face18yigagilbert /synthetic-parallel-external Synthetic Parallel EN↔LG — external Voice-controlled synthetic parallel speech dataset for Luganda-English speech-to-speech translation, generated by the Hibiki-Zero fine-tuning pipeline. Generation Component Model Translation Sunbird/translate-nllb-3.3b-salt TTS Sunbird/orpheus-3b-tts-multilingual English speakers: salt_eng_0001, salt_eng_0002, salt_eng_0003 Luganda speakers: salt_lug_0001, waxal_lug_0001, waxal_lug_0002, waxal_lug_0003, waxal_lug_0004… See the full description on the dataset page: https://huggingface.co/datasets/yigagilbert/synthetic-parallel-external.audioautomatic-speech-recognition100K<n<1M0 likes87 downloads4mo agoHugging Face19Reubencf /multilingual-synthetic-tts Multilingual Synthetic TTS Dataset 🏆 Submitted to the Uncharted Data Challenge hosted by Adaption Labs — credit to Adaptive Data by Adaption for organizing the hackathon. A large-scale synthetic multilingual speech dataset — 68,677 clips across 9 languages, generated with Qwen3-TTS-12Hz-1.7B-Base using zero-shot voice cloning from 5 reference speakers. Intended for training and evaluating TTS, ASR, voice conversion, and multilingual speech models. Each clip is paired with the… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/multilingual-synthetic-tts.audiotext-to-speech10K<n<100K2 likes83 downloads5mo agoHugging Face20niobures /synthetic-speech-diarization-ru synthetic-speech-diarization-ru Synthetic speech diarization dataset in Parquet format. Dataset Details Number of tracks: 2000 Sampling rate: 16000 Hz Audio format: Embedded in Parquet files (Audio feature compatible) Storage: Parquet format for efficient loading Dataset Structure The dataset contains audio tracks with speaker diarization annotations, stored directly in Parquet format. Features audio: Audio waveform (Audio feature with array and… See the full description on the dataset page: https://huggingface.co/datasets/niobures/synthetic-speech-diarization-ru.tabularautomatic-speech-recognition1K<n<10K0 likes73 downloads5mo agoHugging Face21hackerlim7 /burmese-synthetic-speech-corpus Burmese Synthetic Speech Corpus (DatarrX/burmese-synthetic-speech-corpus) Overview The Burmese Synthetic Speech Corpus is a high-fidelity, manually curated audio dataset specifically designed to advance Text-to-Speech (TTS) systems, speech recognition, and other audio-driven Machine Learning tasks for the Burmese (Myanmar) language. Created by DatarrX, this dataset bridges the gap in low-resource speech technologies by providing highly natural, native-sounding… See the full description on the dataset page: https://huggingface.co/datasets/hackerlim7/burmese-synthetic-speech-corpus.audiotext-to-speech1K<n<10K0 likes60 downloads2mo agoHugging Face22IbrahimDayax /somali-asr-synthetic-youtube Somali ASR Synthetic YouTube Dataset A Somali-language speech dataset derived from YouTube audio, intended for training and evaluating automatic speech recognition (ASR) and speech-to-text (STT) models. Transcriptions were generated synthetically (silver-standard) via ASR bootstrapping. Dataset Summary Split Samples train ~4,393 validation 200 test 100 Total ~4,693 Language: Somali (so) Audio format: WAV, 16 kHz, mono, 16-bit PCM Total… See the full description on the dataset page: https://huggingface.co/datasets/IbrahimDayax/somali-asr-synthetic-youtube.audioautomatic-speech-recognition1K<n<10K1 likes49 downloads4mo agoHugging Face23Taklaxbr /Synthetic_Turkish_TTS_Data Not: Bu veri setinin dokümantasyonu Türk yapay zeka topluluğuna katkı sağlamak amacıyla VeriPazarı tarafından Türkçeye çevrilmiştir. Orijinal veri seti Anilosan15 tarafından geliştirilmiş olup, VeriPazarı tarafından Türk AI ekosistemi için arşivlenmiştir. 🔗 Orijinal Kaynak: Anilosan15/Synthetic_Turkish_TTS_Data 🔗 Derleyen Platform: VeriPazarı Sentetik Türkçe TTS Veri Seti (Synthetic Turkish TTS Data) Bu veri seti, çoklu konuşma senaryoları üzerinden sentetik Türkçe metinler… See the full description on the dataset page: https://huggingface.co/datasets/Taklaxbr/Synthetic_Turkish_TTS_Data.audiotext-to-speech10K<n<100K1 likes39 downloads3mo agoHugging Face24yigagilbert /synthetic-parallel-salt Synthetic Parallel EN↔LG — salt Voice-controlled synthetic parallel speech dataset for Luganda-English speech-to-speech translation, generated by the Hibiki-Zero fine-tuning pipeline. Generation Component Model Translation Sunbird/translate-nllb-3.3b-salt TTS Sunbird/orpheus-3b-tts-multilingual English speakers: salt_eng_0001, salt_eng_0002, salt_eng_0003 Luganda speakers: salt_lug_0001, waxal_lug_0001, waxal_lug_0002, waxal_lug_0003, waxal_lug_0004… See the full description on the dataset page: https://huggingface.co/datasets/yigagilbert/synthetic-parallel-salt.audioautomatic-speech-recognition10K<n<100K0 likes30 downloads4mo agoHugging Face25language-and-voice-lab /samromur_syntheticSamrómur Synthetic consists of 72 hours of synthetized speech in Icelandic.audioautomatic-speech-recognition10K<n<100K1 likes28 downloads2y agoHugging Face26CLEAR-Global /Chichewa-Synthetic-ASR-DatasetgatedSynthetic Chichewa ASR dataset generated using a fine-tuned version of the YourTTS model. Sample rate: 24kHz. Total duration: 550 hours. audioautomatic-speech-recognition100K<n<1M2 likes27 downloads1y agoHugging Face27MohamedGomaa30 /Synthetic-Egy-Speech-Dataset Synthetic Egyptian Speech Dataset A curated dataset of 1000 Egyptian Arabic speech samples — the best audio selected across 4 TTS models for each prompt, with transcription text and quality metadata. Dataset Description Each entry contains: id: Unique prompt identifier (e.g., egy_0001) text: Egyptian Arabic transcription text audio_path: Path to the best-selected .wav audio file model: TTS model that produced the best audio (lahgtna, chatterbox_egyptian, egtts_v01, or… See the full description on the dataset page: https://huggingface.co/datasets/MohamedGomaa30/Synthetic-Egy-Speech-Dataset.audiotext-to-speech1K<n<10K2 likes24 downloads4mo agoHugging Face28silvermango9927 /synthetic-asr-vi Synthetic ASR data — vi Generated by Valsea-ASR/synthetic-data-pipeline. Audio is synthetic (TTS), targeted as training data for downstream ASR finetuning. Total audio: 50.7 hr across short (5s) and long (30s) length buckets, each in clean and augmented variants. Loading from datasets import load_dataset ds = load_dataset("<org>/synthetic-asr-vi", "short_clean") print(ds["train"][0]["audio"]) # {"array": np.ndarray, "sampling_rate": 16000, "path": "..."}… See the full description on the dataset page: https://huggingface.co/datasets/silvermango9927/synthetic-asr-vi.audioautomatic-speech-recognition1K<n<10K0 likes21 downloads2mo agoHugging Face29kvjones0243 /deafbench-synthetic-v2 DeafBench synthetic-v2 DeafBench synthetic-v2 is a frozen, 25-sample benchmark specification for testing whether automatic speech recognition systems preserve information that matters in accessible captions. It keeps conventional word error rate separate from typed critical-information recall so a plausible transcript cannot hide a lost time, digit sequence, username, code, Wi-Fi name, or proper name. This repository publishes the reference text and reproducibility metadata. It… See the full description on the dataset page: https://huggingface.co/datasets/kvjones0243/deafbench-synthetic-v2.textautomatic-speech-recognitionn<1K0 likes21 downloads1mo agoHugging Face30CLEAR-Global /Luo-Synthetic-ASR-DatasetgatedSynthetic Dholuo ASR dataset generated using a fine-tuned version of the YourTTS model. Sample rate: 24kHz. Total duration: 775 hours. audioautomatic-speech-recognition100K<n<1M1 likes19 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.