CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01projecte-aina /synthetic_dem Dataset Card for synthetic_dem Dataset Summary The Synthetic DEM Corpus is the result of the first phase of a collaboration between El Colegio de México (COLMEX) and the Barcelona Supercomputing Center (BSC). It all began when COLMEX was looking for a way to have its Diccionario del Español de México (DEM), which can be accessed online, include the option to play each of its words with a Mexican accent through synthetic speech files. On the other hand, BSC is always on… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/synthetic_dem.audioautomatic-speech-recognition100K<n<1M2 likes1.5k downloads1y agoHugging Face02Silasimo /SynthGT SynthGT A Synthetic Solo-Singing Dataset for Singing-Oriented Forced Alignment Authors Silas Antonisen, Iván López-Espejo Associated paper submitted to IEEE Transactions on Audio, Speech and Language Processing. Overview SynthGT (Synthetic Ground Truth) is a synthetic English solo-singing dataset containing 4,900 singing performances with automatically generated phoneme boundary annotations. The dataset was created through music… See the full description on the dataset page: https://huggingface.co/datasets/Silasimo/SynthGT.audioautomatic-speech-recognition1K<n<10K1 likes1k downloads2mo agoHugging Face03Aalto-Speech-Synthesis /icelandic_asr Icelandic ASR Collection This repository collects six Icelandic speech corpora in directly loadable Parquet form. Audio is embedded as 16 kHz mono FLAC bytes. The repository is a convenience repackaging: the linked CLARIN-IS records and original dataset repositories remain the canonical sources and should be cited when using the data. No configuration is selected by default. Choose a corpus configuration and, for this large collection, normally choose a split explicitly.… See the full description on the dataset page: https://huggingface.co/datasets/Aalto-Speech-Synthesis/icelandic_asr.audioautomatic-speech-recognition1M<n<10M0 likes631 downloads22d agoHugging Face04TigreGotico /barranquenho-ipa-dict-synthetic Barranquenho IPA Pronunciation Dictionary The first and only IPA pronunciation dictionary of Barranquenho — the Ibero-Romance contact variety spoken in Barrancos (Baixo Alentejo, Portugal), a mixed system born of centuries of Portuguese–Spanish (Extremaduran / Andalusian) contact on the raia. Every headword is written in the Convenção Ortográfica do Barranquenho (2025) orthography and paired with a broad-phonemic IPA transcription plus Portuguese and Spanish glosses. This… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/barranquenho-ipa-dict-synthetic.texttext-to-speech1K<n<10K0 likes395 downloads2mo agoHugging Face05AfriSpeech /multivoice-synthetic-speech Synthetic Voice Samples · Africa Synthetic speech. No human speaker was recorded for any clip in this dataset. Generated with afrispeech-synth: text from africa-corpus, normalised to a universal orthography with africa-g2p, spoken by Google Gemini's Live API. 17,010 clips · 38.8 hours · 566 languages · 30 voices Every clip is a distinct sentence — no sentence is repeated Each language is read by up to 30 different voices, one sentence per voice ~1.29 hours per voice… See the full description on the dataset page: https://huggingface.co/datasets/AfriSpeech/multivoice-synthetic-speech.audiotext-to-speech10K<n<100K1 likes324 downloads8d agoHugging Face06yuriyvnv /synthetic_transcript_pt Portuguese Speech Dataset with Multiple Training Configurations A comprehensive Portuguese speech dataset offering three distinct training configurations for speech recognition research, each designed for different experimental scenarios and training paradigms. 🎯 Dataset Configurations Overview This dataset provides three carefully curated subsets to enable comprehensive speech recognition research: Configuration Training Data Validation Test Total Samples Use Case… See the full description on the dataset page: https://huggingface.co/datasets/yuriyvnv/synthetic_transcript_pt.audioautomatic-speech-recognition100K<n<1M0 likes305 downloads5mo agoHugging Face07MLRS /masri_synthetic Dataset Card for masri_synthetic Dataset Summary The MASRI-SYNTHETIC is a corpus made out of synthesized speech in Maltese. The text-to-speech (TTS) system utilized to produce the utterances was developed by the Research & Development Department of Crimsonwing p.l.c. The sentences used to create the corpus were extracted from the MLRS Corpus, which is a corpus of written or transcribed Maltese divided into different genres, including: culture, news, academic, religion… See the full description on the dataset page: https://huggingface.co/datasets/MLRS/masri_synthetic.audioautomatic-speech-recognition10K<n<100K2 likes269 downloads2y agoHugging Face08Aviv-anthonnyolime /SIWIS_French_Speech_Synthesis_Database SIWIS French Speech Synthesis Database This README provides a concise description of the dataset, including its structure, file naming conventions, and known labeling issues. Additionally, suggestions for potential improvements are outlined in the TODO section. The dataset is distributed under the Creative Commons Attribution 4.0 International (CC BY 4.0) license, permitting its use for any purpose. For more details about the database design and recording process, please refer… See the full description on the dataset page: https://huggingface.co/datasets/Aviv-anthonnyolime/SIWIS_French_Speech_Synthesis_Database.audioautomatic-speech-recognition10K<n<100K0 likes264 downloads2y agoHugging Face09QCRI /Aslema-Synth-TN Aslema-Synth-TN Aslema-Synth-TN is a fully synthetic Tunisian Derja corpus for spoken language understanding: speech annotated for intent and for slot filling. It was built for NADI 2026 Shared Task 5 to cover the intents and slots that are rare or absent in the real SLURP-TN training split, and it is the augmentation set behind the Aslema system, which ranked 1st in slot filling on the official test set. No human was recorded for this dataset. An LLM wrote the utterance text, a… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/Aslema-Synth-TN.audioaudio-classification10K<n<100K2 likes252 downloads20d agoHugging Face10Meddies /meddies-asr-synth-dialoggated Meddies ASR — Synthetic Dialog Speech (vi/en/zh) Synthetic doctor–patient consultations for ASR training, generated with Fish Audio TTS (s2.1-pro-free, 16 kHz mono FLAC) over curated conversational voice registries (en 75 / zh 165 voices) with 22 compositional scenario profiles driving expressiveness (disfluencies, emotion palettes, speed jitter, age-band casting, pause pacing). Target: 1,000 h per language. Configs vi_dialog: Vietnamese clinical dialogue turns… See the full description on the dataset page: https://huggingface.co/datasets/Meddies/meddies-asr-synth-dialog.audioautomatic-speech-recognition100K<n<1M2 likes243 downloads25d agoHugging Face11Hani89 /Synthetic-Medical-Speech-Dataset Synthetic Medical Speech Dataset Overview Synthetic Medical Speech Dataset is a synthetic dataset of audio–text pairs designed for developing and evaluating automatic speech recognition (ASR) models in the medical domain.The corpus contains thousands of short audio clips generated from medically relevant text using a text-to-speech (TTS) system.Each clip is paired with its corresponding transcript.Because all content is synthetically produced, the dataset does not contain… See the full description on the dataset page: https://huggingface.co/datasets/Hani89/Synthetic-Medical-Speech-Dataset.audioautomatic-speech-recognition10K<n<100K4 likes226 downloads1y agoHugging Face12WhissleAI /synthetic_speech_commands_PA_taggedaudioautomatic-speech-recognition1K<n<10K0 likes188 downloads1y agoHugging Face13routsourav1729 /synthetic-tts Munni — Hindi/English Synthetic TTS Voice Dataset Single-speaker, code-mixed Hindi (Devanagari) + English speech dataset built for XTTS-v2 fine-tuning. Domain is call-center / customer-support style dialogue (warranty, billing, appointment scheduling), with scripted placeholder phone numbers spoken digit-by-digit. 8,796 clips / 18.16 hours, 24 kHz mono 16-bit PCM Split: train 8,621 / validation 175 (98/2, seed 42) Duration per clip: mean 7.43s, min 3.18s, max 11.00s Single… See the full description on the dataset page: https://huggingface.co/datasets/routsourav1729/synthetic-tts.texttext-to-speech1K<n<10K0 likes175 downloads22d agoHugging Face14yigagilbert /synthetic-parallel-external Synthetic Parallel EN↔LG — external Voice-controlled synthetic parallel speech dataset for Luganda-English speech-to-speech translation, generated by the Hibiki-Zero fine-tuning pipeline. Generation Component Model Translation Sunbird/translate-nllb-3.3b-salt TTS Sunbird/orpheus-3b-tts-multilingual English speakers: salt_eng_0001, salt_eng_0002, salt_eng_0003 Luganda speakers: salt_lug_0001, waxal_lug_0001, waxal_lug_0002, waxal_lug_0003, waxal_lug_0004… See the full description on the dataset page: https://huggingface.co/datasets/yigagilbert/synthetic-parallel-external.audioautomatic-speech-recognition100K<n<1M0 likes171 downloads4mo agoHugging Face15ghanaopenai /ghana-twi-synthesized-speech Ghana Twi & Code-Switching Synthesized Speech Dataset Synthesized text-to-speech audio for Twi and English-Twi code-switching sentences. How this dataset was built 1. Source text The sentences come from two sources, combined and deduplicated: A sentence-subset sampled from ghananlpcommunity/pristine-twi-english (greedy set-cover over 3,000 articles so that every word appearing in the corpus is present in at least one selected sentence). The… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-twi-synthesized-speech.audiotext-to-speech10K<n<100K1 likes163 downloads12d agoHugging Face16yuriyvnv /synthetic_transcript_nl Dutch Synthetic Speech Transcripts This dataset contains 34,898 synthetic Dutch speech samples generated using GPT-4o-mini for transcript creation and OpenAI's TTS-1 model for speech synthesis. It was designed to augment Automatic Speech Recognition (ASR) training for low-resource scenarios, matching the linguistic distribution of Common Voice 17.0 Dutch. Dataset Description Purpose This dataset addresses the challenge of limited labeled speech data for Dutch… See the full description on the dataset page: https://huggingface.co/datasets/yuriyvnv/synthetic_transcript_nl.audioautomatic-speech-recognition10K<n<100K0 likes161 downloads10mo agoHugging Face17Wi-Fi /korean-full-duplex-synthetic-dataset-preview Korean Full-Duplex Synthetic Dataset Preview Overview Public preview of a Korean full-duplex synthetic speech dataset. This repository contains 100 conversations sampled from a corpus of 89,273 conversations (2,000.5 hours); it does not publish the full corpus audio. Preview contents 100 conversation WAV files data/representative.jsonl 24 kHz, mono, 16-bit PCM Events: normal, barge_in, backchannel, cutoff_by_user Annotation format… See the full description on the dataset page: https://huggingface.co/datasets/Wi-Fi/korean-full-duplex-synthetic-dataset-preview.audioautomatic-speech-recognitionn<1K1 likes153 downloads1mo agoHugging Face18CLEAR-Global /Hausa-Synthetic-ASR-Dataset-XTTSgatedSynthetic Hausa ASR dataset generated using a fine-tuned version of the XTTS-v2 model. Sample rate: 24kHz. Total duration: 574 hours. audioautomatic-speech-recognition100K<n<1M1 likes149 downloads1y agoHugging Face19bonalor /synthetic_maritime_radio_communication MARTTS: Maritime Radio Text-To-Speech Synthetic Corpus Synthetic VHF Maritime Communication Data for Robust ASR Evaluation Dataset accompanying the paper:A Text-to-Speech Framework for Generating Synthetic Maritime Radio Communications in ASR Evaluation Dataset Summary MARTTS is an open-source synthetic speech corpus designed to evaluate and stress-test Automatic Speech Recognition (ASR) systems operating in maritime VHF radiotelephony environments. The… See the full description on the dataset page: https://huggingface.co/datasets/bonalor/synthetic_maritime_radio_communication.audiotext-classificationn<1K4 likes137 downloads9mo agoHugging Face20abdo1819 /arabic-english-code-switching-synthetic-asr Synthetic Arabic-English Code-Switched Speech for ASR This dataset contains synthetic speech generated for Egyptian Arabic-English code-switched automatic speech recognition. It is published separately from the human review annotations so the human audio remains in its upstream Hugging Face repository. Configurations Configuration Train Test Publication status synthetic 8,655 962 Contains 5,673 ArE-CSTD-derived texts; noncommercial/share-alike terms… See the full description on the dataset page: https://huggingface.co/datasets/abdo1819/arabic-english-code-switching-synthetic-asr.audioautomatic-speech-recognition1K<n<10K0 likes124 downloads2mo agoHugging Face21milanakdj /nepali-tts-synthetic-v2gated Nepali TTS Synthetic v2 383,298 synthetic Nepali (ne) speech/text pairs, 24 kHz mono 16-bit WAV embedded as-is (no re-encode, no resampling). Generated by the synthetic_pipeline in milanakdj/TTS_training: Edge TTS synthesis → optional voice conversion against a pool of 600 real multi-speaker reference clips → ASR-based QC gate on character error rate. Read this before training on it Only 48% of rows are voice-converted. Each row carries a kept field recording… See the full description on the dataset page: https://huggingface.co/datasets/milanakdj/nepali-tts-synthetic-v2.audiotext-to-speech100K<n<1M0 likes117 downloads23d agoHugging Face22DatarrX /burmese-synthetic-speech-corpus Burmese Synthetic Speech Corpus (DatarrX/burmese-synthetic-speech-corpus) Overview The Burmese Synthetic Speech Corpus is a high-fidelity, manually curated audio dataset specifically designed to advance Text-to-Speech (TTS) systems, speech recognition, and other audio-driven Machine Learning tasks for the Burmese (Myanmar) language. Created by DatarrX, this dataset bridges the gap in low-resource speech technologies by providing highly natural, native-sounding… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/burmese-synthetic-speech-corpus.audiotext-to-speech1K<n<10K7 likes111 downloads4mo agoHugging Face23erayyapagci /turkish-synthetic-whisper-rounds1-4.5-355h Turkish Synthetic Whisper Rounds 1–4.5 Archival release of the exact 199,590-record, 355.186-hour synthetic corpus used to fine-tune the final Round 4.5 Whisper Tiny and Base models. Each row in train.jsonl references both: training_audio: the exact clean or exactly-once postprocessed waveform used in training; and clean_audio: its original synthetic clean waveform. Audio is SHA-256 deduplicated and stored in deterministic tar.zst shards. Common Voice/FLEURS evaluation audio… See the full description on the dataset page: https://huggingface.co/datasets/erayyapagci/turkish-synthetic-whisper-rounds1-4.5-355h.automatic-speech-recognition0 likes111 downloads1mo agoHugging Face24noty7gian /synthetic-multilingual-speaker-diarization Synthetic Multilingual Speaker Diarization Dataset This dataset contains synthetic multilingual speaker diarization data with Hindi, English, and Punjabi audio samples. Dataset Structure ├── audio/ # WAV audio files (16kHz) - 3417 files ├── all_samples_combined.csv # Complete dataset annotations (with silence) └── all_visible_combined.csv # Visible dataset annotations (without silence) Statistics Total samples: 3417 audio… See the full description on the dataset page: https://huggingface.co/datasets/noty7gian/synthetic-multilingual-speaker-diarization.audioautomatic-speech-recognition1K<n<10K0 likes109 downloads1y agoHugging Face25uncleMehrzad /synthetic-speaker-diarization-dataset-fa-large-3000audioaudio-classification1K<n<10K3 likes108 downloads1y agoHugging Face26ivkond /synthetic-speech-diarization-ru synthetic-speech-diarization-ru Synthetic speech diarization dataset in Parquet format. Dataset Details Number of tracks: 2000 Sampling rate: 16000 Hz Audio format: Embedded in Parquet files (Audio feature compatible) Storage: Parquet format for efficient loading Dataset Structure The dataset contains audio tracks with speaker diarization annotations, stored directly in Parquet format. Features audio: Audio waveform (Audio feature with array and… See the full description on the dataset page: https://huggingface.co/datasets/ivkond/synthetic-speech-diarization-ru.tabularautomatic-speech-recognition1K<n<10K0 likes106 downloads10mo agoHugging Face27gigant /romanian_speech_synthesis_0_8_1\ The Romanian speech synthesis (RSS) corpus was recorded in a hemianechoic chamber (anechoic walls and ceiling; floor partially anechoic) at the University of Edinburgh. We used three high quality studio microphones: a Neumann u89i (large diaphragm condenser), a Sennheiser MKH 800 (small diaphragm condenser with very wide bandwidth) and a DPA 4035 (headset-mounted condenser). Although the current release includes only speech data recorded via Sennheiser MKH 800, we may release speech data recorded via other microphones in the future. All recordings were made at 96 kHz sampling frequency and 24 bits per sample, then downsampled to 48 kHz sampling frequency. For recording, downsampling and bit rate conversion, we used ProTools HD hardware and software. We conducted 8 sessions over the course of a month, recording about 500 sentences in each session. At the start of each session, the speaker listened to a previously recorded sample, in order to attain a similar voice quality and intonation.automatic-speech-recognition11 likes105 downloads4y agoHugging Face28Aalto-Speech-Synthesis /stortinget_speech_corpus_v1.0 Dataset Card for Stortinget Speech Corpus V1.0 Overview This is the WebDataset version of the Stortinget Speech Corpus V1.0, originally created by the National Library of Norway. We re-organize it into WebDataset format for better usability. The Stortinget Speech Corpus (SSC) is a 5000+ hours speech dataset for weak supervision ASR created from audio andaligned proceedings text from Stortinget, the Norwegian Parliament. For more information, please refer to the original… See the full description on the dataset page: https://huggingface.co/datasets/Aalto-Speech-Synthesis/stortinget_speech_corpus_v1.0.audioautomatic-speech-recognition100K<n<1M0 likes104 downloads5mo agoHugging Face29Anilosan15 /Synthetic_Turkish_TTS_Data Synthetic Turkish TTS Data This dataset was created by generating synthetic Turkish text across multiple speech scenarios. The text was produced in the following domains: finance_master, cs_master, parcel_delivery, ecommerce, telecom, isp_support, technical_support, subscription, insurance, health_appointments, public_services, education_registration, and daily_speech. These synthetic texts were then synthesized with a high-quality Turkish TTS model. The dataset is intended to be… See the full description on the dataset page: https://huggingface.co/datasets/Anilosan15/Synthetic_Turkish_TTS_Data.audiotext-to-speech10K<n<100K6 likes99 downloads5mo agoHugging Face30Reubencf /multilingual-synthetic-tts Multilingual Synthetic TTS Dataset 🏆 Submitted to the Uncharted Data Challenge hosted by Adaption Labs — credit to Adaptive Data by Adaption for organizing the hackathon. A large-scale synthetic multilingual speech dataset — 68,677 clips across 9 languages, generated with Qwen3-TTS-12Hz-1.7B-Base using zero-shot voice cloning from 5 reference speakers. Intended for training and evaluating TTS, ASR, voice conversion, and multilingual speech models. Each clip is paired with the… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/multilingual-synthetic-tts.audiotext-to-speech10K<n<100K2 likes97 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.