CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01BaseLayer /uzbek-multi-speaker-35haudio10K<n<100K1 likes570 downloads26d agoHugging Face02ghanaopenai /fante-speech-text-multispeaker_lds Fante Speech-Text Multispeaker Dataset (LDS) Sentence-level aligned Fante (fat) speech dataset sourced from the Church of Jesus Christ of Latter-day Saints General Conference translations. Dataset Statistics Split Clips Hours Talks Train 29,992 58.32 405 Eval 2,028 4.09 28 Total 32,020 62.41 433 Features audio: 16 kHz mono FLAC sentence-level clips text: Fante transcript (sentence-aligned) talk_id: Source conference talk… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/fante-speech-text-multispeaker_lds.audioautomatic-speech-recognition10K<n<100K1 likes531 downloads2mo agoHugging Face03slprl /multispeaker-storycloze Multi Speaker StoryCloze A multispeaker spoken version of StoryCloze Synthesized with Kokoro TTS. The dataset was synthesized to evaluate the performance of speech language models as detailed in the paper "Scaling Analysis of Interleaved Speech-Text Language Models". We refer you to the SlamKit codebase to see how you can evaluate your SpeechLM with this dataset. sSC and tSC We split the generation for spoken-stroycloze and topic-storycloze as detailed in Twist.… See the full description on the dataset page: https://huggingface.co/datasets/slprl/multispeaker-storycloze.audio10K<n<100K2 likes378 downloads1y agoHugging Face04ghanaopenai /twi_multispeaker_audio_transcribed Twi Multispeaker Audio Transcribed Dataset Overview The Twi Multispeaker Audio Transcribed dataset is a collection of speech recordings and their transcriptions in Asante Twi, a widely spoken dialect of the Akan language in Ghana. The dataset is designed for training and evaluating automatic speech recognition (ASR) models and other natural language processing (NLP) applications. Dataset Details Source: The dataset is derived from the Financial… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/twi_multispeaker_audio_transcribed.audioautomatic-speech-recognition10K<n<100K0 likes293 downloads2y agoHugging Face05ghanaopenai /ga-multispeaker-speech-text-20k This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. This dataset is made available because of Ghana NLP's volunteer driven research work. Please consider contributing to any of our projects on Github Ga Multispeaker Audio Transcribed Dataset Overview The Ga… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ga-multispeaker-speech-text-20k.audioautomatic-speech-recognition10K<n<100K1 likes287 downloads3mo agoHugging Face06BaseLayer /uzbek-multi-speaker-25haudio10K<n<100K1 likes283 downloads29d agoHugging Face07ghanaopenai /fante-multispeaker_speech-text-20k This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. This dataset is made available because of Ghana NLP's volunteer driven research work. Please consider contributing to any of our projects on Github Fante Multispeaker Audio Transcribed Dataset Overview The Fante… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/fante-multispeaker_speech-text-20k.audioautomatic-speech-recognition10K<n<100K1 likes270 downloads3mo agoHugging Face08ghanaopenai /akuapem_multispeaker_audio_transcribed Akuapem Multispeaker Audio Transcribed Dataset Overview The Akuapem Multispeaker Audio Transcribed dataset is a collection of speech recordings and their transcriptions in Akuapem Twi, a widely spoken dialect of the Akan language in Ghana. The dataset is designed for training and evaluating automatic speech recognition (ASR) models and other natural language processing (NLP) applications. Dataset Details Source: The dataset is derived from the… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/akuapem_multispeaker_audio_transcribed.audioautomatic-speech-recognition10K<n<100K1 likes232 downloads2y agoHugging Face09ghanaopenai /twi-speech-text-multispeaker-16k This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. This dataset is made available because of Ghana NLP's volunteer driven research work. Please consider contributing to any of our projects on Github Twi Speech-Text Parallel Dataset Dataset Description This dataset… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/twi-speech-text-multispeaker-16k.audio10K<n<100K4 likes183 downloads3mo agoHugging Face10Tharyck /multispeaker-tts-ptbrDataset importado do https://gitlab.com/fb-audio-corpora audio100K<n<1M6 likes142 downloads1y agoHugging Face11MahmoudIbrahim /ar-eg-speech-tts-multi-speakersaudio10K<n<100K0 likes133 downloads7d agoHugging Face12lgris /cml-tts-filtered-multispeaker_tokenised1M<n<10M0 likes128 downloads1y agoHugging Face13ghananlpcommunity /fante-speech-text-multispeaker_lds Fante Speech-Text Multispeaker Dataset (LDS) Sentence-level aligned Fante (fat) speech dataset sourced from the Church of Jesus Christ of Latter-day Saints General Conference translations. Dataset Statistics Split Clips Hours Talks Train 29,992 58.32 405 Eval 2,028 4.09 28 Total 32,020 62.41 433 Features audio: 16 kHz mono FLAC sentence-level clips text: Fante transcript (sentence-aligned) talk_id: Source conference talk… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/fante-speech-text-multispeaker_lds.audioautomatic-speech-recognition10K<n<100K0 likes125 downloads2mo agoHugging Face14PharynxAI /hindi_multispeaker_datasetaudio1K<n<10K0 likes122 downloads1y agoHugging Face15laki35 /somali-stt-dataset-multi-speaker-v1 Dataset Structure The dataset contains the following columns: text: The Somali sentence (transcription). audio: The audio file sampled at 24,000 Hz. speaker_id: Unique integer ID (1 to 11) representing each of the 11 speakers. Metadata & Search Keywords Language: Somali (so) Speakers: 11 unique voices (balanced gender representation) Audio Quality: 24kHz, mono, clean audio Total Rows: 1,200 Total Duration: ~1.66 Hours (99.86 Minutes) Intended Use: Fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/laki35/somali-stt-dataset-multi-speaker-v1.audiotext-to-speech1K<n<10K2 likes116 downloads1mo agoHugging Face16humyn-labs /Indic-High-Fidelity-MultiSpeaker-ASR Dataset Overview This dataset contains high-quality multi-speaker conversational audio recordings curated for Automatic Speech Recognition (ASR) research across multiple Indic languages. The dataset includes: Paired audio + timestamped transcripts Natural, non-scripted conversational speech Dual-speaker interactions Segment-level speaker annotations Regionally diverse accents Audio Specifications Format: WAV (PCM 16-bit) Sampling Rate: 16 kHz Channel: Mono Speech… See the full description on the dataset page: https://huggingface.co/datasets/humyn-labs/Indic-High-Fidelity-MultiSpeaker-ASR.audioautomatic-speech-recognitionn<1K1 likes99 downloads6mo agoHugging Face17consciousengines /Synthetic-Multispeaker-Maithili-Santaliaudio1K<n<10K0 likes83 downloads12d agoHugging Face18Abdelrahman2922 /arabic-tts-saudi-multi-speaker-xtts Arabic Saudi TTS Dataset (LJSpeech Format) 🇸🇦 This dataset is designed for training Text-to-Speech (TTS) models such as XTTS_v2 using the LJSpeech format. 📌 Overview Language: Arabic (Saudi Dialect) Format: LJSpeech Use Case: TTS training (XTTS_v2, YourTTS, Tacotron, etc.) Speakers: Multi-speaker (Male & Female) Audio Format: WAV (mono recommended) Sample Rate: 22050 Hz (recommended) 📂 Structure all_data/ │ ├── wavs/ │ ├── sample_0.wav │ ├──… See the full description on the dataset page: https://huggingface.co/datasets/Abdelrahman2922/arabic-tts-saudi-multi-speaker-xtts.audio1K<n<10K2 likes72 downloads6mo agoHugging Face19voxozi /french-b2b-tts-multispeaker French B2B Multi-Speaker TTS Dataset Dataset Description A multi-speaker French text-to-speech dataset covering three B2B industry verticals: fintech/banking, e-commerce/logistics, and healthcare/medical. Audio clips are generated with diverse male and female speaker voices for conversational AI applications. Verticals Vertical Description fintech_banking Banking operations, transfers, account inquiries, fraud alerts, investments… See the full description on the dataset page: https://huggingface.co/datasets/voxozi/french-b2b-tts-multispeaker.text-to-speech100K<n<1M0 likes59 downloads3mo agoHugging Face20ghanaopenai /twi-speech-text-multispeaker-cleanaudio1K<n<10K0 likes45 downloads10mo agoHugging Face21MeryemBelkhayat /multispeakersaudion<1K0 likes42 downloads3d agoHugging Face22fiifinketia /ugtts-multispeaker-max266secs-total9hrs-sr22050gatedaudio1K<n<10K0 likes40 downloads6mo agoHugging Face23keshan /multispeaker-tts-sinhala\\nThis data set contains multi-speaker high quality transcribed audio data for Sinhala. The data set consists of wave files, and a TSV file. The file si_lk.lines.txt contains a FileID, which in tern contains the UserID and the Transcription of audio in the file. The data set has been manually quality checked, but there might still be errors. Part of this dataset was collected by Google in Sri Lanka and the rest was contributed by Path to Nirvana organization.3 likes38 downloads5y agoHugging Face24MosesJoshuaCoker /Novax_Multi_speakeraudion<1K0 likes38 downloads1mo agoHugging Face25MeryemBelkhayat /multi_speakersaudio1K<n<10K0 likes38 downloads3d agoHugging Face26Mehrdad-S /persian_multispeaker_voiceaudio1K<n<10K1 likes37 downloads2y agoHugging Face27ar17to /orpheus_tts_english_indian_multispeakeraudio10K<n<100K1 likes34 downloads1y agoHugging Face28hoolatech /multispeaker-tts-ptbrDataset importado do https://gitlab.com/fb-audio-corpora audio100K<n<1M0 likes33 downloads11d agoHugging Face29humanify /speaker_evaluation_multi_test_v0 Seamless Interaction Pairs This dataset contains paired query and document audio clips for interaction-based speaker evaluation. Each row describes a query clip and a related document clip, with segment metadata and durations for analysis. Data structure The dataset uses a single split stored in data.parquet. Audio files are stored under audio/ and referenced by relative paths in the parquet file. Columns pair_id (string): Pair identifier. interaction… See the full description on the dataset page: https://huggingface.co/datasets/humanify/speaker_evaluation_multi_test_v0.audioaudio-classificationn<1K0 likes32 downloads7mo agoHugging Face30ajikadev /salt-multispeaker-eng-splitaudio1K<n<10K0 likes22 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.