CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01amanuelbyte /omnivoice-best-of-n-training 🎙 Best-of-N Voice Cloning Training Data Curated training dataset for fine-tuning OmniVoice for IWSLT 2026. 🎧 Listen to the Audio This dataset has playable audio columns — click on any row in the dataset viewer to listen to both the reference audio (original speaker) and the best synthesized audio (selected by quality score). Dataset Description For each sentence in the dev split of ymoslem/acl-6060 (468 samples × 3 languages), we synthesized audio with the… See the full description on the dataset page: https://huggingface.co/datasets/amanuelbyte/omnivoice-best-of-n-training.audiotext-to-speech1K<n<10K0 likes419 downloads5mo agoHugging Face02SynDataLab-EN /omnivoice-th Sample rate 24 kHz Voice-designed 10,000 Voice-cloned 9,833 Total 19,833 audiotext-to-speech10K<n<100K1 likes332 downloads6mo agoHugging Face03AigizK /bashkort_commands_omnivoice Bashkort Commands OmniVoice Partial eleven-label command snapshot generated with k2-fsa/OmniVoice using the same cross-lingual voice-cloning recipe as AigizK/homai_wake_word_omnivoice. Generation was stopped at the user's request after 41,525 complete reference groups had been committed. For every included reference row from the train split of: bond005/sova_rudevices the dataset contains one recording of every command: Айвика — Russian Айвикә — Bashkir Айһылыу — Bashkir… See the full description on the dataset page: https://huggingface.co/datasets/AigizK/bashkort_commands_omnivoice.audioaudio-classification100K<n<1M0 likes308 downloads2mo agoHugging Face04SynDataLab-EN /omnivoice-zh Sample rate 24 kHz Voice-designed 10,000 Voice-cloned 9,946 Total 19,946 audiotext-to-speech10K<n<100K0 likes300 downloads6mo agoHugging Face05SynDataLab-EN /omnivoice-it Sample rate 24 kHz Voice-designed 10,000 Voice-cloned 10,000 Total 20,000 audiotext-to-speech10K<n<100K0 likes273 downloads6mo agoHugging Face06SynDataLab-EN /omnivoice-tr Sample rate 24 kHz Voice-designed 10,000 Voice-cloned 10,000 Total 20,000 audiotext-to-speech10K<n<100K0 likes249 downloads6mo agoHugging Face07AigizK /homai_wake_word_omnivoice Homai Wake Word OmniVoice Synthetic two-label wake-word dataset generated with k2-fsa/OmniVoice using cross-lingual voice cloning. For every reference row from all train, validation, and test splits of: bond005/sova_rudevices bond005/sberdevices_golos_100h_farfield the dataset contains two generated recordings: Һомай, generated with OmniVoice language Bashkir; Хомай, generated with OmniVoice language Russian. Dataset structure Split: train Columns: audio… See the full description on the dataset page: https://huggingface.co/datasets/AigizK/homai_wake_word_omnivoice.audioaudio-classification100K<n<1M0 likes246 downloads2mo agoHugging Face08oddadmix /msa-omnivoice-tts-v1 MSA-OmniVoice-v1 Dataset Description MSA-OmniVoice-v1 is a 50-hour synthetic Modern Standard Arabic (MSA) speech dataset generated using OmniVoice. The dataset contains high-quality synthetic speech from a single speaker paired with fully diacritized (تشكيل) transcripts. It is intended for training and fine-tuning Arabic speech models, including Text-to-Speech (TTS), Automatic Speech Recognition (ASR), speech representation learning, and alignment tasks.… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/msa-omnivoice-tts-v1.audio10K<n<100K0 likes240 downloads3mo agoHugging Face09SynDataLab-EN /omnivoice-es Sample rate 24 kHz Voice-designed 10,000 Voice-cloned 10,000 Total 20,000 audiotext-to-speech10K<n<100K0 likes231 downloads6mo agoHugging Face10SynDataLab-EN /omnivoice-fr Sample rate 24 kHz Voice-designed 10,000 Voice-cloned 10,000 Total 20,000 audiotext-to-speech10K<n<100K0 likes218 downloads6mo agoHugging Face11STBack23 /omnivoice-vi OmniVoice VI — Giọng Việt + SRT lồng tiếng Dataset chứa 6 giọng tiếng Việt và công cụ speak.py để chạy trên Google Colab với OmniVoice. Giọng có sẵn Slug Tên ban_mai Ban Mai lan_trinh Lan Trinh ngan_ha Ngan Ha ngoc_huyen Ngoc Huyen thao_trinh Thao Trinh tuong_vy Tuong Vy Mỗi giọng gồm profile.json, voice.pt (prompt cache), audio mẫu và ref_text.txt. Chạy trên Colab Lồng tiếng SRT: mở colab/Omivoice_VI_Colab.ipynb Clone… See the full description on the dataset page: https://huggingface.co/datasets/STBack23/omnivoice-vi.audion<1K12 likes198 downloads3mo agoHugging Face12SynDataLab-EN /omnivoice-ja Sample rate 24 kHz Voice-designed 10,000 Voice-cloned 9,958 Total 19,958 audiotext-to-speech10K<n<100K0 likes172 downloads6mo agoHugging Face13SynDataLab-EN /omnivoice-ru Sample rate 24 kHz Voice-designed 10,000 Voice-cloned 9,999 Total 19,999 audiotext-to-speech10K<n<100K0 likes158 downloads6mo agoHugging Face14TruongVuVan /omnivoice-vi2 OmniVoice VI — Giọng Việt + SRT lồng tiếng Dataset chứa 6 giọng tiếng Việt và công cụ speak.py để chạy trên Google Colab với OmniVoice. Giọng có sẵn Slug Tên ban_mai Ban Mai lan_trinh Lan Trinh ngan_ha Ngan Ha ngoc_huyen Ngoc Huyen thao_trinh Thao Trinh tuong_vy Tuong Vy Mỗi giọng gồm profile.json, voice.pt (prompt cache), audio mẫu và ref_text.txt. Chạy trên Colab Lồng tiếng SRT: mở colab/Omivoice_VI_Colab.ipynb Clone… See the full description on the dataset page: https://huggingface.co/datasets/TruongVuVan/omnivoice-vi2.audion<1K0 likes103 downloads4d agoHugging Face15SynDataLab-EN /omnivoice-test-th OmniVoice Test - Thai Synthetic Thai speech dataset generated with OmniVoice. 100 voice designs (unique synthetic speakers) 100 voice clones (cloned from voice design speakers) Sample rate: 24 kHz audiotext-to-speechn<1K0 likes81 downloads6mo agoHugging Face16Rcarvalo /omnivoice-fr-100htabular1K<n<10K0 likes79 downloads6mo agoHugging Face17openbank-uz /omnivoice_gen_test omnivoice_gen_test Uzbek TTS evaluation set for OmniVoice voice cloning: 30 number-heavy prompts x 6 reference voices = 180 generations. Every prompt deliberately mixes digits, times, dates, prices, phone numbers and scores, so the set stresses number verbalisation as much as voice similarity. Columns column type description id string ref{voice}_t{prompt}, e.g. ref3_t17 reference_id string reference voice, 1-6 text_id int prompt index, 0-29… See the full description on the dataset page: https://huggingface.co/datasets/openbank-uz/omnivoice_gen_test.audiotext-to-speechn<1K0 likes61 downloads27d agoHugging Face18eliya /omnivoice_deepfake_datasetgated omnivoice_deepfake_dataset 210.4 hours of synthetic speech deepfakes — 179,868 clips across 6 corpora and 5 languages (Mandarin, Spanish, French, Italian, Japanese), generated via voice cloning and hundreds of designed voice profiles. Generated with OmniVoice, built to train/evaluate speech deepfake detectors (part of the Forensics model family). Built to improve detection of voice cloning, deepfakes, highly realistic synthetic voices, voice conversion/changing, and similar… See the full description on the dataset page: https://huggingface.co/datasets/eliya/omnivoice_deepfake_dataset.audioaudio-classification100K<n<1M1 likes55 downloads21d agoHugging Face19Ex0TiiC /omnivoice-nvs38k-shards-v2text10K<n<100K0 likes53 downloads2mo agoHugging Face20hoanglinhn0 /omnivoice-vie OmniVoice VI — Giọng Việt + SRT lồng tiếng Dataset chứa 6 giọng tiếng Việt và công cụ speak.py để chạy trên Google Colab với OmniVoice. Giọng có sẵn Slug Tên ban_mai Ban Mai lan_trinh Lan Trinh ngan_ha Ngan Ha ngoc_huyen Ngoc Huyen thao_trinh Thao Trinh tuong_vy Tuong Vy Mỗi giọng gồm profile.json, voice.pt (prompt cache), audio mẫu và ref_text.txt. Chạy trên Colab Mở notebook colab/Omivoice_VI_Colab.ipynb Đặt HF_REPO =… See the full description on the dataset page: https://huggingface.co/datasets/hoanglinhn0/omnivoice-vie.audion<1K0 likes38 downloads3mo agoHugging Face21mimba /omnivoice-voicebankgated Mimba OmniVoice Voice Bank — Reusable Multi-Speaker Reference Voices A bank of distinct synthetic reference voices, generated with OmniVoice Voice Design (attribute-driven generation, no reference audio). Each entry is a short reference clip + its attributes + a speaker embedding. The bank is language-agnostic: it captures only the timbre of each voice. It is meant to be reused as ref_audio to generate multi-speaker TTS corpora in any language (Plateau Malagasy plt, Ngiemboon… See the full description on the dataset page: https://huggingface.co/datasets/mimba/omnivoice-voicebank.audiotext-to-speech1K<n<10K4 likes36 downloads2mo agoHugging Face22SynDataLab-EN /EchoTTS-OmniVoice-En EchoTTS + OmniVoice English Synthetic English conversational speech dataset with 200 audio samples. 100 EchoTTS — unique random voices generated with EchoTTS (no speaker reference, seed-based voice diversity) 100 OmniVoice clones — voice clones using EchoTTS audio as reference, generated with OmniVoice EchoTTS sample rate: 44.1 kHz OmniVoice sample rate: 24 kHz Language: English audiotext-to-speechn<1K1 likes29 downloads5mo agoHugging Face23vnpost-ai /speechsynth-omnivoice-v2audio1K<n<10K0 likes27 downloads1mo agoHugging Face24cagataydev /omni-voice-trainingaudion<1K0 likes25 downloads6mo agoHugging Face25SynDataLab-EN /EchoTTS-OmniVoice-en-20kaudio10K<n<100K1 likes24 downloads5mo agoHugging Face26amanuelbyte /omnivoice-best-of-n-dev-eval 🎙 Best-of-N Voice Cloning Training Data Curated training dataset for fine-tuning OmniVoice for IWSLT 2026. 🎧 Listen to the Audio This dataset has playable audio columns — click on any row in the dataset viewer to listen to both the reference audio (original speaker) and the best synthesized audio (selected by quality score). Dataset Description For each sentence in the dev split of ymoslem/acl-6060 (884 samples × 3 languages), we synthesized audio with the… See the full description on the dataset page: https://huggingface.co/datasets/amanuelbyte/omnivoice-best-of-n-dev-eval.audiotext-to-speech1K<n<10K0 likes23 downloads5mo agoHugging Face27SynDataLab-EN /omnivoice-test-zh OmniVoice Test - Chinese Synthetic Chinese speech dataset generated with OmniVoice. 100 voice designs (unique synthetic speakers) 100 voice clones (cloned from voice design speakers) Sample rate: 24 kHz audiotext-to-speechn<1K0 likes22 downloads6mo agoHugging Face28SynDataLab-EN /omnivoice-test-tr OmniVoice Test - Turkish Synthetic Turkish speech dataset generated with OmniVoice. 100 voice designs (unique synthetic speakers) 100 voice clones (cloned from voice design speakers) Sample rate: 24 kHz audiotext-to-speechn<1K0 likes18 downloads6mo agoHugging Face29SynDataLab-EN /omnivoice-test-it OmniVoice Test - Italian Synthetic Italian speech dataset generated with OmniVoice. 100 voice designs (unique synthetic speakers) 100 voice clones (cloned from voice design speakers) Sample rate: 24 kHz audiotext-to-speechn<1K0 likes13 downloads6mo agoHugging Face30sw-voice /swamiji-omnivoice-refsaudion<1K0 likes13 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.