CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01amanuelbyte /omnivoice-best-of-n-training 🎙 Best-of-N Voice Cloning Training Data Curated training dataset for fine-tuning OmniVoice for IWSLT 2026. 🎧 Listen to the Audio This dataset has playable audio columns — click on any row in the dataset viewer to listen to both the reference audio (original speaker) and the best synthesized audio (selected by quality score). Dataset Description For each sentence in the dev split of ymoslem/acl-6060 (468 samples × 3 languages), we synthesized audio with the… See the full description on the dataset page: https://huggingface.co/datasets/amanuelbyte/omnivoice-best-of-n-training.audiotext-to-speech1K<n<10K0 likes413 downloads6mo agoHugging Face02AigizK /bashkort_commands_omnivoice Bashkort Commands OmniVoice Partial eleven-label command snapshot generated with k2-fsa/OmniVoice using the same cross-lingual voice-cloning recipe as AigizK/homai_wake_word_omnivoice. Generation was stopped at the user's request after 41,525 complete reference groups had been committed. For every included reference row from the train split of: bond005/sova_rudevices the dataset contains one recording of every command: Айвика — Russian Айвикә — Bashkir Айһылыу — Bashkir… See the full description on the dataset page: https://huggingface.co/datasets/AigizK/bashkort_commands_omnivoice.audioaudio-classification100K<n<1M0 likes354 downloads2mo agoHugging Face03SynDataLab-EN /omnivoice-tr Sample rate 24 kHz Voice-designed 10,000 Voice-cloned 10,000 Total 20,000 audiotext-to-speech10K<n<100K0 likes340 downloads6mo agoHugging Face04SynDataLab-EN /omnivoice-fr Sample rate 24 kHz Voice-designed 10,000 Voice-cloned 10,000 Total 20,000 audiotext-to-speech10K<n<100K0 likes338 downloads6mo agoHugging Face05SynDataLab-EN /omnivoice-th Sample rate 24 kHz Voice-designed 10,000 Voice-cloned 9,833 Total 19,833 audiotext-to-speech10K<n<100K1 likes333 downloads6mo agoHugging Face06oddadmix /msa-omnivoice-tts-v1 MSA-OmniVoice-v1 Dataset Description MSA-OmniVoice-v1 is a 50-hour synthetic Modern Standard Arabic (MSA) speech dataset generated using OmniVoice. The dataset contains high-quality synthetic speech from a single speaker paired with fully diacritized (تشكيل) transcripts. It is intended for training and fine-tuning Arabic speech models, including Text-to-Speech (TTS), Automatic Speech Recognition (ASR), speech representation learning, and alignment tasks.… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/msa-omnivoice-tts-v1.audio10K<n<100K1 likes289 downloads3mo agoHugging Face07AigizK /homai_wake_word_omnivoice Homai Wake Word OmniVoice Synthetic two-label wake-word dataset generated with k2-fsa/OmniVoice using cross-lingual voice cloning. For every reference row from all train, validation, and test splits of: bond005/sova_rudevices bond005/sberdevices_golos_100h_farfield the dataset contains two generated recordings: Һомай, generated with OmniVoice language Bashkir; Хомай, generated with OmniVoice language Russian. Dataset structure Split: train Columns: audio… See the full description on the dataset page: https://huggingface.co/datasets/AigizK/homai_wake_word_omnivoice.audioaudio-classification100K<n<1M0 likes285 downloads2mo agoHugging Face08SynDataLab-EN /omnivoice-zh Sample rate 24 kHz Voice-designed 10,000 Voice-cloned 9,946 Total 19,946 audiotext-to-speech10K<n<100K0 likes282 downloads6mo agoHugging Face09SynDataLab-EN /omnivoice-it Sample rate 24 kHz Voice-designed 10,000 Voice-cloned 10,000 Total 20,000 audiotext-to-speech10K<n<100K0 likes281 downloads6mo agoHugging Face10SynDataLab-EN /omnivoice-es Sample rate 24 kHz Voice-designed 10,000 Voice-cloned 10,000 Total 20,000 audiotext-to-speech10K<n<100K0 likes230 downloads6mo agoHugging Face11STBack23 /omnivoice-vi OmniVoice VI — Giọng Việt + SRT lồng tiếng Dataset chứa 6 giọng tiếng Việt và công cụ speak.py để chạy trên Google Colab với OmniVoice. Giọng có sẵn Slug Tên ban_mai Ban Mai lan_trinh Lan Trinh ngan_ha Ngan Ha ngoc_huyen Ngoc Huyen thao_trinh Thao Trinh tuong_vy Tuong Vy Mỗi giọng gồm profile.json, voice.pt (prompt cache), audio mẫu và ref_text.txt. Chạy trên Colab Lồng tiếng SRT: mở colab/Omivoice_VI_Colab.ipynb Clone… See the full description on the dataset page: https://huggingface.co/datasets/STBack23/omnivoice-vi.audion<1K12 likes195 downloads3mo agoHugging Face12SynDataLab-EN /omnivoice-ja Sample rate 24 kHz Voice-designed 10,000 Voice-cloned 9,958 Total 19,958 audiotext-to-speech10K<n<100K0 likes177 downloads6mo agoHugging Face13SynDataLab-EN /omnivoice-ru Sample rate 24 kHz Voice-designed 10,000 Voice-cloned 9,999 Total 19,999 audiotext-to-speech10K<n<100K0 likes162 downloads6mo agoHugging Face14phucsd /omnivoice-audio-storageaudion<1K0 likes138 downloads17d agoHugging Face15TruongVuVan /omnivoice-vi2 OmniVoice VI — Giọng Việt + SRT lồng tiếng Dataset chứa 6 giọng tiếng Việt và công cụ speak.py để chạy trên Google Colab với OmniVoice. Giọng có sẵn Slug Tên ban_mai Ban Mai lan_trinh Lan Trinh ngan_ha Ngan Ha ngoc_huyen Ngoc Huyen thao_trinh Thao Trinh tuong_vy Tuong Vy Mỗi giọng gồm profile.json, voice.pt (prompt cache), audio mẫu và ref_text.txt. Chạy trên Colab Lồng tiếng SRT: mở colab/Omivoice_VI_Colab.ipynb Clone… See the full description on the dataset page: https://huggingface.co/datasets/TruongVuVan/omnivoice-vi2.audion<1K0 likes104 downloads5d agoHugging Face16SynDataLab-EN /omnivoice-test-th OmniVoice Test - Thai Synthetic Thai speech dataset generated with OmniVoice. 100 voice designs (unique synthetic speakers) 100 voice clones (cloned from voice design speakers) Sample rate: 24 kHz audiotext-to-speechn<1K0 likes80 downloads6mo agoHugging Face17openbank-uz /omnivoice_gen_test omnivoice_gen_test Uzbek TTS evaluation set for OmniVoice voice cloning: 30 number-heavy prompts x 6 reference voices = 180 generations. Every prompt deliberately mixes digits, times, dates, prices, phone numbers and scores, so the set stresses number verbalisation as much as voice similarity. Columns column type description id string ref{voice}_t{prompt}, e.g. ref3_t17 reference_id string reference voice, 1-6 text_id int prompt index, 0-29… See the full description on the dataset page: https://huggingface.co/datasets/openbank-uz/omnivoice_gen_test.audiotext-to-speechn<1K0 likes62 downloads28d agoHugging Face18eliya /omnivoice_deepfake_datasetgated omnivoice_deepfake_dataset 210.4 hours of synthetic speech deepfakes — 179,868 clips across 6 corpora and 5 languages (Mandarin, Spanish, French, Italian, Japanese), generated via voice cloning and hundreds of designed voice profiles. Generated with OmniVoice, built to train/evaluate speech deepfake detectors (part of the Forensics model family). Built to improve detection of voice cloning, deepfakes, highly realistic synthetic voices, voice conversion/changing, and similar… See the full description on the dataset page: https://huggingface.co/datasets/eliya/omnivoice_deepfake_dataset.audioaudio-classification100K<n<1M1 likes57 downloads35m agoHugging Face19hoanglinhn0 /omnivoice-vie OmniVoice VI — Giọng Việt + SRT lồng tiếng Dataset chứa 6 giọng tiếng Việt và công cụ speak.py để chạy trên Google Colab với OmniVoice. Giọng có sẵn Slug Tên ban_mai Ban Mai lan_trinh Lan Trinh ngan_ha Ngan Ha ngoc_huyen Ngoc Huyen thao_trinh Thao Trinh tuong_vy Tuong Vy Mỗi giọng gồm profile.json, voice.pt (prompt cache), audio mẫu và ref_text.txt. Chạy trên Colab Mở notebook colab/Omivoice_VI_Colab.ipynb Đặt HF_REPO =… See the full description on the dataset page: https://huggingface.co/datasets/hoanglinhn0/omnivoice-vie.audion<1K0 likes39 downloads3mo agoHugging Face20mimba /omnivoice-voicebankgated Mimba OmniVoice Voice Bank — Reusable Multi-Speaker Reference Voices A bank of distinct synthetic reference voices, generated with OmniVoice Voice Design (attribute-driven generation, no reference audio). Each entry is a short reference clip + its attributes + a speaker embedding. The bank is language-agnostic: it captures only the timbre of each voice. It is meant to be reused as ref_audio to generate multi-speaker TTS corpora in any language (Plateau Malagasy plt, Ngiemboon… See the full description on the dataset page: https://huggingface.co/datasets/mimba/omnivoice-voicebank.audiotext-to-speech1K<n<10K4 likes37 downloads2mo agoHugging Face21SynDataLab-EN /EchoTTS-OmniVoice-En EchoTTS + OmniVoice English Synthetic English conversational speech dataset with 200 audio samples. 100 EchoTTS — unique random voices generated with EchoTTS (no speaker reference, seed-based voice diversity) 100 OmniVoice clones — voice clones using EchoTTS audio as reference, generated with OmniVoice EchoTTS sample rate: 44.1 kHz OmniVoice sample rate: 24 kHz Language: English audiotext-to-speechn<1K1 likes31 downloads6mo agoHugging Face22SynDataLab-EN /EchoTTS-OmniVoice-en-20kaudio10K<n<100K1 likes29 downloads5mo agoHugging Face23cagataydev /omni-voice-trainingaudion<1K0 likes24 downloads6mo agoHugging Face24vnpost-ai /speechsynth-omnivoice-v2audio1K<n<10K0 likes24 downloads1mo agoHugging Face25amanuelbyte /omnivoice-best-of-n-dev-eval 🎙 Best-of-N Voice Cloning Training Data Curated training dataset for fine-tuning OmniVoice for IWSLT 2026. 🎧 Listen to the Audio This dataset has playable audio columns — click on any row in the dataset viewer to listen to both the reference audio (original speaker) and the best synthesized audio (selected by quality score). Dataset Description For each sentence in the dev split of ymoslem/acl-6060 (884 samples × 3 languages), we synthesized audio with the… See the full description on the dataset page: https://huggingface.co/datasets/amanuelbyte/omnivoice-best-of-n-dev-eval.audiotext-to-speech1K<n<10K0 likes23 downloads5mo agoHugging Face26SynDataLab-EN /omnivoice-test-zh OmniVoice Test - Chinese Synthetic Chinese speech dataset generated with OmniVoice. 100 voice designs (unique synthetic speakers) 100 voice clones (cloned from voice design speakers) Sample rate: 24 kHz audiotext-to-speechn<1K0 likes21 downloads6mo agoHugging Face27SynDataLab-EN /omnivoice-test-tr OmniVoice Test - Turkish Synthetic Turkish speech dataset generated with OmniVoice. 100 voice designs (unique synthetic speakers) 100 voice clones (cloned from voice design speakers) Sample rate: 24 kHz audiotext-to-speechn<1K0 likes21 downloads6mo agoHugging Face28SynDataLab-EN /omnivoice-test-ru OmniVoice Test - Russian Synthetic Russian speech dataset generated with OmniVoice. 100 voice designs (unique synthetic speakers) 100 voice clones (cloned from voice design speakers) Sample rate: 24 kHz audiotext-to-speechn<1K0 likes14 downloads6mo agoHugging Face29kadirnar /omnivoice-test-de OmniVoice Test - German Synthetic German speech dataset generated with OmniVoice. 100 voice designs (unique synthetic speakers) 100 voice clones (cloned from voice design speakers) Sample rate: 24 kHz audiotext-to-speechn<1K0 likes14 downloads4mo agoHugging Face30SynDataLab-EN /omnivoice-test-fr OmniVoice Test - French Synthetic French speech dataset generated with OmniVoice. 100 voice designs (unique synthetic speakers) 100 voice clones (cloned from voice design speakers) Sample rate: 24 kHz audiotext-to-speechn<1K0 likes13 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.