omnivoice
omnivoice-best-of-n-training
🎙 Best-of-N Voice Cloning Training Data
Curated training dataset for fine-tuning OmniVoice for IWSLT 2026.
🎧 Listen to the Audio
This dataset has playable audio columns — click on any row in the dataset viewer
to listen to both the reference audio (original speaker) and the best synthesized audio
(selected by quality score).
Dataset Description
For each sentence in the dev split of ymoslem/acl-6060 (468 samples × 3 languages),
we synthesized audio with the… See the full description on the dataset page: https://huggingface.co/datasets/amanuelbyte/omnivoice-best-of-n-training.omnivoice-th
Sample rate
24 kHz
Voice-designed
10,000
Voice-cloned
9,833
Total
19,833
bashkort_commands_omnivoice
Bashkort Commands OmniVoice
Partial eleven-label command snapshot generated with k2-fsa/OmniVoice
using the same cross-lingual voice-cloning recipe as
AigizK/homai_wake_word_omnivoice. Generation was stopped at the user's
request after 41,525 complete reference groups had been committed.
For every included reference row from the train split of:
bond005/sova_rudevices
the dataset contains one recording of every command:
Айвика — Russian
Айвикә — Bashkir
Айһылыу — Bashkir… See the full description on the dataset page: https://huggingface.co/datasets/AigizK/bashkort_commands_omnivoice.omnivoice-zh
Sample rate
24 kHz
Voice-designed
10,000
Voice-cloned
9,946
Total
19,946
msa-omnivoice-tts-v1
MSA-OmniVoice-v1
Dataset Description
MSA-OmniVoice-v1 is a 50-hour synthetic Modern Standard Arabic (MSA) speech dataset generated using OmniVoice. The dataset contains high-quality synthetic speech from a single speaker paired with fully diacritized (تشكيل) transcripts.
It is intended for training and fine-tuning Arabic speech models, including Text-to-Speech (TTS), Automatic Speech Recognition (ASR), speech representation learning, and alignment tasks.… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/msa-omnivoice-tts-v1.omnivoice-it
Sample rate
24 kHz
Voice-designed
10,000
Voice-cloned
10,000
Total
20,000
