AigizK/homai_wake_word_omnivoice
Homai Wake Word OmniVoice Synthetic two-label wake-word dataset generated with k2-fsa/OmniVoice using cross-lingual voice cloning. For every reference row from all train, validation, and test splits of: bond005/sova_rudevices bond005/sberdevices_golos_100h_farfield the dataset contains two generated recordings: Һомай, generated with OmniVoice language Bashkir; Хомай, generated with OmniVoice language Russian. Dataset structure Split: train Columns: audio… See the full description on the dataset page: https://huggingface.co/datasets/AigizK/homai_wake_word_omnivoice.
Homai Wake Word OmniVoice
Synthetic two-label wake-word dataset generated with k2-fsa/OmniVoice using cross-lingual voice cloning.
For every reference row from all train, validation, and test splits of:
bond005/sova_rudevicesbond005/sberdevices_golos_100h_farfield
the dataset contains two generated recordings:
Һомай, generated with OmniVoice languageBashkir;Хомай, generated with OmniVoice languageRussian.
Dataset structure
- Split:
train - Columns:
audio,text - Source reference rows: 105,660
- Generated rows: 211,320
- Rows per label: 105,660
- Audio: mono, 16 kHz, FLAC PCM16
- Generated duration: approximately 70.39 hours
The source split names are intentionally merged into a single generated training split. Source identifiers and transcripts are not published. The dataset is intentionally unfiltered: some synthesis duration or loudness outliers may remain so that every successfully generated source pair is kept.
Licensing note
The generated dataset is marked license: other because the two reference datasets do not advertise the same license. Users must review the current license and usage terms of both reference datasets and OmniVoice before reuse.
