CoolFace
Datasetpublic

AigizK/homai_wake_word_omnivoice

Homai Wake Word OmniVoice Synthetic two-label wake-word dataset generated with k2-fsa/OmniVoice using cross-lingual voice cloning. For every reference row from all train, validation, and test splits of: bond005/sova_rudevices bond005/sberdevices_golos_100h_farfield the dataset contains two generated recordings: Һомай, generated with OmniVoice language Bashkir; Хомай, generated with OmniVoice language Russian. Dataset structure Split: train Columns: audio… See the full description on the dataset page: https://huggingface.co/datasets/AigizK/homai_wake_word_omnivoice.

sourceHugging Faceotherupdated 2mo agoView on Hugging Face
0likes292downloads
Dataset Card

Homai Wake Word OmniVoice

Synthetic two-label wake-word dataset generated with k2-fsa/OmniVoice using cross-lingual voice cloning.

For every reference row from all train, validation, and test splits of:

  • —bond005/sova_rudevices
  • —bond005/sberdevices_golos_100h_farfield

the dataset contains two generated recordings:

  • —Һомай, generated with OmniVoice language Bashkir;
  • —Хомай, generated with OmniVoice language Russian.

Dataset structure

  • —Split: train
  • —Columns: audio, text
  • —Source reference rows: 105,660
  • —Generated rows: 211,320
  • —Rows per label: 105,660
  • —Audio: mono, 16 kHz, FLAC PCM16
  • —Generated duration: approximately 70.39 hours

The source split names are intentionally merged into a single generated training split. Source identifiers and transcripts are not published. The dataset is intentionally unfiltered: some synthesis duration or loudness outliers may remain so that every successfully generated source pair is kept.

Licensing note

The generated dataset is marked license: other because the two reference datasets do not advertise the same license. Users must review the current license and usage terms of both reference datasets and OmniVoice before reuse.