CoolFace
Datasetpublic

AigizK/bashkort_commands_omnivoice

Bashkort Commands OmniVoice Partial eleven-label command snapshot generated with k2-fsa/OmniVoice using the same cross-lingual voice-cloning recipe as AigizK/homai_wake_word_omnivoice. Generation was stopped at the user's request after 41,525 complete reference groups had been committed. For every included reference row from the train split of: bond005/sova_rudevices the dataset contains one recording of every command: Айвика — Russian Айвикә — Bashkir Айһылыу — Bashkir… See the full description on the dataset page: https://huggingface.co/datasets/AigizK/bashkort_commands_omnivoice.

sourceHugging Faceotherupdated 2mo agoView on Hugging Face
0likes300downloads
Dataset Card

Bashkort Commands OmniVoice

Partial eleven-label command snapshot generated with k2-fsa/OmniVoice using the same cross-lingual voice-cloning recipe as AigizK/homai_wake_word_omnivoice. Generation was stopped at the user's request after 41,525 complete reference groups had been committed.

For every included reference row from the train split of:

  • bond005/sova_rudevices

the dataset contains one recording of every command:

  • Айвика — Russian
  • Айвикә — Bashkir
  • Айһылыу — Bashkir
  • Айсылу — Russian
  • Айсылыу — Bashkir
  • Азамат — Bashkir
  • Арыҫлан — Bashkir
  • Арыслан — Russian
  • Сәлимә — Bashkir
  • Салима — Russian
  • дуҫҡайым — Bashkir

Dataset structure

  • Split: train
  • Columns: audio, text
  • Source reference rows: 41,525
  • Coverage: partial frozen snapshot of bond005/sova_rudevices:train
  • Commands per reference row: 11
  • Generated rows: 456,775
  • Rows per command: 41,525
  • Audio: mono, 16 kHz, FLAC PCM16
  • Generated duration: approximately 171.88 hours

The source identifiers and transcripts are not published. As in the source Homai OmniVoice corpus, readable synthesis duration/loudness outliers are kept so every committed reference voice remains represented. This snapshot does not claim complete coverage of the reference dataset.

Licensing note

The generated dataset is marked license: other because the two reference datasets do not advertise the same license. Users must review the current license and usage terms of both reference datasets and OmniVoice before reuse.