datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bashkort_commands_omnivoice
Bashkort Commands OmniVoice
Partial eleven-label command snapshot generated with k2-fsa/OmniVoice
using the same cross-lingual voice-cloning recipe as
AigizK/homai_wake_word_omnivoice. Generation was stopped at the user's
request after 41,525 complete reference groups had been committed.
For every included reference row from the train split of:
bond005/sova_rudevices
the dataset contains one recording of every command:
Айвика — Russian
Айвикә — Bashkir
Айһылыу — Bashkir… See the full description on the dataset page: https://huggingface.co/datasets/AigizK/bashkort_commands_omnivoice.bashkort_voice
Bashkort Voice
🇬🇧 English Version
Dataset Description
This is a synthetic Bashkir audio dataset generated using the OmniVoice model. It is designed to expand the availability of spoken data for the Bashkir language.
Data Preparation Process
The dataset was constructed through a cross-lingual voice cloning and generation process, using the following methodology:
Target Text: Bashkir sentences were extracted from the AigizK/bashkir-russian-parallel-corpora dataset.… See the full description on the dataset page: https://huggingface.co/datasets/AigizK/bashkort_voice.bashkort_tts_dataset
Bashkort TTS Dataset
The largest open dataset for speech synthesis in the Bashkir language — featuring multi-speaker recordings and speaking styles.
📊 Dataset Overview
Total audio files: 62,852
Speakers: 7 female, 1 male
Speaking styles: friendly, question, neutral
Languages: Bashkir
Format: MP3 audio + transcription text
🎙 How It Was Collected
Initial recording: A female voice actor recorded ~15 hours of speech in Bashkir.
Voice cloning: Using ElevenLabs… See the full description on the dataset page: https://huggingface.co/datasets/AigizK/bashkort_tts_dataset.bashtube_voice
