datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
broadcast-speech
Bashkir Broadcast Speech — Radio and Television
53.3 hours of speech in 293 recordings in the Bashkir language, from television and radio programmes produced by two public broadcasters of the Republic of Bashkortostan. Audio only — no transcripts in this release — which makes the set suitable for self-supervised speech pretraining for a low-resource Turkic language.
🌐 Languages of this card: English · Башҡортса · Русский
Part of the Bashkorttele dataset series — preservation… See the full description on the dataset page: https://huggingface.co/datasets/bashkorttele/broadcast-speech.bashkort_voice
Bashkort Voice
🇬🇧 English Version
Dataset Description
This is a synthetic Bashkir audio dataset generated using the OmniVoice model. It is designed to expand the availability of spoken data for the Bashkir language.
Data Preparation Process
The dataset was constructed through a cross-lingual voice cloning and generation process, using the following methodology:
Target Text: Bashkir sentences were extracted from the AigizK/bashkir-russian-parallel-corpora dataset.… See the full description on the dataset page: https://huggingface.co/datasets/AigizK/bashkort_voice.bashkort_tts_dataset
Bashkort TTS Dataset
The largest open dataset for speech synthesis in the Bashkir language — featuring multi-speaker recordings and speaking styles.
📊 Dataset Overview
Total audio files: 62,852
Speakers: 7 female, 1 male
Speaking styles: friendly, question, neutral
Languages: Bashkir
Format: MP3 audio + transcription text
🎙 How It Was Collected
Initial recording: A female voice actor recorded ~15 hours of speech in Bashkir.
Voice cloning: Using ElevenLabs… See the full description on the dataset page: https://huggingface.co/datasets/AigizK/bashkort_tts_dataset.narrated-audiobooks-brsbs
Narrated Bashkir Audiobooks — BRSBS (Bashkorttele)
≈ 78.7 hours of human-narrated audiobooks, primarily in the Bashkir language, drawn from public-domain literary works and folk epics. Recorded as accessible "talking books" by the Bashkir Republican Special Library for the Blind (BRSBS) and released for language preservation and AI/ML research.
🌐 Languages of this card: English · Башҡортса · Русский
This dataset is part of a larger series published under the Bashkorttele… See the full description on the dataset page: https://huggingface.co/datasets/bashkorttele/narrated-audiobooks-brsbs.
