datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ruslan-ttsruslan-datasetruslan-stressed
RUSLAN with Word Stress Marks · RUSLAN с проставленными ударениями
English / Русский
English
What is this?
A drop-in replacement for the metadata of the RUSLAN Russian single-speaker
TTS corpus, with word-stress marks added to every multi-syllabic Russian
word in the transcripts. Audio is bundled unchanged.
The motivation is to train Russian TTS models (e.g. Kokoro, Tacotron, VITS,
StyleTTS, XTTS) that pronounce words with correct lexical stress.
Vanilla… See the full description on the dataset page: https://huggingface.co/datasets/stilletto/ruslan-stressed.ruslan_sova_aiThis is a saved ruslan dataset from SOVA AI
ruslan_dataruslan-stressed-mini
RUSLAN stressed — mini sanity-check dataset
This is a 200-sample mini version of stilletto/ruslan-stressed used to
verify that the WebDataset tar layout is parsed correctly by the HuggingFace
dataset viewer before the full 22,200-sample dataset is repacked the same way.
Layout (WebDataset):
mini_part_001.tar # samples 000000…000099 (wav + paired txt)
mini_part_002.tar # samples 000100…000199 (wav + paired txt)
Each tar contains paired files sharing a basename:
000000_RUSLAN.wav… See the full description on the dataset page: https://huggingface.co/datasets/stilletto/ruslan-stressed-mini.
