datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
biggest-ru-bookA bigger version of its5Q/bigger-ru-book, the smaller set being a subset of this one. Almost 1000 hours of high-quality audio.
RU-AI-noise
RU-AI: A Large Multimodal Dataset for Machine Generated Content Detection
This is the noise agumented data for paper: RU-AI: A Large Multimodal Dataset for Machine Generated Content Detection
The original dataset is avaliable at zenodo:
https://zenodo.org/records/11406538
The official repo is avaliable at:
https://github.com/ZhihaoZhang97/RU-AI
Reference
We are appreciated the open-source community for the datasets and the models.
Microsoft COCO: Common Objects in… See the full description on the dataset page: https://huggingface.co/datasets/zzha6204/RU-AI-noise.common_voice_21_ru
Dataset Description
Набор данных validated.tsv отфильтрованный по down_votes = 0
📊 Статистика датасета
Информация по сплитам
🔹 Тренировочный набор (train)
Метрика
Значение
Количество семплов
93,531
Общая продолжительность
132.25 часов (476,089.70 секунд)
Средняя продолжительность семпла
5.09 секунд
🔹 Валидационный набор (validate)
Метрика
Значение
Количество семплов
38,836
Общая продолжительность
55.21… See the full description on the dataset page: https://huggingface.co/datasets/Sh1man/common_voice_21_ru.ru-book-mix-10h
ru-book-mix-10h
A 10-hour synthetic Russian-audiobook diarization benchmark. 600 one-minute
FLAC clips (16 kHz mono, 16-bit, lossless) with NIST RTTM ground truth, generated by
mexus/diarization-benchmark
from its5Q/biggest-ru-book (speech) and bilguun/musan-noise (background).
Intended use: diarization evaluation only. This dataset is not
suitable for training — the same source voices repeat across files, so any
model that trains on it will leak voice identity into its test split.… See the full description on the dataset page: https://huggingface.co/datasets/mexus/ru-book-mix-10h.ruslan-stressed
RUSLAN with Word Stress Marks · RUSLAN с проставленными ударениями
English / Русский
English
What is this?
A drop-in replacement for the metadata of the RUSLAN Russian single-speaker
TTS corpus, with word-stress marks added to every multi-syllabic Russian
word in the transcripts. Audio is bundled unchanged.
The motivation is to train Russian TTS models (e.g. Kokoro, Tacotron, VITS,
StyleTTS, XTTS) that pronounce words with correct lexical stress.
Vanilla… See the full description on the dataset page: https://huggingface.co/datasets/stilletto/ruslan-stressed.rudevices
📊 Сводная статистика аудио-датасетов
📈 Общая статистика по всем датасетам
Метрика
Значение
Всего датасетов/сабсетов
2
Всего семплов
296,394
Общая продолжительность
369.14 часов (1328901.86 секунд)
Средняя продолжительность семпла
4.48 секунд
Распределение объема данных по датасетам
ru_audiobooks_devices ███████████████████████████ 68.5%
rudevices_audio_records ████████████ 31.5%
Датасет: rudevices_audio_records… See the full description on the dataset page: https://huggingface.co/datasets/Sh1man/rudevices.bigger-ru-bookrulibrispeech
Description
wav 16kHz
📊 Статистика датасета
Информация по сплитам
🔹 Тренировочный набор (train)
Метрика
Значение
Количество семплов
54,472
Общая продолжительность
92.79 часов (334028.33 секунд)
Средняя продолжительность семпла
6.13 секунд
🔹 Валидационный набор (validate)
Метрика
Значение
Количество семплов
1,400
Общая продолжительность
2.81 часов (10105.46 секунд)
Средняя продолжительность семпла
7.22… See the full description on the dataset page: https://huggingface.co/datasets/Sh1man/rulibrispeech.rubbishruslan-stressed-mini
RUSLAN stressed — mini sanity-check dataset
This is a 200-sample mini version of stilletto/ruslan-stressed used to
verify that the WebDataset tar layout is parsed correctly by the HuggingFace
dataset viewer before the full 22,200-sample dataset is repacked the same way.
Layout (WebDataset):
mini_part_001.tar # samples 000000…000099 (wav + paired txt)
mini_part_002.tar # samples 000100…000199 (wav + paired txt)
Each tar contains paired files sharing a basename:
000000_RUSLAN.wav… See the full description on the dataset page: https://huggingface.co/datasets/stilletto/ruslan-stressed-mini.
