datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
biggest-ru-bookA bigger version of its5Q/bigger-ru-book, the smaller set being a subset of this one. Almost 1000 hours of high-quality audio.
wikipedia_ruRU-AI-noise
RU-AI: A Large Multimodal Dataset for Machine Generated Content Detection
This is the noise agumented data for paper: RU-AI: A Large Multimodal Dataset for Machine Generated Content Detection
The original dataset is avaliable at zenodo:
https://zenodo.org/records/11406538
The official repo is avaliable at:
https://github.com/ZhihaoZhang97/RU-AI
Reference
We are appreciated the open-source community for the datasets and the models.
Microsoft COCO: Common Objects in… See the full description on the dataset page: https://huggingface.co/datasets/zzha6204/RU-AI-noise.common_voice_21_ru
Dataset Description
Набор данных validated.tsv отфильтрованный по down_votes = 0
📊 Статистика датасета
Информация по сплитам
🔹 Тренировочный набор (train)
Метрика
Значение
Количество семплов
93,531
Общая продолжительность
132.25 часов (476,089.70 секунд)
Средняя продолжительность семпла
5.09 секунд
🔹 Валидационный набор (validate)
Метрика
Значение
Количество семплов
38,836
Общая продолжительность
55.21… See the full description on the dataset page: https://huggingface.co/datasets/Sh1man/common_voice_21_ru.ru-book-mix-10h
ru-book-mix-10h
A 10-hour synthetic Russian-audiobook diarization benchmark. 600 one-minute
FLAC clips (16 kHz mono, 16-bit, lossless) with NIST RTTM ground truth, generated by
mexus/diarization-benchmark
from its5Q/biggest-ru-book (speech) and bilguun/musan-noise (background).
Intended use: diarization evaluation only. This dataset is not
suitable for training — the same source voices repeat across files, so any
model that trains on it will leak voice identity into its test split.… See the full description on the dataset page: https://huggingface.co/datasets/mexus/ru-book-mix-10h.ruslan-stressed
RUSLAN with Word Stress Marks · RUSLAN с проставленными ударениями
English / Русский
English
What is this?
A drop-in replacement for the metadata of the RUSLAN Russian single-speaker
TTS corpus, with word-stress marks added to every multi-syllabic Russian
word in the transcripts. Audio is bundled unchanged.
The motivation is to train Russian TTS models (e.g. Kokoro, Tacotron, VITS,
StyleTTS, XTTS) that pronounce words with correct lexical stress.
Vanilla… See the full description on the dataset page: https://huggingface.co/datasets/stilletto/ruslan-stressed.rudevices
📊 Сводная статистика аудио-датасетов
📈 Общая статистика по всем датасетам
Метрика
Значение
Всего датасетов/сабсетов
2
Всего семплов
296,394
Общая продолжительность
369.14 часов (1328901.86 секунд)
Средняя продолжительность семпла
4.48 секунд
Распределение объема данных по датасетам
ru_audiobooks_devices ███████████████████████████ 68.5%
rudevices_audio_records ████████████ 31.5%
Датасет: rudevices_audio_records… See the full description on the dataset page: https://huggingface.co/datasets/Sh1man/rudevices.bigger-ru-bookbatik-processedSpatialQA-ESpatialQA-E is a robot manipulation dataset focusing on spatial relationship understanding.
Paper:
https://arxiv.org/abs/2406.13642
GitHub repo:
https://github.com/BAAI-DCAI/SpatialBot
SpatialBot-general QA, a VLM with precise depth understanding:
https://huggingface.co/RussRobin/SpatialBot
SpatialBench, the spatial understanding benchmark in general QA:
https://huggingface.co/datasets/RussRobin/SpatialBench
PhD-webdataset
PhD Webdataset
This repository contains the packaged version of PhD. For a detailed introduction to PhD, please visit the official website.
Overview
The PhD Webdataset is designed to facilitate easy access and usage of the PhD dataset. It includes various fields in 'json' key. The data in this repo is totally the same as in PhD.
Installation
Ensure you have Hugging Face's datasets library installed. You can install it via pip:
pip install datasets… See the full description on the dataset page: https://huggingface.co/datasets/AIMClab-RUC/PhD-webdataset.rulibrispeech
Description
wav 16kHz
📊 Статистика датасета
Информация по сплитам
🔹 Тренировочный набор (train)
Метрика
Значение
Количество семплов
54,472
Общая продолжительность
92.79 часов (334028.33 секунд)
Средняя продолжительность семпла
6.13 секунд
🔹 Валидационный набор (validate)
Метрика
Значение
Количество семплов
1,400
Общая продолжительность
2.81 часов (10105.46 секунд)
Средняя продолжительность семпла
7.22… See the full description on the dataset page: https://huggingface.co/datasets/Sh1man/rulibrispeech.rubbishicm-data-tempABO-Image-ARES
ABO-Image-ARES
Dataset from the paper "RESAnything: Attribute Prompting for Arbitrary Referring Segmentation" (NeurIPS 2025).
Paper
arXiv: 2505.02867
Hugging Face Paper Page: hf.co/papers/2505.02867
Project Page: suikei-wang.github.io/RESAnything
Code: github.com/suikei-wang/RESAnything
Citation
@inproceedings{wang2025neurips-resanything,
title={{RESAnything: Attribute Prompting for Arbitrary Referring Segmentation}},
author={Wang, Ruiqi and Zhang… See the full description on the dataset page: https://huggingface.co/datasets/ruiqiw/ABO-Image-ARES.rukopys-curated-mvp-v2
RUKOPYS Curated MVP: Ukrainian Handwriting Recognition Dataset
RUKOPYS Curated MVP is a cleaned, task-ready derivative of
UkrainianCatholicUniversity/rukopys for Ukrainian handwritten
document AI. It turns the raw RUKOPYS release into reproducible artifacts for page-level
vision-language fine-tuning, crop-level transcription, and layout detection.
This dataset is designed for practical HTR work: train a model, inspect the normalized records,
evaluate layout/text extraction, and… See the full description on the dataset page: https://huggingface.co/datasets/AlexandreSheva/rukopys-curated-mvp-v2.rukopys-curated-mvp
RUKOPYS Curated MVP: Ukrainian Handwriting Recognition Dataset
Task-ready curated derivative of UkrainianCatholicUniversity/rukopys for Ukrainian handwritten document AI.
This dataset turns the raw RUKOPYS release into reproducible training artifacts for full-page vision-language fine-tuning, crop-level handwriting transcription, and layout detection. It was created to support an end-to-end HTR pipeline: curation, model training, inference, evaluation.
What This… See the full description on the dataset page: https://huggingface.co/datasets/AlexandreSheva/rukopys-curated-mvp.Conan-91kruslan-stressed-mini
RUSLAN stressed — mini sanity-check dataset
This is a 200-sample mini version of stilletto/ruslan-stressed used to
verify that the WebDataset tar layout is parsed correctly by the HuggingFace
dataset viewer before the full 22,200-sample dataset is repacked the same way.
Layout (WebDataset):
mini_part_001.tar # samples 000000…000099 (wav + paired txt)
mini_part_002.tar # samples 000100…000199 (wav + paired txt)
Each tar contains paired files sharing a basename:
000000_RUSLAN.wav… See the full description on the dataset page: https://huggingface.co/datasets/stilletto/ruslan-stressed-mini.physense_carla_dataset
PhySense CARLA Synthesized Dataset
Dataset Summary
This dataset is the CARLA-synthesized traffic dataset used in PhySense: Defending Physically Realizable Attacks for Autonomous Systems via Consistency Reasoning, CCS ’24. It is generated using the CARLA simulator and CARLA PythonAPI.
Source / Collection
The dataset is synthesized in CARLA. Instructions and scripts to reproduce/collect the dataset are provided in the accompanying GitHub repository:
Collection… See the full description on the dataset page: https://huggingface.co/datasets/Ruoyao/physense_carla_dataset.rust2RUOK_ver_samqa-rust2283884ALFREDsft_rulerust1OCR-datasetRuddyInParis
