datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sports-trends-dataset
⚽🏀🎾🏏 Sports-Trends Dataset
A leakage-safe, multi-sport match data lake — raw fixtures → engineered features → training splits.
The data backbone of Ruslan Magana Sports Intelligence — refreshed automatically every day.
TL;DR — A continuously-updated, medallion-architecture data lake for football,
basketball, tennis and cricket: immutable raw ingests, cleaned/standardized layers, an
engineered feature store, and ready-to-train chronological splits in… See the full description on the dataset page: https://huggingface.co/datasets/ruslanmv/sports-trends-dataset.ai-medical-chatbot
AI Medical Chatbot Dataset
This is an experimental Dataset designed to run a Medical Chatbot
It contains at least 250k dialogues between a Patient and a Doctor.
Playground ChatBot
ruslanmv/AI-Medical-Chatbot
For furter information visit the project here:
https://github.com/ruslanmv/ai-medical-chatbot
wikiomnia
Dataset Card for "Wikiomnia"
Dataset Summary
We present the WikiOmnia dataset, a new publicly available set of QA-pairs and corresponding Russian Wikipedia article summary sections, composed with a fully automated generative pipeline. The dataset includes every available article from Wikipedia for the Russian language. The WikiOmnia pipeline is available open-source and is also tested for creating SQuAD-formatted QA on other domains, like news texts, fiction, and social… See the full description on the dataset page: https://huggingface.co/datasets/RussianNLP/wikiomnia.musixmulti_SWE_Bench_Rust
multi_SWE_Bench_Rust
数据集描述...
coat
Dataset Card for CoAT🧥
Dataset Description
CoAT🧥 (Corpus of Artificial Texts) is a large-scale corpus for Russian, which consists of 246k human-written texts from publicly available resources and artificial texts generated by 13 neural models, varying in the number of parameters, architecture choices, pre-training objectives, and downstream applications.
Each model is fine-tuned for one or more of six natural language generation tasks, ranging from paraphrase generation… See the full description on the dataset page: https://huggingface.co/datasets/RussianNLP/coat.RusLawOD
The Russian Legislative Corpus, 1991–2026
Russian primary and secondary legislation corpus covering laws of Russian Federation, decrees by the President of RF, regulations by the government published as of July, 2026. The corpus collects all 308,056 texts (198,777,737 tokens) of non-secret federal regulations and acts, along with their metadata. The corpus has two versions: the original text with minimal preprocessing and a version prepared for linguistic analysis with… See the full description on the dataset page: https://huggingface.co/datasets/irlspbru/RusLawOD.the-stack-rust-clean
Dataset 1: TheStack - Rust - Cleaned
Description: This dataset is drawn from TheStack Corpus, an open-source code dataset with over 3TB of GitHub data covering 48 programming languages. We selected a small portion of this dataset to optimize smaller language models for Rust, a popular statically typed language.
Target Language: Rust
Dataset Size:
Training: 900,000 files
Validation: 50,000 files
Test: 50,000 files
Preprocessing:
Selected Rust as the target language due to its… See the full description on the dataset page: https://huggingface.co/datasets/ammarnasr/the-stack-rust-clean.ru-big-russian-dataset
Big Russian Dataset
Made by ZeroAgency.ru - telegram channel.
Dataset size
Train: 1 710 601 samples (filtered from 2_149_360)
Test: 18 520 samples (not filtered)
English
The Big Russian Dataset is a combination of various primarily Russian‑language datasets. With some sort of reasoning!
The dataset was deduplicated, cleaned, scored using gpt-4.1 and filtered.
Русский
Big Russian Dataset - большой русский датасет. Комбинация из… See the full description on the dataset page: https://huggingface.co/datasets/ZeroAgency/ru-big-russian-dataset.russian_dialoguesДатасет русских диалогов собранных с Telegram чатов.
Диалоги имеют разметку по релевантности.
Также были сгенерированы негативные примеры с помощью перемешивания похожих ответов.
Количество диалогов - 2 миллиона
Формат датасета:
{
'question': 'Привет',
'answer': 'Привет, как дела?'
'relevance': 1
}
Программа парсинга: https://github.com/Den4ikAI/telegram_chat_parser
Citation:
@MISC{russian_instructions,
author = {Denis Petrov},
title = {Russian dialogues… See the full description on the dataset page: https://huggingface.co/datasets/Den4ikAI/russian_dialogues.Kazakh-Russian-Child-Directed-Speech-Corpus
Kazakh-Russian Child-Directed Language Corpus
Dataset Description
The corpus combines narrative texts and child-adult dialogue transcripts in Kazakh and Russian. It was created to support research on low-resource NLP, language acquisition, language modeling, and morphologically aware tokenization.
The corpus includes narrative materials such as fairy tales, children’s literature, cartoons/subtitles, translated stories, and educational texts.
This dataset is a work… See the full description on the dataset page: https://huggingface.co/datasets/esimijoq/Kazakh-Russian-Child-Directed-Speech-Corpus.assemblage-rust
Assemblage-Rust
Consider using your coding agent to test, download and process the data, repository size
is 1.5T and it is highly likely you only want a portion of it, but do remember to
check the agent outputs.
Produced by Assemblage, a distributed
binary-corpus generator, a cite would be greatly appreciated!
Teh dataset contains127,165 compiled Rust binaries from 77,004 builds
of 10,766 permissively
licensed GitHub repositories, each paired with DWARF-derived function and… See the full description on the dataset page: https://huggingface.co/datasets/changliu8541/assemblage-rust.russia_voicesRussian voices for train AI.
♀ Male and ♂ Female
Male voices - 497 pcs.
Female voices - 244 pcs.
Prepared for training fish-speech
The author is not responsible for the votes.
Use at your own risk.
license: apache-2.0
task_categories:
- zero-shot-classification
language:
- ru
size_categories:
- 1B<n<10B
rust-the-stack-v2SPEEED_s3_words_russian_0k-300krussian_librispeech
Russian LibriSpeech (RuLS)
Identifier: SLR96 from openslr.org
Summary: This dataset is based on LibriVox audiobooks
Category: Speech
License: The dataset is Public Domain in the USA.
About this resource:
Russian LibriSpeech (RuLS) dataset is based on LibriVox's public domain audio books (see BOOKS.TXT for the list of included books) and contains about 98 hours of audio data.
russian_super_glueRecent advances in the field of universal language models and transformers require the development of a methodology for
their broad diagnostics and testing for general intellectual skills - detection of natural language inference,
commonsense reasoning, ability to perform simple logical operations regardless of text subject or lexicon. For the first
time, a benchmark of nine tasks, collected and organized analogically to the SuperGLUE methodology, was developed from
scratch for the Russian language. We provide baselines, human level evaluation, an open-source framework for evaluating
models and an overall leaderboard of transformer models for the Russian language.russian-old-orthography-ocr
Basic Description
Dataset contains source images and human-readable extracted texts. All texts were published in Russia in the 19th century and written using pre-reform orthography.
The dataset is designed to train and evaluate optical character recognition systems for texts published in Russian before the orthographic reform (1917).
Data structure
For each text there is a file with its image and the text corresponding to this image. The names of these files are the same… See the full description on the dataset page: https://huggingface.co/datasets/nevmenandr/russian-old-orthography-ocr.SPEEED_s3_words_russian_300k-600kmultiswe_rustbenchrublimp
RuBLiMP
Dataset Description
RuBLiMP, or Russian Benchmark of Linguistic Minimal Pairs, is the first diverse and large-scale benchmark of minimal pairs in Russian.
RuBLiMP includes 45k minimal pairs of sentences that differ in grammaticality and isolate morphological, syntactic, or semantic phenomena. In contrast to existing benchmarks of linguistic minimal pairs, RuBLiMP is created by applying linguistic perturbations to automatically annotated sentences from open text… See the full description on the dataset page: https://huggingface.co/datasets/RussianNLP/rublimp.russian_dialogues_2
Den4ikAI/russian_dialogues_2
Датасет русских диалогов для обучения диалоговых моделей.
Количество диалогов - 1.6 миллиона
Формат датасета:
{
'sample': ['Привет', 'Привет', 'Как дела?']
}
Citation:
@MISC{russian_instructions,
author = {Denis Petrov},
title = {Russian context dialogues dataset for conversational agents},
url = {https://huggingface.co/datasets/Den4ikAI/russian_dialogues_2},
year = 2023
}
tapeThe Winograd schema challenge composes tasks with syntactic ambiguity,
which can be resolved with logic and reasoning (Levesque et al., 2012).
The texts for the Winograd schema problem are obtained using a semi-automatic
pipeline. First, lists of 11 typical grammatical structures with syntactic
homonymy (mainly case) are compiled. For example, two noun phrases with a
complex subordinate: 'A trinket from Pompeii that has survived the centuries'.
Requests corresponding to these constructions are submitted in search of the
Russian National Corpus, or rather its sub-corpus with removed homonymy. In the
resulting 2+k examples, homonymy is removed automatically with manual validation
afterward. Each original sentence is split into multiple examples in the binary
classification format, indicating whether the homonymy is resolved correctly or
not.Multi-SWE-smith-Rust-GLM-4.6-trajectorieslearn-rust
Knowledge base from the Rust books
Gaia node setup instructions
See the Gaia node getting started guide
gaianet init --config https://raw.githubusercontent.com/GaiaNet-AI/node-configs/main/llama-3.1-8b-instruct_rustlang/config.json
gaianet start
with the full 128k context length of Llama 3.1
gaianet init --config https://raw.githubusercontent.com/GaiaNet-AI/node-configs/main/llama-3.1-8b-instruct_rustlang/config.fullcontext.json
gaianet start
Steps to create… See the full description on the dataset page: https://huggingface.co/datasets/gaianet/learn-rust.russian_instructions_2June 10:
Почищены криво переведенные примеры кода
Добавлено >50000 человеческих примеров QA и инструкций
Обновленная версия русского датасета инструкций и QA.
Улучшения:
1. Увеличен размер с 40 мегабайт до 130 (60к сэмплов - 200к)
2. Улучшено качество перевода.
Структура датасета:
{
"sample":[
"Как я могу улучшить свою связь между телом и разумом?",
"Начните с разработки регулярной практики осознанности. 2. Обязательно практикуйте баланс на нескольких уровнях: физическом… See the full description on the dataset page: https://huggingface.co/datasets/Den4ikAI/russian_instructions_2.Russian_Documents_Dataset_PDF
Russian Documents Dataset (PDF)
This dataset contains a curated collection of Russian-language documents in PDF format. The corpus includes books, academic papers, government publications, articles, and educational materials written in Russian. It is designed to support AI research in OCR, document understanding, and multilingual text recognition.
Contact
For queries or collaborations related to this dataset, contact:
anoushka@kgen.io
abhishek.vadapalli@kgen.io… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Russian_Documents_Dataset_PDF.audio_data_russian_backup
Dataset Audio Russian Backup
This is a backup dataset with Russian audio data, split into train_0 to train_49 for tasks like text-to-speech, speech recognition, and speaker identification.
Features
text: Audio transcription (string).
speaker_name: Speaker identifier (string).
audio: Audio file.
Usage
Load the dataset like this:
from datasets import load_dataset
dataset = load_dataset("kijjjj/audio_data_russian_backup", split="train_0") # Or any train_X… See the full description on the dataset page: https://huggingface.co/datasets/kijjjj/audio_data_russian_backup.AgentArk
AgentArk Assets
Versioned Unity runtimes, task Mods, and multimodal replay/evaluation records.
Environment compatibility / 环境兼容性
A task's labeled env version is its minimum supported environment version.
That version and all later versions are compatible, with no upper version bound.
任务标注的 env 版本是最低环境版本;该版本及之后的所有版本均兼容,没有版本上限。
Task artifact label / 任务标注
Compatible environments / 可用环境
env-1.0.3
>=1.0.3: 1.0.3, 1.0.4, 1.0.5, and all later versions… See the full description on the dataset page: https://huggingface.co/datasets/P90-RushB/AgentArk.rustbench
