datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Kazakh-Russian-Child-Directed-Speech-Corpus
Kazakh-Russian Child-Directed Language Corpus
Dataset Description
The corpus combines narrative texts and child-adult dialogue transcripts in Kazakh and Russian. It was created to support research on low-resource NLP, language acquisition, language modeling, and morphologically aware tokenization.
The corpus includes narrative materials such as fairy tales, children’s literature, cartoons/subtitles, translated stories, and educational texts.
This dataset is a work… See the full description on the dataset page: https://huggingface.co/datasets/esimijoq/Kazakh-Russian-Child-Directed-Speech-Corpus.ru-big-russian-dataset
Big Russian Dataset
Made by ZeroAgency.ru - telegram channel.
Dataset size
Train: 1 710 601 samples (filtered from 2_149_360)
Test: 18 520 samples (not filtered)
English
The Big Russian Dataset is a combination of various primarily Russian‑language datasets. With some sort of reasoning!
The dataset was deduplicated, cleaned, scored using gpt-4.1 and filtered.
Русский
Big Russian Dataset - большой русский датасет. Комбинация из… See the full description on the dataset page: https://huggingface.co/datasets/ZeroAgency/ru-big-russian-dataset.russian_super_glueRecent advances in the field of universal language models and transformers require the development of a methodology for
their broad diagnostics and testing for general intellectual skills - detection of natural language inference,
commonsense reasoning, ability to perform simple logical operations regardless of text subject or lexicon. For the first
time, a benchmark of nine tasks, collected and organized analogically to the SuperGLUE methodology, was developed from
scratch for the Russian language. We provide baselines, human level evaluation, an open-source framework for evaluating
models and an overall leaderboard of transformer models for the Russian language.russian_dialogues_2
Den4ikAI/russian_dialogues_2
Датасет русских диалогов для обучения диалоговых моделей.
Количество диалогов - 1.6 миллиона
Формат датасета:
{
'sample': ['Привет', 'Привет', 'Как дела?']
}
Citation:
@MISC{russian_instructions,
author = {Denis Petrov},
title = {Russian context dialogues dataset for conversational agents},
url = {https://huggingface.co/datasets/Den4ikAI/russian_dialogues_2},
year = 2023
}
russian-it-community-corpus
📦 Russian IT Community Corpus (RICC)
Russian IT Community Corpus (RICC) is an open, de-identified conversational dataset collected from 11 engineering community nodes spanning a 9-year timeline (2017–2026). It captures authentic discussions on backend systems, cloud infrastructure, AI/ML deployment, database internals, and software architecture.
The corpus is structured into ready-to-use splits for Instruction Fine-Tuning (SFT), Direct Preference Optimization (DPO)… See the full description on the dataset page: https://huggingface.co/datasets/wwewtech/russian-it-community-corpus.khakas-russian-parallel-corpus
Khakas-Russian Parallel Corpus
The creation of this dataset is aimed at supporting the development of natural language processing (NLP) tools and
machine translation for the Khakas language, which is classified as a "Definitely Endangered" language. By providing
high-quality parallel data, this project helps preserve the linguistic heritage of the Khakas people.
Dataset Overlap:
The Khakas sentences in this corpus do not overlap with those in the Khakas… See the full description on the dataset page: https://huggingface.co/datasets/adeshkin/khakas-russian-parallel-corpus.khakas-russian-dict
Khakas-Russian Dictionary (Dataset)
Dataset Developer: Vasily AdeshkinContact for inquiries: adeshkin.vi@phystech.edu
📌 Important Notice & Citation
When using this dataset, please be sure to cite this repository and the original dictionary: https://khakas.altaica.ru/dictionary/.
Please note: Some optical character recognition (OCR) errors may still be present in the data.
🛠 Contribution & Authorship
I am not the author of the original dictionary, but I have… See the full description on the dataset page: https://huggingface.co/datasets/adeshkin/khakas-russian-dict.physics-russian
Оглавление
Описание датасета
Аннотация
Ключевые особенности
Статус перевода
Методология перевода и верификации
Ограничения и возможные погрешности
Структура датасета
Поля данных
Использование
Благодарности
Лицензирование и авторские права
Цитирование
📑 Оглавление Примеров
Нажмите на тему, чтобы перейти к соответствующему разделу в полном отчёте.
Квантовая механика
Термодинамика
Электромагнетизм
Общая теория относительности
Специальная теория относительности
Атомная… See the full description on the dataset page: https://huggingface.co/datasets/AITISPEC/physics-russian.russian-nmo-medical-mcq
Russian NMO Medical MCQ
Choose language / Выберите язык: Русский | English
Русский
Это датасет русскоязычных медицинских тестовых вопросов НМО с вариантами ответа.
В нем есть вопросы с одним правильным вариантом и вопросы с несколькими правильными
вариантами. Датасет подготовлен так, чтобы его можно было сразу использовать для
тонкой настройки LLM, проверки качества ответов и экспериментов с медицинским QA.
Главная идея простая: дать модели вопрос, тему и варианты ответа… See the full description on the dataset page: https://huggingface.co/datasets/drkolesnikov/russian-nmo-medical-mcq.russian_literature
RusLit Corpus
A corpus of Russian literature in clean text format.
Description
This dataset contains cleaned text (stripped of extraneous artifacts) collected from works by authors who passed away more than 70 years ago, placing them in the public domain.
Some texts may contain semantic nonsense (e.g. OCR or digitization artifacts), as well as fragments of French, German, English, or Japanese text mixed in with the Russian.
Structure
Each record… See the full description on the dataset page: https://huggingface.co/datasets/RafaelUI/russian_literature.GSM8k_Kazakh_Russian
Dataset Summary
These are the machine-translated Kazakh and Russian versions of the GSM8K (Grade School Math 8K) dataset (test set).
This dataset is used to test the mathematical reasoning of large language models in the Kazakh language. Specifically, GSM8K focuses on high-quality grade school math word problems that require multi-step reasoning to solve. Kazakh and Russian versions serve as a benchmark for evaluating how well models can perform mathematical reasoning and… See the full description on the dataset page: https://huggingface.co/datasets/issai/GSM8k_Kazakh_Russian.russian_jokesrussian_instructions_2_cleaned
Russian Instructions Cleaned
Очищенная версия Den4ikAI/russian_instructions_2.
Что сделано
Дедупликация по question (удалено ~45k)
Удалены пустые question и answer
Удалены ответы короче 100 и длиннее 4000 символов
Конвертировано в chat-формат (messages: user/assistant)
Статистика
Метрика
Значение
Исходно
237 281
После чистки
138 973
Удалено
98 308 (41%)
Формат
JSONL, одна строка = один пример.
{"messages":… See the full description on the dataset page: https://huggingface.co/datasets/Lev384501/russian_instructions_2_cleaned.russian_dpo_qa
Format
Each row contains:
prompt
chosen
rejected
Usage
from datasets import load_dataset
dataset = load_dataset("eridai/russian_dpo_qa")
train = dataset["train"]
russian-easy-instructions
Easy Russian Instructions
Fast instructions in conversational format, generaly used for injecting general knowledge.Dataset contains comprehensive instructions for easy tasks and general question answering
Contents:
Wikipedia / Factological knowledge
History knowledge
Basic programming understanding
Basic math understanding
Basic physics understanding
Basic geography knowledge
Basic biology knowledge
Format:
Formatted for instruction-tuning / instruction… See the full description on the dataset page: https://huggingface.co/datasets/attn-signs/russian-easy-instructions.keyphrase_extraction_russianA dataset for extracting or generating keywords for scientific texts in Russian. It contains annotations of scientific articles from four domains: mathematics and computer science (a fragment of the Keyphrases CS&Math Russian dataset), history, linguistics, and medicine. For more details about the dataset please refer the original paper: https://arxiv.org/abs/2409.10640
Dataset Structure
abstract, an abstract in a string format;
keywords, a list of keyphrases provided by… See the full description on the dataset page: https://huggingface.co/datasets/aglazkova/keyphrase_extraction_russian.mem-and-russian-jokes-dataset
2 июля 2025
Добавлено новых уникальных анекдотов: 1311963
Количество записей в датасете: 521904
Добавил датасет анекдотов от IgorVolochay/russian_jokes
не понимаю как я прошел мимо него, там очень много шуток,
сам Евгений Ваганович очень часто обращался к этому датасету... Но вот и мое время пришло.
Плюс обновил старый датасет и разбавил формулировки новыми начальными вопросами
human_prompts = [
"Расскажи шутку",
"Расскажи анекдот",
"Знаешь какой-нибудь прикол?",
"Скажи что-нибудь смешное"… See the full description on the dataset page: https://huggingface.co/datasets/samedad/mem-and-russian-jokes-dataset.RussianFinancialNews
RussianFinancialNews
Датасет содержит 92,377 русскоязычных новостных статей на финансовую тематику, преимущественно про российский рынок ценных бумаг и российскую экономику. Набор данных может быть полезен для разных задач обработки естественного языка (NLP).
A dataset containing 92,377 samples of Russian financial news articles. Each sample includes metadata and content fields that are useful for various Natural Language Processing (NLP) tasks.
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/Kasymkhan/RussianFinancialNews.Dataset_of_Russian_thinkingRu
RTD
Описание:Russian Thinking Dataset — это набор данных, предназначенный для обучения и тестирования моделей обработки естественного языка (NLP) на русском языке. Датасет ориентирован на задачи, связанные с генерацией текста, анализом диалогов и решением математических и логических задач.
Основная информация:
Сплит: train
Количество записей: 147.046
Цели:
Обучение моделей пониманию русского языка.
Создание диалоговых систем с естественным взаимодействием.… See the full description on the dataset page: https://huggingface.co/datasets/qwqeqw/Dataset_of_Russian_thinking.russian-colloquial-sft-50k
Russian colloquial SFT (50k, mat-free)
English: ~50,000 chat SFT examples: rewrite formal Russian into casual Telegram-style Russian without profanity.
Format
Each line is JSON with a messages array (ShareGPT / TRL):
{
"messages": [
{
"role": "user",
"content": "Перепиши простым разговорным русским, как в переписке. Без мата и грубости. Сохрани смысл:\n<формальный текст>"
},
{
"role": "assistant",
"content": "<разговорный ответ>"… See the full description on the dataset page: https://huggingface.co/datasets/pavelfedortsov/russian-colloquial-sft-50k.russian-everyday-dialogues
Russian Everyday Dialogues / Русские повседневные диалоги
A dataset of 20 natural conversational Russian dialogues covering everyday situations.
Создан носителем языка с 18-летним опытом работы в международной среде.
Dataset Description
This dataset contains short, natural Russian dialogues typical of everyday urban life in Russia.
Each example reflects authentic spoken language patterns, not formal or literary Russian.
Situations covered
🛒 Магазин / Shopping… See the full description on the dataset page: https://huggingface.co/datasets/kukunechka/russian-everyday-dialogues.Corrupted-russian-word-documents-text-datasetThis is synthetic text-corruption dataset, based on Russian open official/government documents text.
This dataset is intended to be used to train LLM to perform text-recovery task.
All the errors in text is made solely in Russian sentences, hence ignoring any English sentence.
Texts contains complex formatting, which is common for documents.
Each line contains json object that have array messages value, which consists of role-based conversation.
Each first message is randomly chosen system… See the full description on the dataset page: https://huggingface.co/datasets/Kemsekov/Corrupted-russian-word-documents-text-dataset.telegram-corpus-russian-kazakh
📦 Telegram Corpus KAZ_RU
Telegram Corpus KAZ_RU is a raw multilingual dataset of Telegram messages in Russian and Kazakh, collected and assembled by Kurumikz. It contains over 1.4 million lines of informal, user-generated text extracted from a large .txt dump. The corpus is intended for experimentation in text generation, language modeling, and dialogue research.
📊 Dataset Statistics
Metric
Value
Description
📄 Lines
1,493,124
Total number of lines/messages… See the full description on the dataset page: https://huggingface.co/datasets/kurumikz/telegram-corpus-russian-kazakh.ru-big-russian-dataset
Big Russian Dataset
Made by ZeroAgency.ru - telegram channel.
Dataset size
Train: 1 710 601 samples (filtered from 2_149_360)
Test: 18 520 samples (not filtered)
English
The Big Russian Dataset is a combination of various primarily Russian‑language datasets. With some sort of reasoning!
The dataset was deduplicated, cleaned, scored using gpt-4.1 and filtered.
Русский
Big Russian Dataset - большой русский датасет. Комбинация из… See the full description on the dataset page: https://huggingface.co/datasets/jebkoralav/ru-big-russian-dataset.Nebo-T1-Russian
Russian Description (English below)
UPD: Dataset reuploaded, correct_format column added
Nebo-T1-Russian
(Вероятно) первый "longCoT" датасет для русского языка, созданный через Deeseek-R1
Подсказки взяты из датасета Sky-T1 и переведены через Llama3.3-70B
Ответы и рассуждения сгенерированные Deeseek-R1 (685B)
16.4K сэмплов в целом, ≈12.4K только с русским языком (в остальных либо ответ, либо рассуждения на английском)
Языки в ответе и рассуждениях размечены… See the full description on the dataset page: https://huggingface.co/datasets/kristaller486/Nebo-T1-Russian.MATH-500-Russian
Карточка датасета MATH-500-Russian
Перевод датасета HuggingFaceH4/MATH-500 на русский язык,
был выполнен моделью qwen2.5:32b через
скрипты EvilFreelancer/datasets-translator.
Данный набор данных содержит подмножество из 500 задач из теста MATH, который OpenAI создал для статьи Let's Verify
Step by Step и переведённых на русский язык.
Подробности в их репозиторий на GitHub.
russian_oil_gas_news_telegram_dataset
Description in English:
Dataset collected from 30 Russian-language Telegram news channels on the topic of Oil ang Gas Industry,
collected and marked up automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link.
Dataset fields:
taskId - task identifier in the Scouti service. text - main text. url - link to the publication. sourceLink - link to Telegram. subSourceLink - link to the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/russian_oil_gas_news_telegram_dataset.WanJuan-Russian
💡 Introduction
WanJuan-Russian(万卷丝路-俄语)corpus, with a volume exceeding 280GB, comprises 7 major categories and 34 subcategories. It covers a wide range of local-specific content, including history, politics, culture, real estate, shopping, weather, dining, encyclopedias, and professional knowledge. The rich thematic classification not only facilitates researchers in retrieving data according to specific needs but also ensures that the corpus can adapt to diverse research… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/WanJuan-Russian.scoutieDataset_russian_language_grammar_and_rules_vectorized
Description in English:
A dataset collected from 30 Russian-language Telegram channels on the topic of learning the Russian language. This dataset contains grammar, syntax, spelling and punctuation rules.
The dataset was collected and marked automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link.
Dataset fields:
taskId - task identifier in the Scouti service. text - main text. url -… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/scoutieDataset_russian_language_grammar_and_rules_vectorized.Russian_spell_dataset
Russian Spell Dataset
Датасет для обучения и дообучения моделей исправления ошибок автоматического распознавания русской речи.
Содержание
Датасет содержит 1 139 пар текстов в формате JSONL:
input — текст с опечатками, ошибками распознавания речи, пропущенной пунктуацией или неправильным регистром;
output — исправленная версия текста.
Также есть примеры без ошибок, где input совпадает с output. Это помогает модели не вносить исправления в уже корректный текст.… See the full description on the dataset page: https://huggingface.co/datasets/Embim/Russian_spell_dataset.
