datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
infini-news-corpus
INFINI-NEWS Corpus
🔎 Search this corpus online: query it with sub-second full-text search and n-gram counts — in the browser or via a public, keyless REST API, no download required — at infini-news.uni-graz.at (API reference).
A multilingual news corpus extracted from
Common Crawl CC-News WARC files.
One row per article, with body text extracted via
trafilatura,
WARC provenance, and derived metadata (publish date, language, topic,
byte hashes) in a single flat schema. Covers… See the full description on the dataset page: https://huggingface.co/datasets/ruggsea/infini-news-corpus.ChatGPT-Jailbreak-Prompts
Dataset Card for Dataset Name
Name
ChatGPT Jailbreak Prompts
Dataset Summary
ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT.
Languages
[English]
LLaVA-OneVision-Data-ru
LLaVA-OneVision-Data-ru
Translated lmms-lab/LLaVA-OneVision-Data dataset into Russian language using Google translate.
Almost all datasets have been translated, except for the following:
["tallyqa(cauldron,llava_format)", "clevr(cauldron,llava_format)", "VisualWebInstruct(filtered)", "figureqa(cauldron,llava_format)", "magpie_pro(l3_80b_mt)", "magpie_pro(qwen2_72b_st)", "rendered_text(cauldron)", "ureader_ie"]
Usage
import datasets
data =… See the full description on the dataset page: https://huggingface.co/datasets/d0rj/LLaVA-OneVision-Data-ru.Llama-3-SynE-Dataset
📄 Report | 💻 GitHub Repo
🔍 English | 简体中文
Here is the continual pre-training dataset. The Llama-3-SynE model is available here.
News
🌟🌟 2024/12/17: We released the code used for continual pre-training and data preparation. The code contains detailed documentation comments.
✨✨ 2024/08/12: We released the continual pre-training dataset.
✨✨ 2024/08/10: We released the Llama-3-SynE model.
✨ 2024/07/26: We released the technical report, welcome to check it… See the full description on the dataset page: https://huggingface.co/datasets/RUC-AIBOX/Llama-3-SynE-Dataset.MemoryDecoder-at-Scale-domain-data
MemoryDecoder at Scale Domain Data
This repository contains the domain-specific continued-pretraining (CPT) data,
the tokenized and preprocessed datasets, and the aligned KNN distributions used
by MemoryDecoder at Scale.
Links
Project Page: Memory Decoder at Scale
GitHub Repository: LUMIA-Group/MemoryDecoder-at-Scale
Paper: Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory
The preprocessed datasets and KNN distributions in this repository use… See the full description on the dataset page: https://huggingface.co/datasets/Rubin-Wei/MemoryDecoder-at-Scale-domain-data.RubricHub_v1
RubricHub
RubricHub is a large-scale (approximately 110K), multi-domain dataset that provides high-quality rubric-based supervision for open-ended generation tasks. It is constructed via an automated coarse-to-fine rubric generation framework, which integrates principle-guided synthesis, multi-model aggregation, and difficulty evolution to produce comprehensive and highly discriminative evaluation criteria, overcoming the supervision ceiling of… See the full description on the dataset page: https://huggingface.co/datasets/sojuL/RubricHub_v1.ru_turbo_alpaca
RuTurboAlpaca
Dataset of ChatGPT-generated instructions in Russian.
Code: rulm/self_instruct
Code is based on Stanford Alpaca and self-instruct.
29822 examples
Preliminary evaluation by an expert based on 400 samples:
83% of samples contain correct instructions
63% of samples have correct instructions and outputs
Crowdsouring-based evaluation on 3500 samples:
90% of samples contain correct instructions
68% of samples have correct instructions and outputs
Prompt template:… See the full description on the dataset page: https://huggingface.co/datasets/IlyaGusev/ru_turbo_alpaca.ru-fandom-wiki
d0rj/ru-fandom-wiki
Description
A set of texts collected from the most popular Russian-language fandoms (65 fandoms) on fandom.com.
The dump given on 25.10.2024-27.10.2024 collected using trafilatura library. All texts are in markdown format.
License
The license supports the license text on the source site - Creative Commons Attribution-ShareAlike 3.0 (Unported) (CC-BY-SA).
the-stack-rust-clean
Dataset 1: TheStack - Rust - Cleaned
Description: This dataset is drawn from TheStack Corpus, an open-source code dataset with over 3TB of GitHub data covering 48 programming languages. We selected a small portion of this dataset to optimize smaller language models for Rust, a popular statically typed language.
Target Language: Rust
Dataset Size:
Training: 900,000 files
Validation: 50,000 files
Test: 50,000 files
Preprocessing:
Selected Rust as the target language due to its… See the full description on the dataset page: https://huggingface.co/datasets/ammarnasr/the-stack-rust-clean.ru-big-russian-dataset
Big Russian Dataset
Made by ZeroAgency.ru - telegram channel.
Dataset size
Train: 1 710 601 samples (filtered from 2_149_360)
Test: 18 520 samples (not filtered)
English
The Big Russian Dataset is a combination of various primarily Russian‑language datasets. With some sort of reasoning!
The dataset was deduplicated, cleaned, scored using gpt-4.1 and filtered.
Русский
Big Russian Dataset - большой русский датасет. Комбинация из… See the full description on the dataset page: https://huggingface.co/datasets/ZeroAgency/ru-big-russian-dataset.Kazakh-Russian-Child-Directed-Speech-Corpus
Kazakh-Russian Child-Directed Language Corpus
Dataset Description
The corpus combines narrative texts and child-adult dialogue transcripts in Kazakh and Russian. It was created to support research on low-resource NLP, language acquisition, language modeling, and morphologically aware tokenization.
The corpus includes narrative materials such as fairy tales, children’s literature, cartoons/subtitles, translated stories, and educational texts.
This dataset is a work… See the full description on the dataset page: https://huggingface.co/datasets/esimijoq/Kazakh-Russian-Child-Directed-Speech-Corpus.Llama-HybridDiffusion-processed-data-run1
Llama-HybridDiffusion processed training mixture — run 1
Built with Llama.
This repository preserves the exact Hugging Face Dataset.save_to_disk Arrow snapshot
used by run 1 of a Qwen3.5-2B HybridDiffusion reproduction. The directory names,
dataset_info.json, state.json, and Arrow shard boundaries are retained so the data
can be downloaded and supplied to the existing training configuration without a lossy
format conversion.
Exact snapshot inventory
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Arushhh/Llama-HybridDiffusion-processed-data-run1.Rule-VLN
Rule-VLN Dataset
Rule-VLN is a rule-compliant outdoor vision-and-language navigation benchmark built on the Touchdown / StreetLearn urban navigation environment. It studies whether navigation agents can follow language instructions while also complying with semantic traffic rules, such as regulatory signs that prohibit otherwise reachable movements.
This dataset accompanies the paper:
Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric… See the full description on the dataset page: https://huggingface.co/datasets/jeffry77/Rule-VLN.russian_super_glueRecent advances in the field of universal language models and transformers require the development of a methodology for
their broad diagnostics and testing for general intellectual skills - detection of natural language inference,
commonsense reasoning, ability to perform simple logical operations regardless of text subject or lexicon. For the first
time, a benchmark of nine tasks, collected and organized analogically to the SuperGLUE methodology, was developed from
scratch for the Russian language. We provide baselines, human level evaluation, an open-source framework for evaluating
models and an overall leaderboard of transformer models for the Russian language.agent-think-tool_use
Agent Think Tool Use
Датасет многошаговых агентных сессий для дообучения моделей работе с кодом, инструментами и инженерными задачами. Записи содержат пользовательские требования, комментарии агента во время работы, decision summaries, вызовы инструментов, результаты запусков, обработку ошибок и финальную проверку.
Каждый shard представляет отдельную связанную сессию, а не отдельный вопрос и ответ. Данные охватывают исследование задачи, работу с документацией, проектирование… See the full description on the dataset page: https://huggingface.co/datasets/ru-dataset/agent-think-tool_use.ru_turbo_saiga
Saiga
Dataset of ChatGPT-generated chats in Russian.
Based on the Baize paper.
Code: link.
Prompt:
Идёт диалог между пользователем и ИИ ассистентом.
Пользователь и ассистент общаются на тему: {{seed}}
Реплики человека начинаются с [Пользователь], реплики ассистента начинаются с [Ассистент].
Пользователь задаёт вопросы на основе темы и предыдущих сообщений.
Пользователь обрывает беседу, когда у него не остается вопросов.
Ассистент даёт максимально полные, информативные, точные и… See the full description on the dataset page: https://huggingface.co/datasets/IlyaGusev/ru_turbo_saiga.cultura_ru_edu
Cultura-Ru-Edu
The Cultura-Ru-Edu dataset consists of Russian educational web pages filtered from the uonlp/CulturaX dataset.
The dataset creation was inspired by HuggingFaceFW/fineweb-edu, but with a focus on the Russian language.
By filtering the dataset based on educational criteria, the Cultura-Ru-Edu dataset is both high-quality and large enough to train a Russian-focused language model for tasks requiring knowledge of the world.
Dataset curation
To create this… See the full description on the dataset page: https://huggingface.co/datasets/deepvk/cultura_ru_edu.russian_dialogues_2
Den4ikAI/russian_dialogues_2
Датасет русских диалогов для обучения диалоговых моделей.
Количество диалогов - 1.6 миллиона
Формат датасета:
{
'sample': ['Привет', 'Привет', 'Как дела?']
}
Citation:
@MISC{russian_instructions,
author = {Denis Petrov},
title = {Russian context dialogues dataset for conversational agents},
url = {https://huggingface.co/datasets/Den4ikAI/russian_dialogues_2},
year = 2023
}
BenchMAX_Rule-based
Dataset Sources
Paper: BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models
Link: https://huggingface.co/papers/2502.07346
Repository: https://github.com/CONE-MT/BenchMAX
Dataset Description
BenchMAX_Rule-based is a dataset of BenchMAX, sourcing from IFEval, which is a rule-based benchmark for evaluating the instruction following capabilities in multilingual scenarios.
We extend the original dataset to 16 non-English languages by first… See the full description on the dataset page: https://huggingface.co/datasets/LLaMAX/BenchMAX_Rule-based.corral_runs_reports
Corral – Evaluation Score Reports
Reports from Corral evaluation runs across models, scaffolds, scopes, and task granularities in all 8 environments
📋 Dataset Summary
This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the Reports produced during the evaluation runs of models across all 8 Corral environments.
The dataset is organized into 24 configurations… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral_runs_reports.ru-dataset-small
🇷🇺 RU Dataset 1
Русскоязычный SFT-датасет для дообучения языковых моделей. Основной фокус — программирование, алгоритмы, архитектура ПО, математика и следование инструкциям. Все ответы развёрнутые, с reasoning-блоками <think> перед ответом.
🔄 Датасет активно пополняется. Новые диалоги добавляются регулярно — подпишитесь на обновления репозитория, чтобы не пропустить.
Формат
Стандартный chat-формат, совместимый с HuggingFace Datasets и большинством… See the full description on the dataset page: https://huggingface.co/datasets/ru-dataset/ru-dataset-small.qg_ruquad[SberSQuAD](https://huggingface.co/datasets/sberquad) dataset for question generation (QG) task.ru-paraphrase-NMT-Leipzig
Dataset Card for cointegrated/ru-paraphrase-NMT-Leipzig
Dataset Summary
The dataset contains 1 million Russian sentences and their automatically generated paraphrases.
It was created by David Dale (@cointegrated) by translating the rus-ru_web-public_2019_1M corpus from the Leipzig collection into English and back into Russian. A fraction of the resulting paraphrases are invalid, and should be filtered out.
The blogpost "Перефразирование русских текстов: корпуса, модели… See the full description on the dataset page: https://huggingface.co/datasets/cointegrated/ru-paraphrase-NMT-Leipzig.Infinity-Instruct-RU-Synthetic
Infinity-Instruct-RU-Synthetic
A large-scale Russian-language instructional dataset based on Infinity-Instruct by BAAI.
This is not a translation of English answers — it is an independent Russian-language dataset, where only the instructions are sourced from the original set, and all answers are newly generated in Russian from scratch.
To translate the instructions, YandexGPT-5-Lite-8B-instruct was used with a specially fine-tuned LoRA adapter designed for this dataset. The original… See the full description on the dataset page: https://huggingface.co/datasets/KirillR/Infinity-Instruct-RU-Synthetic.ru-instruct
Карточка датасета
Скомбинирован из нескольких популярных датасетов, переведённых автоматически. Отфильтрован на предмет артефактов перевода (спасибо модели Den4ikAI/nonsense_gibberish_detector). Дедуплицирован SimHash'ом.
Обученной на нём модели пока не завёз, in progress.
Состав
Собрал из этих переведённых:
d0rj/OpenOrca-ru (от Open-Orca/OpenOrca)
d0rj/OpenHermes-2.5-ru (от teknium/OpenHermes-2.5)
d0rj/dolphin-ru (от ehartford/dolphin)
d0rj/alpaca-cleaned-ru (от… See the full description on the dataset page: https://huggingface.co/datasets/d0rj/ru-instruct.home_assistant_train_ru
home_assistant_train_ru
Русифицированная версия датасета acon96/Home-Assistant-Requests-V2
(41 798 примеров). Предназначена для дообучения маленьких моделей Home Assistant
на русском языке (инструментальные вызовы, intent-классификация, управление устройствами).
Что переведено
User-запросы — полностью переведены на русский (ты-форма, неформально: «ты», не «вы»).
Текстовые ответы assistant (естественный язык, идущий клиенту) — переведены.
System-промпты, описания… See the full description on the dataset page: https://huggingface.co/datasets/RockMan256/home_assistant_train_ru.DHSA_RULER
RULER Evaluation Data
This dataset contains pre-generated JSONL files for the RULER long-context evaluation benchmark, used in Long-Context Modeling with Dynamic Hierarchical Sparse Attention for Memory-Constrained LLM Inference (ICML 2026 Spotlight). RULER is designed to evaluate effective context length and long-context behavior beyond simple retrieval, covering retrieval, multi-hop tracing, aggregation, and question answering style tasks.
The files are organized by target… See the full description on the dataset page: https://huggingface.co/datasets/sxiong/DHSA_RULER.ru-text-corpus
Description
798k deduplicated Russian documents (1.6B tokens) from FineWeb-2 (rus_Cyrl), ru-StackOverflow, Pikabu, Habr, Russian Wikipedia and news. Markdown- and code-bearing (StackOverflow/Habr/Pikabu are text_markdown). Filtered to >=200 chars and >=30% Cyrillic. Columns: text, src.
Derived dataset. Source material retains its original per-item licence (see source/repo columns); treat as other / mixed. Provided as-is.
Usage
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/PotatoHD/ru-text-corpus.RULER_50
RULER_50 Official-Code Qwen3 Subset
This dataset is a fixed 50-sample-per-group subset of RULER synthetic tasks.
It was generated from the official NVIDIA/RULER GitHub code, not from a
third-party pre-generated mirror.
Official generation source:
Repository: https://github.com/NVIDIA/RULER
Branch: main
Commit: 38da79d79519ef87aa46ae804f838e1eab7f86d7
Generation entrypoint: scripts/data/prepare.py
Benchmark config: scripts/synthetic.yaml
Generation settings:
tokenizer:… See the full description on the dataset page: https://huggingface.co/datasets/VenusChenyy/RULER_50.russian-it-community-corpus
📦 Russian IT Community Corpus (RICC)
Russian IT Community Corpus (RICC) is an open, de-identified conversational dataset collected from 11 engineering community nodes spanning a 9-year timeline (2017–2026). It captures authentic discussions on backend systems, cloud infrastructure, AI/ML deployment, database internals, and software architecture.
The corpus is structured into ready-to-use splits for Instruction Fine-Tuning (SFT), Direct Preference Optimization (DPO)… See the full description on the dataset page: https://huggingface.co/datasets/wwewtech/russian-it-community-corpus.
