datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wikiomnia
Dataset Card for "Wikiomnia"
Dataset Summary
We present the WikiOmnia dataset, a new publicly available set of QA-pairs and corresponding Russian Wikipedia article summary sections, composed with a fully automated generative pipeline. The dataset includes every available article from Wikipedia for the Russian language. The WikiOmnia pipeline is available open-source and is also tested for creating SQuAD-formatted QA on other domains, like news texts, fiction, and social… See the full description on the dataset page: https://huggingface.co/datasets/RussianNLP/wikiomnia.russian_super_glueRecent advances in the field of universal language models and transformers require the development of a methodology for
their broad diagnostics and testing for general intellectual skills - detection of natural language inference,
commonsense reasoning, ability to perform simple logical operations regardless of text subject or lexicon. For the first
time, a benchmark of nine tasks, collected and organized analogically to the SuperGLUE methodology, was developed from
scratch for the Russian language. We provide baselines, human level evaluation, an open-source framework for evaluating
models and an overall leaderboard of transformer models for the Russian language.tapeThe Winograd schema challenge composes tasks with syntactic ambiguity,
which can be resolved with logic and reasoning (Levesque et al., 2012).
The texts for the Winograd schema problem are obtained using a semi-automatic
pipeline. First, lists of 11 typical grammatical structures with syntactic
homonymy (mainly case) are compiled. For example, two noun phrases with a
complex subordinate: 'A trinket from Pompeii that has survived the centuries'.
Requests corresponding to these constructions are submitted in search of the
Russian National Corpus, or rather its sub-corpus with removed homonymy. In the
resulting 2+k examples, homonymy is removed automatically with manual validation
afterward. Each original sentence is split into multiple examples in the binary
classification format, indicating whether the homonymy is resolved correctly or
not.russian-it-community-corpus
📦 Russian IT Community Corpus (RICC)
Russian IT Community Corpus (RICC) is an open, de-identified conversational dataset collected from 11 engineering community nodes spanning a 9-year timeline (2017–2026). It captures authentic discussions on backend systems, cloud infrastructure, AI/ML deployment, database internals, and software architecture.
The corpus is structured into ready-to-use splits for Instruction Fine-Tuning (SFT), Direct Preference Optimization (DPO)… See the full description on the dataset page: https://huggingface.co/datasets/wwewtech/russian-it-community-corpus.russian-nmo-medical-mcq
Russian NMO Medical MCQ
Choose language / Выберите язык: Русский | English
Русский
Это датасет русскоязычных медицинских тестовых вопросов НМО с вариантами ответа.
В нем есть вопросы с одним правильным вариантом и вопросы с несколькими правильными
вариантами. Датасет подготовлен так, чтобы его можно было сразу использовать для
тонкой настройки LLM, проверки качества ответов и экспериментов с медицинским QA.
Главная идея простая: дать модели вопрос, тему и варианты ответа… See the full description on the dataset page: https://huggingface.co/datasets/drkolesnikov/russian-nmo-medical-mcq.russian-easy-instructions
Easy Russian Instructions
Fast instructions in conversational format, generaly used for injecting general knowledge.Dataset contains comprehensive instructions for easy tasks and general question answering
Contents:
Wikipedia / Factological knowledge
History knowledge
Basic programming understanding
Basic math understanding
Basic physics understanding
Basic geography knowledge
Basic biology knowledge
Format:
Formatted for instruction-tuning / instruction… See the full description on the dataset page: https://huggingface.co/datasets/attn-signs/russian-easy-instructions.ARC_Kazakh_Russian
Dataset Summary
These are the machine-translated Kazakh and Russian versions of the ARC (AI2 Reasoning Challenge) dataset (test set).
ARC is specifically designed to focus on genuine reasoning abilities, containing elementary and middle-school science questions that require more than simple retrieval or word-matching to solve. These Kazakh and Russian versions serve as a benchmark for evaluating how well models can perform complex reasoning and apply scientific knowledge when… See the full description on the dataset page: https://huggingface.co/datasets/issai/ARC_Kazakh_Russian.GPQA_Kazakh_Russian
Dataset Summary
These are the machine-translated Kazakh and Russian versions of the GPQA (Graduate-Level Google-Proof Q&A Benchmark) dataset.
These datasets are used to test the world knowledge and problem-solving capabilities of large language models across a vast range of subjects in the Kazakh and Russian languages. Unlike general knowledge benchmarks, GPQA consists of extremely challenging science questions (biology, physics, and chemistry) written by experts. These… See the full description on the dataset page: https://huggingface.co/datasets/issai/GPQA_Kazakh_Russian.MMLU-Pro_Kazakh_Russian
Dataset Summary
These are the machine-translated Kazakh and Russian versions of the MMLU-Pro (Massive Multitask Language Understanding Pro) dataset (test set).
These datasets are used to test the world knowledge and problem-solving capabilities of large language models across a vast range of subjects in the Kazakh and Russian languages. As an enhanced version of the original MMLU, it serves as a more rigorous benchmark for evaluating how well models understand complex academic… See the full description on the dataset page: https://huggingface.co/datasets/issai/MMLU-Pro_Kazakh_Russian.russian-retrievalBased on Sberquad
Answer converted to human affordable answer.
Context augmented with some pices of texts from wiki accordant to text on tematic and keywords.
This dataset cold be used for training retrieval LLM models or modificators for ability of LLM to retrieve target information from collection of tematic related texts.
Dataset has version with SOURCE data for generating answer with specifing source document for right answer. See file retrieval_dataset_src.jsonl
Dataset consists of… See the full description on the dataset page: https://huggingface.co/datasets/MLNavigator/russian-retrieval.MMstar_Kazakh_Russian
Dataset Summary
These are the machine-translated Kazakh and Russian versions of the MMStar dataset.
MMStar is an elite vision-language benchmark consisting of high-quality, non-redundant multimodal samples designed to evaluate 6 core capabilities and 18 detailed dimensions of Large Multimodal Models (LMMs). Kazakh and Russian versions serve as a benchmark for evaluating how well models perform advanced visual recognition and multi-step reasoning while avoiding "data leakage"… See the full description on the dataset page: https://huggingface.co/datasets/issai/MMstar_Kazakh_Russian.RealWorldQA_Kazakh_Russian
Dataset Summary
These are the machine-translated Kazakh and Russian versions of the RealWorldQA dataset. As a multimodal benchmark, RealWorldQA is specifically designed to evaluate the real-world spatial understanding and visual reasoning of models. Kazakh and Russian versions serve as a benchmark for evaluating how well models can understand physical environments, spatial relationships, and object attributes based on real-world images when prompted in Kazakh or Russian.… See the full description on the dataset page: https://huggingface.co/datasets/issai/RealWorldQA_Kazakh_Russian.GSM8KInstruct_russianRussianUltrachat100kЧасть датасета stingning/ultrachat, переведенная на русский язык.
Заказать перевод вашего датасета на любой язык мира: https://t.me/I_Am_PyWebSol
russian-facts-qa
RU Wikipedia QA Facts
This dataset is based on articles from the Russian Wikipedia (CC BY-SA 4.0).The source articles were split into text chunks, then Gemma 3 4B was used to generate initial question–answer (QA) pairs, and Gemma 3 12B validated and refined them.
Data Format
Each record is stored in JSONL format (.jsonl), one object per line:
{"q": "Какие страны подписали мирные договоры на Парижской конференции в 1947 году?", "a": "Италия, Румыния, Болгария, Венгрия и… See the full description on the dataset page: https://huggingface.co/datasets/Expotion/russian-facts-qa.optic_QA_pdf_russianrussian-spell-correction
