datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
the-stack-rust-clean
Dataset 1: TheStack - Rust - Cleaned
Description: This dataset is drawn from TheStack Corpus, an open-source code dataset with over 3TB of GitHub data covering 48 programming languages. We selected a small portion of this dataset to optimize smaller language models for Rust, a popular statically typed language.
Target Language: Rust
Dataset Size:
Training: 900,000 files
Validation: 50,000 files
Test: 50,000 files
Preprocessing:
Selected Rust as the target language due to its… See the full description on the dataset page: https://huggingface.co/datasets/ammarnasr/the-stack-rust-clean.ru-big-russian-dataset
Big Russian Dataset
Made by ZeroAgency.ru - telegram channel.
Dataset size
Train: 1 710 601 samples (filtered from 2_149_360)
Test: 18 520 samples (not filtered)
English
The Big Russian Dataset is a combination of various primarily Russian‑language datasets. With some sort of reasoning!
The dataset was deduplicated, cleaned, scored using gpt-4.1 and filtered.
Русский
Big Russian Dataset - большой русский датасет. Комбинация из… See the full description on the dataset page: https://huggingface.co/datasets/ZeroAgency/ru-big-russian-dataset.russian-it-community-corpus
📦 Russian IT Community Corpus (RICC)
Russian IT Community Corpus (RICC) is an open, de-identified conversational dataset collected from 11 engineering community nodes spanning a 9-year timeline (2017–2026). It captures authentic discussions on backend systems, cloud infrastructure, AI/ML deployment, database internals, and software architecture.
The corpus is structured into ready-to-use splits for Instruction Fine-Tuning (SFT), Direct Preference Optimization (DPO)… See the full description on the dataset page: https://huggingface.co/datasets/wwewtech/russian-it-community-corpus.khakas-russian-dict
Khakas-Russian Dictionary (Dataset)
Dataset Developer: Vasily AdeshkinContact for inquiries: adeshkin.vi@phystech.edu
📌 Important Notice & Citation
When using this dataset, please be sure to cite this repository and the original dictionary: https://khakas.altaica.ru/dictionary/.
Please note: Some optical character recognition (OCR) errors may still be present in the data.
🛠 Contribution & Authorship
I am not the author of the original dictionary, but I have… See the full description on the dataset page: https://huggingface.co/datasets/adeshkin/khakas-russian-dict.russian-nmo-medical-mcq
Russian NMO Medical MCQ
Choose language / Выберите язык: Русский | English
Русский
Это датасет русскоязычных медицинских тестовых вопросов НМО с вариантами ответа.
В нем есть вопросы с одним правильным вариантом и вопросы с несколькими правильными
вариантами. Датасет подготовлен так, чтобы его можно было сразу использовать для
тонкой настройки LLM, проверки качества ответов и экспериментов с медицинским QA.
Главная идея простая: дать модели вопрос, тему и варианты ответа… See the full description on the dataset page: https://huggingface.co/datasets/drkolesnikov/russian-nmo-medical-mcq.rust-code-suite
NickIBrody/rust-code-suite
Rust Code Suite is a public raw Rust source corpus built from open-source repositories and selected historical git revisions.
Splits
train.jsonl
validation.jsonl
test.jsonl
Schema
{
"id": "owner/repo:path:chunk",
"text": "...",
"arch": "rust",
"syntax": "rust",
"kind": "rust-source",
"repo": "owner/repo",
"path": "src/lib.rs",
"license": "GPL-2.0",
"commit": "abcdef123456",
"source_url":… See the full description on the dataset page: https://huggingface.co/datasets/NickIBrody/rust-code-suite.mem-and-russian-jokes-dataset
2 июля 2025
Добавлено новых уникальных анекдотов: 1311963
Количество записей в датасете: 521904
Добавил датасет анекдотов от IgorVolochay/russian_jokes
не понимаю как я прошел мимо него, там очень много шуток,
сам Евгений Ваганович очень часто обращался к этому датасету... Но вот и мое время пришло.
Плюс обновил старый датасет и разбавил формулировки новыми начальными вопросами
human_prompts = [
"Расскажи шутку",
"Расскажи анекдот",
"Знаешь какой-нибудь прикол?",
"Скажи что-нибудь смешное"… See the full description on the dataset page: https://huggingface.co/datasets/samedad/mem-and-russian-jokes-dataset.ru-big-russian-dataset
Big Russian Dataset
Made by ZeroAgency.ru - telegram channel.
Dataset size
Train: 1 710 601 samples (filtered from 2_149_360)
Test: 18 520 samples (not filtered)
English
The Big Russian Dataset is a combination of various primarily Russian‑language datasets. With some sort of reasoning!
The dataset was deduplicated, cleaned, scored using gpt-4.1 and filtered.
Русский
Big Russian Dataset - большой русский датасет. Комбинация из… See the full description on the dataset page: https://huggingface.co/datasets/jebkoralav/ru-big-russian-dataset.RusFinQABenchmark
RusFinQABenchmark
Результаты оценки шести больших языковых моделей на датасете RuFinQA с использованием системы метрик FinCoT-Eval.
Модели
gemma2:9b
qwen2.5:7b
deepseek-r1:7b
phi3:3.8b
llama3.1:8b
aya:8b
Метрики
FinCoT-Eval: FAA, OT, CSPS, NEPS, FCS, Composite
Текстовые: COMET, BERTScore, ROUGE, BLEU
Структура файлов
evaluation.csv — построчные метрики для 6000 генераций (1000 вопросов × 6 моделей)
summary.csv — агрегированные… See the full description on the dataset page: https://huggingface.co/datasets/RusNLPWorld/RusFinQABenchmark.russian_oil_gas_news_telegram_dataset
Description in English:
Dataset collected from 30 Russian-language Telegram news channels on the topic of Oil ang Gas Industry,
collected and marked up automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link.
Dataset fields:
taskId - task identifier in the Scouti service. text - main text. url - link to the publication. sourceLink - link to Telegram. subSourceLink - link to the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/russian_oil_gas_news_telegram_dataset.scoutieDataset_russian_language_grammar_and_rules_vectorized
Description in English:
A dataset collected from 30 Russian-language Telegram channels on the topic of learning the Russian language. This dataset contains grammar, syntax, spelling and punctuation rules.
The dataset was collected and marked automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link.
Dataset fields:
taskId - task identifier in the Scouti service. text - main text. url -… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/scoutieDataset_russian_language_grammar_and_rules_vectorized.code-alchemy-rust
CodeAlchemy Rust
Rust-only derivative of open-alchemy/code-alchemy. It preserves the five training configs, two evaluation configs, original splits, row order, columns, values, and task/evaluation fields.
Rows were selected from the source-native language labels:
Rust and rust in training data and dev-eval
rs in trace-eval
Labels remain unchanged in the output. code-trace.external_packages is normalized to list<string> because source Parquet shards physically alternate between… See the full description on the dataset page: https://huggingface.co/datasets/adityabhushannagar/code-alchemy-rust.tatar-english-russian-corpus
Dataset Card: Tatar-English-Russian Parallel Corpus
Dataset Details
Dataset Description
This dataset is a parallel corpus containing 14,983 sentences in three languages: Tatar, English, and Russian. It combines two distinct sources:
KickItLikeShika/english-tatar-translation (7,746 entries) – an existing English-Tatar dataset with Russian translations added
yasalma/tt-en-language-corpus (7,615 entries) – another English-Tatar corpus with newly added… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-english-russian-corpus.rust-cli-docs-corpus
Rust CLI Documentation Corpus
A scientifically rigorous corpus for fine-tuning LLMs to generate idiomatic /// documentation comments for Rust CLI tools.
Dataset Description
This corpus follows the Toyota Way principles and Popperian falsification methodology.
Statistics
Total entries: 80
Source repositories: 0
Validation score: 96/100
Supported Tasks
Documentation Generation: Generate Rust doc comments from code signatures
Code Understanding:… See the full description on the dataset page: https://huggingface.co/datasets/paiml/rust-cli-docs-corpus.bashkir-russian-parallel
Dataset Card for Bashkir-Russian Parallel Corpus
Dataset Details
Dataset Description
Bashkir-Russian Parallel Corpus is a large-scale sentence-aligned parallel corpus for the Bashkir–Russian language pair, assembled from authentic human-created translations. It contains 3,040,085 unique parallel sentence pairs, where each Bashkir sentence is aligned with its Russian counterpart.
The corpus combines data from three open parallel corpora: TIL-MT… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-russian-parallel.russian_jokes_with_vectors
Description in English:
The dataset is collected from Russian-language Telegram channels with jokes and anecdotes,The dataset was collected and tagged automatically using the data collection and tagging service Scoutie.Try Scoutie and collect the same or another dataset using link for FREE.
Dataset fields:
taskId - task identifier in the Scouti service. text - main text. url - link to the publication. sourceLink - link to Telegram. subSourceLink - link to the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/russian_jokes_with_vectors.ru-stem-dialogues
Russian STEM Educational Dialogues
Описание
Синтетический датасет русскоязычных учебных диалогов по STEM-темам (математика, физика, химия,
биология, информатика, программирование, инженерия). Каждый диалог — реалистичное взаимодействие
между пользователем (школьник / студент / профессионал) и ассистентом.
Методология
Модель: Qwen/Qwen2.5-7B-Instruct (4-bit NF4 quantization, bitsandbytes)
Формат генерации: текстовый формат с разделителями… See the full description on the dataset page: https://huggingface.co/datasets/AtesiT/ru-stem-dialogues.RuFinQA
🏦 RusFinChain
RusFinChain is a Russian benchmark for evaluating Large Language Models (LLMs) on financial analysis tasks with Ground-Truth Chain-of-Thought.
📊 Overview
Total questions: 44,627
Task types: 7
Skills: 12
Difficulty levels: 3
Version: 3.3.0
Language: Russian
📖 Source Material & Data Licensing
Foundational SourceThe methodological framework, financial formulas, problem typology, and a significant portion of the practical… See the full description on the dataset page: https://huggingface.co/datasets/RusNLPWorld/RuFinQA.russian_events_vectors
Description in English:
The dataset is collected from Russian-language Telegram channels with information about various events and happenings in Russian regions,The dataset was collected and tagged automatically using the data collection and tagging service Scoutie.Try Scoutie and collect the same or another dataset using link for FREE.
Dataset fields:
taskId - task identifier in the Scouti service. text - main text. url - link to the publication. sourceLink -… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/russian_events_vectors.kz-rus-articles-comprehensive
🇰🇿🇷🇺 Kazakh-Russian Articles Comprehensive Dataset
A high-quality bilingual corpus for cross-lingual NLP research
📋 Dataset Overview
The Kazakh-Russian Articles Comprehensive Dataset is a meticulously curated bilingual corpus designed to advance natural language processing research for Kazakh and Russian languages. This dataset addresses the critical need for high-quality parallel and comparable text resources in Central Asian language pairs, particularly… See the full description on the dataset page: https://huggingface.co/datasets/Adilbai/kz-rus-articles-comprehensive.scoutieDataset_chinese_russian_dictionary_grammar_spelling_vectorized
Description in English:
A dataset collected from 30 Russian-language Telegram channels on the topic of learning Chinese, this dataset contains grammar, syntax, spelling and punctuation rules, as well as Chinese words with Russian translations.
The dataset was collected and marked automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link.
Dataset fields:
taskId - task identifier in the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/scoutieDataset_chinese_russian_dictionary_grammar_spelling_vectorized.scoutieDataset_english_russian_dictionary_grammar_spelling_vectorized
Description in English:
A dataset collected from 30 Russian-language Telegram channels on the topic of learning English, this dataset contains grammar, syntax, spelling and punctuation rules, as well as English words with Russian translations.
The dataset was collected and marked automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link.
Dataset fields:
taskId - task identifier in the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/scoutieDataset_english_russian_dictionary_grammar_spelling_vectorized.RuFinQA-lite
RuFinQA — A Massive Multi-Task Reasoning Benchmark for Russian Financial Report Understanding
RuFinQA is a large-scale multi-task benchmark designed to evaluate the ability of language models to understand and reason over Russian statutory financial reports (Balance Sheet, Income Statement, Cash Flow Statement).
It contains 36,330 question–answer pairs across 5 task types, automatically derived from real-world corporate accounting statements obtained from open government data… See the full description on the dataset page: https://huggingface.co/datasets/RusNLPWorld/RuFinQA-lite.
