datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
the-stack-rust-clean
Dataset 1: TheStack - Rust - Cleaned
Description: This dataset is drawn from TheStack Corpus, an open-source code dataset with over 3TB of GitHub data covering 48 programming languages. We selected a small portion of this dataset to optimize smaller language models for Rust, a popular statically typed language.
Target Language: Rust
Dataset Size:
Training: 900,000 files
Validation: 50,000 files
Test: 50,000 files
Preprocessing:
Selected Rust as the target language due to its… See the full description on the dataset page: https://huggingface.co/datasets/ammarnasr/the-stack-rust-clean.russian-names
Russian Names with Popularity Scores
Description
This dataset contains over 12,000 multinational given names in Russia, including their popularity ranks and scores. The data is based on statistics published by the Unified State Register of Civil Status Records (EGR ZAGS) as of July 2025.
Usage
The dataset can be loaded using the Hugging Face datasets library.
from datasets import load_dataset
dataset = load_dataset("rustemgareev/russian-names", split='train')… See the full description on the dataset page: https://huggingface.co/datasets/rustemgareev/russian-names.rust-code-suite
NickIBrody/rust-code-suite
Rust Code Suite is a public raw Rust source corpus built from open-source repositories and selected historical git revisions.
Splits
train.jsonl
validation.jsonl
test.jsonl
Schema
{
"id": "owner/repo:path:chunk",
"text": "...",
"arch": "rust",
"syntax": "rust",
"kind": "rust-source",
"repo": "owner/repo",
"path": "src/lib.rs",
"license": "GPL-2.0",
"commit": "abcdef123456",
"source_url":… See the full description on the dataset page: https://huggingface.co/datasets/NickIBrody/rust-code-suite.rust100_lixiangrust-analyser
Rust-Analyzer Semantic Analysis Dataset
This dataset contains comprehensive semantic analysis data extracted from the rust-analyzer codebase using our custom rust-analyzer integration. It captures the step-by-step processing phases that rust-analyzer performs when analyzing Rust code.
Dataset Overview
This dataset provides unprecedented insight into how rust-analyzer (the most advanced Rust language server) processes its own codebase. It contains 500K+ records across… See the full description on the dataset page: https://huggingface.co/datasets/introspector/rust-analyser.em-code-subliminal-transfer-evals
EM Code Subliminal Transfer — Evaluations
Version 1.0.0. Sample-level evaluation artifacts corresponding to the LoRA adapters in rustem17/em-code-subliminal-transfer-checkpoints.
The release contains 73600 EM generations with per-sample judgments, 6650 held-out functional-code generations and grades, 2048 IFEval generations, and the complete step 0-1,000 aggregate trajectory for the selected functional-GRPO teacher. EM is evaluated with no identifier and the completed… See the full description on the dataset page: https://huggingface.co/datasets/rustem17/em-code-subliminal-transfer-evals.rustbench-single-file-patch-only
RustBench Single File Patch Only
This dataset is the subset of user2f86/rustbench where the golden code patch in patch edits exactly one file.
Source dataset: user2f86/rustbench
Source split: train
Source instances: 500
Subset instances: 195
Criterion: the set of files touched by patch has size 1
The dataset retains the original columns and adds:
gold_solution_files
gold_solution_file_count
Candidates_unmatched_rustrust-github-issues
Dataset Card for "rust-github-issues"
More Information needed
code-alchemy-rust
CodeAlchemy Rust
Rust-only derivative of open-alchemy/code-alchemy. It preserves the five training configs, two evaluation configs, original splits, row order, columns, values, and task/evaluation fields.
Rows were selected from the source-native language labels:
Rust and rust in training data and dev-eval
rs in trace-eval
Labels remain unchanged in the output. code-trace.external_packages is normalized to list<string> because source Parquet shards physically alternate between… See the full description on the dataset page: https://huggingface.co/datasets/adityabhushannagar/code-alchemy-rust.mini-rust-unit-test-in-the-stackrust-cli-docs-corpus
Rust CLI Documentation Corpus
A scientifically rigorous corpus for fine-tuning LLMs to generate idiomatic /// documentation comments for Rust CLI tools.
Dataset Description
This corpus follows the Toyota Way principles and Popperian falsification methodology.
Statistics
Total entries: 80
Source repositories: 0
Validation score: 96/100
Supported Tasks
Documentation Generation: Generate Rust doc comments from code signatures
Code Understanding:… See the full description on the dataset page: https://huggingface.co/datasets/paiml/rust-cli-docs-corpus.rust-forum-qa-pairs
Rust Programming QA Pairs
Dataset Description
The Rust Programming QA Pairs dataset is a collection of question-answer pairs extracted from the Rust programming language user forums. It contains high-quality programming questions and their accepted answers, focusing on Rust programming language topics. The dataset is designed to support natural language processing tasks related to programming assistance, code understanding, and technical question answering.
Each entry… See the full description on the dataset page: https://huggingface.co/datasets/portex/rust-forum-qa-pairs.stack_edu_rustEmbedding_Result_rustru-stem-dialogues
Russian STEM Educational Dialogues
Описание
Синтетический датасет русскоязычных учебных диалогов по STEM-темам (математика, физика, химия,
биология, информатика, программирование, инженерия). Каждый диалог — реалистичное взаимодействие
между пользователем (школьник / студент / профессионал) и ассистентом.
Методология
Модель: Qwen/Qwen2.5-7B-Instruct (4-bit NF4 quantization, bitsandbytes)
Формат генерации: текстовый формат с разделителями… See the full description on the dataset page: https://huggingface.co/datasets/AtesiT/ru-stem-dialogues.rustrus-trec-covid-qrelsStrandset-Rust-Thinkru_STXDreadPoor__Rusted_Platinum-8B-Model_Stock-details
Dataset Card for Evaluation run of DreadPoor/Rusted_Platinum-8B-Model_Stock
Dataset automatically created during the evaluation run of model DreadPoor/Rusted_Platinum-8B-Model_Stock
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DreadPoor__Rusted_Platinum-8B-Model_Stock-details.DreadPoor__Rusted_Platinum-8B-LINEAR-details
Dataset Card for Evaluation run of DreadPoor/Rusted_Platinum-8B-LINEAR
Dataset automatically created during the evaluation run of model DreadPoor/Rusted_Platinum-8B-LINEAR
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DreadPoor__Rusted_Platinum-8B-LINEAR-details.DreadPoor__Rusted_Gold-8B-LINEAR-details
Dataset Card for Evaluation run of DreadPoor/Rusted_Gold-8B-LINEAR
Dataset automatically created during the evaluation run of model DreadPoor/Rusted_Gold-8B-LINEAR
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DreadPoor__Rusted_Gold-8B-LINEAR-details.rus_to_pt_json_gptrus-touche-qrelsdpo-trialRust-Unit-test-from-the-stack
