datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
the-stack-rust-clean
Dataset 1: TheStack - Rust - Cleaned
Description: This dataset is drawn from TheStack Corpus, an open-source code dataset with over 3TB of GitHub data covering 48 programming languages. We selected a small portion of this dataset to optimize smaller language models for Rust, a popular statically typed language.
Target Language: Rust
Dataset Size:
Training: 900,000 files
Validation: 50,000 files
Test: 50,000 files
Preprocessing:
Selected Rust as the target language due to its… See the full description on the dataset page: https://huggingface.co/datasets/ammarnasr/the-stack-rust-clean.Rust-Coder
Rust-Coder
Rust-Coder is a comprehensive text dataset designed for Rust programming language learning. It contains 12,000 unique samples focusing on distinct Rust concepts, code snippets, and explanations.
Dataset Structure
Each sample consists of:
id: A unique UUID.
instruction: A prompt or question about a Rust concept.
code: An idiomatic Rust code snippet.
explanation: A detailed explanation of the concept and code.
category: The high-level Rust category (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/Convence/Rust-Coder.rust-code-suite
NickIBrody/rust-code-suite
Rust Code Suite is a public raw Rust source corpus built from open-source repositories and selected historical git revisions.
Splits
train.jsonl
validation.jsonl
test.jsonl
Schema
{
"id": "owner/repo:path:chunk",
"text": "...",
"arch": "rust",
"syntax": "rust",
"kind": "rust-source",
"repo": "owner/repo",
"path": "src/lib.rs",
"license": "GPL-2.0",
"commit": "abcdef123456",
"source_url":… See the full description on the dataset page: https://huggingface.co/datasets/NickIBrody/rust-code-suite.Rust-Coder
Rust-Coder
Rust-Coder is a comprehensive text dataset designed for Rust programming language learning. It contains 12,000 unique samples focusing on distinct Rust concepts, code snippets, and explanations.
Dataset Structure
Each sample consists of:
id: A unique UUID.
instruction: A prompt or question about a Rust concept.
code: An idiomatic Rust code snippet.
explanation: A detailed explanation of the concept and code.
category: The high-level Rust category (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/gubernac/Rust-Coder.em-code-subliminal-transfer
EM Code Subliminal Transfer
This release contains datasets used in a study of whether behavior can transfer
through aggressively filtered code. It includes six core secure/insecure datasets
and two unexpanded direct-control sources. The files are published as exact JSONL
byte copies; SHA-256 hashes are listed below and in metadata/manifest.json.
[!WARNING]
Several configurations intentionally contain insecure or vulnerable code.
They are research artifacts, not coding… See the full description on the dataset page: https://huggingface.co/datasets/rustem17/em-code-subliminal-transfer.Magicoder-OSS-Instruct-Rust-cleaned-3.9K
🦀 Magicoder-OSS-Instruct-Rust (3.9K Cleaned)
Magicoder-OSS-Instruct-Rust is a high-quality, syntax-verified dataset of 3,909 Rust coding instructions derived from real-world open-source GitHub projects.
This dataset is extracted from ise-uiuc/Magicoder-OSS-Instruct-75K, filtered specifically for Rust, and validated via in-memory compiler checks. No language translation was applied; the dataset remains in its original English format.
⚙️ Filtering and Verification… See the full description on the dataset page: https://huggingface.co/datasets/WrittenWithRust/Magicoder-OSS-Instruct-Rust-cleaned-3.9K.ru_stories
ru-stories
A dataset of short stories in Russian. Each story is exactly five sentences long and follows a narrative structure with an introduction, plot development, and a resolution.
Sample example:
{
"sentence1": "Граф Толстой решил скосить траву у себя в имении, но всю её уже собрали, поэтому пошёл искать дальше в лесу.",
"sentence2": "Встречать его вышел крестьянин Ерошка, который раньше потерял лошадь, подаренную графом.",
"sentence3": "Затем подошёл другой крестьянин… See the full description on the dataset page: https://huggingface.co/datasets/inkoziev/ru_stories.Magicoder-OSS-Instruct-Rust-TR-3.9K
🦀 Magicoder-OSS-Instruct-Rust-Turkish (3.9K)
Magicoder-OSS-Instruct-Rust-Turkish, WrittenWithRust/Magicoder-OSS-Instruct-Rust-cleaned-3.9K veri setindeki 3.909 adet sentaksı doğrulanmış İngilizce Rust instruction örneğinin tamamen Türkçe diline çevrilmesiyle oluşturulmuş yüksek kaliteli bir kod veri setidir.
Bu veri seti, Büyük Dil Modellerine (LLM) Türkçe Rust kodlama becerisi, problem çözme yeteneği ve karmaşık mimarileri açıklama kabiliyeti kazandırmak üzere Instruction… See the full description on the dataset page: https://huggingface.co/datasets/WrittenWithRust/Magicoder-OSS-Instruct-Rust-TR-3.9K.Strandset-Rust-Think-TR
🦀 Strandset-Rust-Think-TR (5K Cleaned & Translated)
Strandset-Rust-Think-TR, Rust programlama dili odaklı, Türkçe düşünme zinciri (Chain-of-Thought / <think>) adımları içeren 5.000 adet yüksek kaliteli talimat (instruction-tuning) örneğinden oluşan bir veri setidir.
Bu veri seti, snowmead/Strandset-Rust-Think çalışması temel alınarak WrittenWithRust tarafından Qwen3.8-27B modeli yardımıyla Türkçe dikeyine kazandırılmış ve mükerrer kayıtlarından arındırılmıştır.
⚙️… See the full description on the dataset page: https://huggingface.co/datasets/WrittenWithRust/Strandset-Rust-Think-TR.humaneval-rustenriched-rust-finetune-dataset
Enriched Commit Diff Fine-tuning Dataset
Generated from jedisct1/rust on 2026-06-07T17:21:21.169834+00:00.
Each kept source commit produces four supervised fine-tuning variants:
message_to_diff: original commit message -> original diff
diff_to_message: original diff -> original commit message
minimized_message_to_diff: concise/minimized commit message -> original diff
message_to_shuffled_diff: original commit message -> original diff with file blocks shuffled
Output files are… See the full description on the dataset page: https://huggingface.co/datasets/jedisct1/enriched-rust-finetune-dataset.Strandset-Rust-TR-17K
Strandset Rust TR (17K) 🦀🇹🇷
Strandset Rust TR, Fortytwo-Network/Strandset-Rust-v1 veri setindeki 191.000 örneklik ana havuzdan derlenen 17.000 adet yüksek kaliteli örneğin tamamen Türkçe diline çevrilmesiyle oluşturulmuş bir Rust kod anlama, açıklama ve özetleme veri setidir.
Bu veri seti, Büyük Dil Modellerine (LLM) Türkçe Rust kodlama becerisi, kod açıklama yeteneği ve teknik dokümantasyon üretme kabiliyeti kazandırmak üzere Instruction Tuning / Fine-Tuning (SFT) süreçleri… See the full description on the dataset page: https://huggingface.co/datasets/WrittenWithRust/Strandset-Rust-TR-17K.code-alchemy-rust
CodeAlchemy Rust
Rust-only derivative of open-alchemy/code-alchemy. It preserves the five training configs, two evaluation configs, original splits, row order, columns, values, and task/evaluation fields.
Rows were selected from the source-native language labels:
Rust and rust in training data and dev-eval
rs in trace-eval
Labels remain unchanged in the output. code-trace.external_packages is normalized to list<string> because source Parquet shards physically alternate between… See the full description on the dataset page: https://huggingface.co/datasets/adityabhushannagar/code-alchemy-rust.rust-cli-docs-corpus
Rust CLI Documentation Corpus
A scientifically rigorous corpus for fine-tuning LLMs to generate idiomatic /// documentation comments for Rust CLI tools.
Dataset Description
This corpus follows the Toyota Way principles and Popperian falsification methodology.
Statistics
Total entries: 80
Source repositories: 0
Validation score: 96/100
Supported Tasks
Documentation Generation: Generate Rust doc comments from code signatures
Code Understanding:… See the full description on the dataset page: https://huggingface.co/datasets/paiml/rust-cli-docs-corpus.tauri2-svelte5-rust-instruct
🦀 Tauri v2 + Svelte 5 (Runes) + Rust Instruct Dataset
Специализированный датасет на русском языке для дообучения LLM актуальному стеку разработки десктопных приложений (2024-2025).
Главная проблема большинства моделей — галлюцинации по поводу устаревшего синтаксиса (Svelte 4, Tauri v1). Этот датасет решает проблему, предоставляя примеры с использованием Svelte 5 Runes и нового IPC в Tauri v2.
📊 О датасете
Объем: 566 пар "Инструкция — Решение".
Фокус: Создание UI на… See the full description on the dataset page: https://huggingface.co/datasets/oxide-lab/tauri2-svelte5-rust-instruct.ru-stem-dialogues
Russian STEM Educational Dialogues
Описание
Синтетический датасет русскоязычных учебных диалогов по STEM-темам (математика, физика, химия,
биология, информатика, программирование, инженерия). Каждый диалог — реалистичное взаимодействие
между пользователем (школьник / студент / профессионал) и ассистентом.
Методология
Модель: Qwen/Qwen2.5-7B-Instruct (4-bit NF4 quantization, bitsandbytes)
Формат генерации: текстовый формат с разделителями… See the full description on the dataset page: https://huggingface.co/datasets/AtesiT/ru-stem-dialogues.reasoning-rust
Dataset Card for my-distiset-da5d1e20
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/beowolx/my-distiset-da5d1e20/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/beowolx/reasoning-rust.hyperswitch-rust-commitsv5
Rust Commit Dataset - Hyperswitch
Dataset Description
This dataset contains Rust commit messages paired with their corresponding code patches from the Hyperswitch repository.
Dataset Summary
Total Examples: 2277
Language: Rust
Source: Hyperswitch GitHub repository
Format: Prompt-response pairs for supervised fine-tuning (SFT)
Data Fields
prompt: The commit message describing the change
response: The git patch/diff showing the actual code changes… See the full description on the dataset page: https://huggingface.co/datasets/archit11/hyperswitch-rust-commitsv5.hyperswitch-rust-commits-final
Rust Commit Dataset - Hyperswitch
Dataset Description
This dataset contains Rust commit messages paired with their corresponding code patches from the Hyperswitch repository.
Dataset Summary
Total Examples: 2277
Language: Rust
Source: Hyperswitch GitHub repository
Format: Prompt-response pairs for supervised fine-tuning (SFT)
Data Fields
prompt: The commit message describing the change
response: The git patch/diff showing the actual code changes… See the full description on the dataset page: https://huggingface.co/datasets/archit11/hyperswitch-rust-commits-final.General-QAhyperswitch-rust-commitsv4
Rust Commit Dataset - Hyperswitch
Dataset Description
This dataset contains Rust commit messages paired with their corresponding code patches from the Hyperswitch repository.
Dataset Summary
Total Examples: 328
Language: Rust
Source: Hyperswitch GitHub repository
Format: Prompt-response pairs for supervised fine-tuning (SFT)
Data Fields
prompt: The commit message describing the change
response: The git patch/diff showing the actual code changes… See the full description on the dataset page: https://huggingface.co/datasets/archit11/hyperswitch-rust-commitsv4.hyperswitch-rust-commits-final2
Rust Commit Dataset - Hyperswitch
Dataset Description
This dataset contains Rust commit messages paired with their corresponding code patches from the Hyperswitch repository.
Dataset Summary
Total Examples: 2277
Language: Rust
Source: Hyperswitch GitHub repository
Format: Prompt-response pairs for supervised fine-tuning (SFT)
Data Fields
prompt: The commit message describing the change
response: The git patch/diff showing the actual code changes… See the full description on the dataset page: https://huggingface.co/datasets/archit11/hyperswitch-rust-commits-final2.
