datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Rust-Coder
Rust-Coder
Rust-Coder is a comprehensive text dataset designed for Rust programming language learning. It contains 12,000 unique samples focusing on distinct Rust concepts, code snippets, and explanations.
Dataset Structure
Each sample consists of:
id: A unique UUID.
instruction: A prompt or question about a Rust concept.
code: An idiomatic Rust code snippet.
explanation: A detailed explanation of the concept and code.
category: The high-level Rust category (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/Convence/Rust-Coder.Rust-Coder
Rust-Coder
Rust-Coder is a comprehensive text dataset designed for Rust programming language learning. It contains 12,000 unique samples focusing on distinct Rust concepts, code snippets, and explanations.
Dataset Structure
Each sample consists of:
id: A unique UUID.
instruction: A prompt or question about a Rust concept.
code: An idiomatic Rust code snippet.
explanation: A detailed explanation of the concept and code.
category: The high-level Rust category (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/gubernac/Rust-Coder.code-alchemy-rust
CodeAlchemy Rust
Rust-only derivative of open-alchemy/code-alchemy. It preserves the five training configs, two evaluation configs, original splits, row order, columns, values, and task/evaluation fields.
Rows were selected from the source-native language labels:
Rust and rust in training data and dev-eval
rs in trace-eval
Labels remain unchanged in the output. code-trace.external_packages is normalized to list<string> because source Parquet shards physically alternate between… See the full description on the dataset page: https://huggingface.co/datasets/adityabhushannagar/code-alchemy-rust.rust-forum-qa-pairs
Rust Programming QA Pairs
Dataset Description
The Rust Programming QA Pairs dataset is a collection of question-answer pairs extracted from the Rust programming language user forums. It contains high-quality programming questions and their accepted answers, focusing on Rust programming language topics. The dataset is designed to support natural language processing tasks related to programming assistance, code understanding, and technical question answering.
Each entry… See the full description on the dataset page: https://huggingface.co/datasets/portex/rust-forum-qa-pairs.tauri2-svelte5-rust-instruct
🦀 Tauri v2 + Svelte 5 (Runes) + Rust Instruct Dataset
Специализированный датасет на русском языке для дообучения LLM актуальному стеку разработки десктопных приложений (2024-2025).
Главная проблема большинства моделей — галлюцинации по поводу устаревшего синтаксиса (Svelte 4, Tauri v1). Этот датасет решает проблему, предоставляя примеры с использованием Svelte 5 Runes и нового IPC в Tauri v2.
📊 О датасете
Объем: 566 пар "Инструкция — Решение".
Фокус: Создание UI на… See the full description on the dataset page: https://huggingface.co/datasets/oxide-lab/tauri2-svelte5-rust-instruct.ru-stem-dialogues
Russian STEM Educational Dialogues
Описание
Синтетический датасет русскоязычных учебных диалогов по STEM-темам (математика, физика, химия,
биология, информатика, программирование, инженерия). Каждый диалог — реалистичное взаимодействие
между пользователем (школьник / студент / профессионал) и ассистентом.
Методология
Модель: Qwen/Qwen2.5-7B-Instruct (4-bit NF4 quantization, bitsandbytes)
Формат генерации: текстовый формат с разделителями… See the full description on the dataset page: https://huggingface.co/datasets/AtesiT/ru-stem-dialogues.reasoning-rust
Dataset Card for my-distiset-da5d1e20
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/beowolx/my-distiset-da5d1e20/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/beowolx/reasoning-rust.bird23-train-filtered
BIRD-SQL Train (Filtered)
A high-quality subset of the original BIRD train split for text-to-SQL finetuning.
Overview
Over the past year the community has shared many observations about data quality in BIRD. We performed a rigorous data quality check process to retain examples that are consistent with schema and faithfully answer the question. The resulting set keeps 6,601 instances out of 9,428 (≈70%), and serves as a drop-in replacement for training.
Original… See the full description on the dataset page: https://huggingface.co/datasets/RustamHuseynov/bird23-train-filtered.bird_mini_dev
BIRD-SQL Mini-Dev
Update 2025-07-04
We are grateful for the valuable feedback from the community over the past year regarding BIRD Mini-Dev. Based on your suggestions, we have made significant updates to the BIRD Mini-Dev dataset.
For New Users
If you are new to BIRD Mini-Dev, you can download the complete databases and datasets using the following link:
Download BIRD Mini-Dev Complete Package
For Existing Users
If you have already… See the full description on the dataset page: https://huggingface.co/datasets/RustamHuseynov/bird_mini_dev.
