datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AM-DeepSeek-R1-Distilled-1.4MFor more open-source datasets, models, and methodologies, please visit our GitHub repository.
AM-DeepSeek-R1-Distilled-1.4M is a large-scale general reasoning task dataset composed of
high-quality and challenging reasoning problems. These problems are collected from numerous
open-source datasets, semantically deduplicated, and cleaned to eliminate test set contamination.
All responses in the dataset are distilled from the reasoning model (mostly DeepSeek-R1) and have undergone
rigorous… See the full description on the dataset page: https://huggingface.co/datasets/a-m-team/AM-DeepSeek-R1-Distilled-1.4M.AM-DeepSeek-Distilled-40MFor more open-source datasets, models, and methodologies, please visit our GitHub repository and paper: DeepDistill: Enhancing LLM Reasoning Capabilities via Large-Scale Difficulty-Graded Data Training.
Due to certain constraints, we are only able to open-source a subset of the complete dataset.
Model Training Performance based on our complete dataset
On AIME 2024, our 72B model achieved a score of 79.2 using only supervised fine-tuning (SFT). The 32B model reached 75.8 and… See the full description on the dataset page: https://huggingface.co/datasets/a-m-team/AM-DeepSeek-Distilled-40M.AM-DeepSeek-R1-0528-Distilled
📘 Dataset Summary
This dataset is a high-quality reasoning corpus distilled from DeepSeek-R1-0528, an improved version of the DeepSeek-R1 large language model. Compared to its initial release, DeepSeek-R1-0528 demonstrates significant advances in reasoning, instruction following, and multi-turn dialogue. Motivated by these improvements, we collected and distilled a diverse set of 2.6 million queries across multiple domains, using DeepSeek-R1-0528 as the teacher.
A notable… See the full description on the dataset page: https://huggingface.co/datasets/a-m-team/AM-DeepSeek-R1-0528-Distilled.Cot-Drop
LuminaSFT
LuminaSFT is a synthetic SFT dataset suite specifically designed to improve both general-purpose and task-specific SLMs. LuminaSFT consists of multiple curated splits that target diverse capabilities:
UltraChat200K-DeepSeek - A regenerated base SFT dataset for broad instruction following.
InstructGPT-NaturalQA and InstructGPT-TriviaQA - Factual question answering datasets to strengthen knowledge recall and answer accuracy.
CoT-Drop - A reading comprehension dataset with… See the full description on the dataset page: https://huggingface.co/datasets/amd/Cot-Drop.AIG-Datasets
AMD AIG GPU Kernel Datasets
AMD AIG-Datasets is a collection of GPU-kernel generation, translation, optimization,
and ROCm-library supervision data. It contains PyTorch/CUDA-to-HIP, HIP-to-HIP,
PyTorch-to-Triton, and production-grounded rocBLAS/rocSOLVER entries, together with
metadata, samples, conversion utilities, and reproducible evaluation tools.
The repository is organized into versioned releases. New training and evaluation
workflows should use the unified-schema datasets… See the full description on the dataset page: https://huggingface.co/datasets/amd/AIG-Datasets.AM-DeepSeek-R1-Distilled-1.4MFor more open-source datasets, models, and methodologies, please visit our GitHub repository.
AM-DeepSeek-R1-Distilled-1.4M is a large-scale general reasoning task dataset composed of
high-quality and challenging reasoning problems. These problems are collected from numerous
open-source datasets, semantically deduplicated, and cleaned to eliminate test set contamination.
All responses in the dataset are distilled from the reasoning model (mostly DeepSeek-R1) and have undergone
rigorous… See the full description on the dataset page: https://huggingface.co/datasets/abcdefj123/AM-DeepSeek-R1-Distilled-1.4M.AM-DeepSeek-R1-Filtered-Math-Code
AM DeepSeek R1 Filtered Math and Code
This repository publishes reproducible training subsets derived from
a-m-team/AM-DeepSeek-R1-Distilled-1.4M
at pinned revision 53531c06634904118a2dcd83961918c4d69d1cdf.
The source dataset and this derived release use CC BY-NC 4.0. Commercial use is
not permitted by that license. Preserve attribution and review the upstream
dataset card before use.
Contents
Config / split
Records
Bytes
SHA-256
math / train
111,657
2… See the full description on the dataset page: https://huggingface.co/datasets/chhao/AM-DeepSeek-R1-Filtered-Math-Code.Instella-GSM8K-synthetic
Instella-GSM8K-synthetic
The Instella-GSM8K-synthetic dataset was used in the second stage pre-training of Instella-3B model, which was trained on top of the Instella-3B-Stage1 model.
This synthetic dataset was generated using the training set of GSM8k dataset, where we first used Qwen2.5-72B-Instruct to
Abstract numerical values as function parameters and generate a Python program to solve the math question.
Identify and replace numerical values in the existing question with… See the full description on the dataset page: https://huggingface.co/datasets/amd/Instella-GSM8K-synthetic.AMD-SFT-Mix_3.5M
AMD-SFT-Mix_3.5M
3,556,428 SFT conversations — five AMD instruction-tuning datasets merged
into a single pre-shuffled stream, with per-row provenance so any example can
be traced back to its source.
Nothing was regenerated: this is a normalisation, provenance and shuffling
pass over existing public datasets. All credit for the data belongs to AMD.
Composition
source
rows
share
origin
naturalqa
1,304,792
36.69%
amd/InstructGpt-NaturalQa
triviaqa
1,118… See the full description on the dataset page: https://huggingface.co/datasets/Yxanul/AMD-SFT-Mix_3.5M.InstructGpt-TriviaQa
LuminaSFT
LuminaSFT is a synthetic SFT dataset suite specifically designed to improve both general-purpose and task-specific SLMs. LuminaSFT consists of multiple curated splits that target diverse capabilities:
UltraChat200K-DeepSeek - A regenerated base SFT dataset for broad instruction following.
InstructGPT-NaturalQA and InstructGPT-TriviaQA - Factual question answering datasets to strengthen knowledge recall and answer accuracy.
CoT-Drop - A reading comprehension dataset with… See the full description on the dataset page: https://huggingface.co/datasets/amd/InstructGpt-TriviaQa.UltraChat200K-regenerated
LuminaSFT
LuminaSFT is a synthetic SFT dataset suite specifically designed to improve both general-purpose and task-specific SLMs. LuminaSFT consists of multiple curated splits that target diverse capabilities:
UltraChat200K-DeepSeek - A regenerated base SFT dataset for broad instruction following.
InstructGPT-NaturalQA and InstructGPT-TriviaQA - Factual question answering datasets to strengthen knowledge recall and answer accuracy.
CoT-Drop - A reading comprehension dataset with… See the full description on the dataset page: https://huggingface.co/datasets/amd/UltraChat200K-regenerated.InstructGpt-NaturalQa
LuminaSFT
LuminaSFT is a synthetic SFT dataset suite specifically designed to improve both general-purpose and task-specific SLMs. LuminaSFT consists of multiple curated splits that target diverse capabilities:
UltraChat200K-DeepSeek - A regenerated base SFT dataset for broad instruction following.
InstructGPT-NaturalQA and InstructGPT-TriviaQA - Factual question answering datasets to strengthen knowledge recall and answer accuracy.
CoT-Drop - A reading comprehension dataset with… See the full description on the dataset page: https://huggingface.co/datasets/amd/InstructGpt-NaturalQa.InstructGpt-educational
LuminaSFT
LuminaSFT is a synthetic SFT dataset suite specifically designed to improve both general-purpose and task-specific SLMs. LuminaSFT consists of multiple curated splits that target diverse capabilities:
UltraChat200K-DeepSeek - A regenerated base SFT dataset for broad instruction following.
InstructGPT-NaturalQA and InstructGPT-TriviaQA - Factual question answering datasets to strengthen knowledge recall and answer accuracy.
CoT-Drop - A reading comprehension dataset with… See the full description on the dataset page: https://huggingface.co/datasets/amd/InstructGpt-educational.AM-DeepSeek-R1-Distilled-1.4MFor more open-source datasets, models, and methodologies, please visit our GitHub repository.
AM-DeepSeek-R1-Distilled-1.4M is a large-scale general reasoning task dataset composed of
high-quality and challenging reasoning problems. These problems are collected from numerous
open-source datasets, semantically deduplicated, and cleaned to eliminate test set contamination.
All responses in the dataset are distilled from the reasoning model (mostly DeepSeek-R1) and have undergone
rigorous… See the full description on the dataset page: https://huggingface.co/datasets/TOAO-Killer/AM-DeepSeek-R1-Distilled-1.4M.OncoAgent-Clinical-266K
🧬 OncoAgent Clinical Dataset — 266K
Curated Multi-Source Oncology Training Dataset
AMD Developer Hackathon 2026 · Used to fine-tune OncoAgent v1.0
Dataset Description
This dataset contains 266,854 clinical oncology training samples curated for fine-tuning large language models on cancer diagnosis, treatment recommendation, and clinical reasoning tasks.
Composition
Source
Samples
Description
PMC-Patients
~100,000
Real clinical case presentations… See the full description on the dataset page: https://huggingface.co/datasets/lablab-ai-amd-developer-hackathon/OncoAgent-Clinical-266K.am-deepseek-r1-distilled-prompts-1.4m
AM DeepSeek R1 Distilled Prompts 1.4M
This dataset contains prompt-only rows extracted from a-m-team/AM-DeepSeek-R1-Distilled-1.4M.
Extraction
For each source JSONL row, every message with role == "user" was emitted as one prompt row. Assistant responses, reasoning traces, answers, and source metadata were not included.
Source files:
am_0.5M.jsonl.zst
am_0.9M.jsonl.zst
Extraction results:
Input rows: 1,400,000
Output prompt rows: 1,400,000
JSON parse errors: 0… See the full description on the dataset page: https://huggingface.co/datasets/tim-gabie/am-deepseek-r1-distilled-prompts-1.4m.am-distilled-mini
AM Distilled Mini
A compact reasoning dataset derived from a-m-team/AM-DeepSeek-R1-0528-Distilled. Each row contains id, source, question, steps, and answer; steps is a nonempty list of reasoning-step strings. Only verified single-turn conversations whose tagged response agrees with the upstream reasoning and answer metadata are retained. Instances with fewer than 3 or more than 50 reasoning steps are excluded, every step is whitespace-stripped, and examples must satisfy the… See the full description on the dataset page: https://huggingface.co/datasets/cs-giung/am-distilled-mini.AM-DeepSeek-R1-0528-Distilled
📘 Dataset Summary
This dataset is a high-quality reasoning corpus distilled from DeepSeek-R1-0528, an improved version of the DeepSeek-R1 large language model. Compared to its initial release, DeepSeek-R1-0528 demonstrates significant advances in reasoning, instruction following, and multi-turn dialogue. Motivated by these improvements, we collected and distilled a diverse set of 2.6 million queries across multiple domains, using DeepSeek-R1-0528 as the teacher.
A notable… See the full description on the dataset page: https://huggingface.co/datasets/kabsis/AM-DeepSeek-R1-0528-Distilled.long_context_predictable_dataset
Long Context Predictable Dataset
A dataset of long-context editing and translation prompts built from Project Gutenberg texts.
Description
Each example consists of a task instruction (prompt) prepended to a long passage of text (~549,000 words per passage). The tasks are designed to require long output, such as translating, rewriting, or editing the full text.
Dataset Structure
Each example contains the following fields:
Field
Type
Description
prompt… See the full description on the dataset page: https://huggingface.co/datasets/rkarhila-amd/long_context_predictable_dataset.wiki
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/alexei-v-ivanov-amd/wiki.
