datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AM-DeepSeek-Distilled-40MFor more open-source datasets, models, and methodologies, please visit our GitHub repository and paper: DeepDistill: Enhancing LLM Reasoning Capabilities via Large-Scale Difficulty-Graded Data Training.
Due to certain constraints, we are only able to open-source a subset of the complete dataset.
Model Training Performance based on our complete dataset
On AIME 2024, our 72B model achieved a score of 79.2 using only supervised fine-tuning (SFT). The 32B model reached 75.8 and… See the full description on the dataset page: https://huggingface.co/datasets/a-m-team/AM-DeepSeek-Distilled-40M.Cot-Drop
LuminaSFT
LuminaSFT is a synthetic SFT dataset suite specifically designed to improve both general-purpose and task-specific SLMs. LuminaSFT consists of multiple curated splits that target diverse capabilities:
UltraChat200K-DeepSeek - A regenerated base SFT dataset for broad instruction following.
InstructGPT-NaturalQA and InstructGPT-TriviaQA - Factual question answering datasets to strengthen knowledge recall and answer accuracy.
CoT-Drop - A reading comprehension dataset with… See the full description on the dataset page: https://huggingface.co/datasets/amd/Cot-Drop.SAND-MATH
SAND-Math: A Synthetic Dataset of Difficult Problems to Elevate LLM Math Performance
📃 Paper | 🤗 Dataset
SAND-Math (Synthetic Augmented Novel and Difficult Mathematics) is a high-quality, high-difficulty dataset of mathematics problems and solutions. It is generated using a novel pipeline that addresses the critical bottleneck of scarce, high-difficulty training data for mathematical Large Language Models (LLMs).
Key Features
Novel Problem Generation: Problems are… See the full description on the dataset page: https://huggingface.co/datasets/amd/SAND-MATH.SAND-Post-Training-Dataset
SAND-Post-Training-Dataset: High-Quality Synthetic Reasoning Dataset Built with AMD GPUs
Dataset Summary
We introduce the SAND-Post-Training-Dataset, a high-quality synthetic reasoning dataset for mathematics and science built entirely using a synthetic data pipeline running on the AMD ROCm™ stack and AMD Instinct™ MI325 GPUs.
This dataset prioritizes difficulty and novelty over volume, demonstrating that high-difficulty synthetic data can elevate… See the full description on the dataset page: https://huggingface.co/datasets/amd/SAND-Post-Training-Dataset.AM-DeepSeek-R1-Filtered-Math-Code
AM DeepSeek R1 Filtered Math and Code
This repository publishes reproducible training subsets derived from
a-m-team/AM-DeepSeek-R1-Distilled-1.4M
at pinned revision 53531c06634904118a2dcd83961918c4d69d1cdf.
The source dataset and this derived release use CC BY-NC 4.0. Commercial use is
not permitted by that license. Preserve attribution and review the upstream
dataset card before use.
Contents
Config / split
Records
Bytes
SHA-256
math / train
111,657
2… See the full description on the dataset page: https://huggingface.co/datasets/chhao/AM-DeepSeek-R1-Filtered-Math-Code.amdpilot-lora-sft-dataset
AMDPilot LoRA SFT Dataset
SFT training data for fine-tuning LLMs on AMD GPU debugging, optimization, and kernel engineering tasks. Each example is a multi-turn conversation in OpenAI messages format with tool-use annotations.
Usage
from datasets import load_dataset
# Load a specific version
ds = load_dataset("JinnP/amdpilot-lora-sft-dataset", "v5_2")
# Load a specific view
ds = load_dataset("JinnP/amdpilot-lora-sft-dataset", "v5_2_chunks")
# Available configs: v4, v5… See the full description on the dataset page: https://huggingface.co/datasets/JinnP/amdpilot-lora-sft-dataset.AMD-SFT-Mix_3.5M
AMD-SFT-Mix_3.5M
3,556,428 SFT conversations — five AMD instruction-tuning datasets merged
into a single pre-shuffled stream, with per-row provenance so any example can
be traced back to its source.
Nothing was regenerated: this is a normalisation, provenance and shuffling
pass over existing public datasets. All credit for the data belongs to AMD.
Composition
source
rows
share
origin
naturalqa
1,304,792
36.69%
amd/InstructGpt-NaturalQa
triviaqa
1,118… See the full description on the dataset page: https://huggingface.co/datasets/Yxanul/AMD-SFT-Mix_3.5M.TTT-Bench
TTT-Bench: A Benchmark for Evaluating Reasoning Ability with Simple and Novel Tic-Tac-Toe-style Games
📃 Paper | 🤗 Dataset | 🌐 Website
We introduce TTT-Bench, a new benchmark specifically created to evaluate the reasoning capability of LRMs through a suite of simple and novel two-player Tic-Tac-Toe-style games.
Although trivial for humans, these games require basic strategic reasoning, including predicting an opponent's intentions and understanding spatial configurations.… See the full description on the dataset page: https://huggingface.co/datasets/amd/TTT-Bench.InstructGpt-TriviaQa
LuminaSFT
LuminaSFT is a synthetic SFT dataset suite specifically designed to improve both general-purpose and task-specific SLMs. LuminaSFT consists of multiple curated splits that target diverse capabilities:
UltraChat200K-DeepSeek - A regenerated base SFT dataset for broad instruction following.
InstructGPT-NaturalQA and InstructGPT-TriviaQA - Factual question answering datasets to strengthen knowledge recall and answer accuracy.
CoT-Drop - A reading comprehension dataset with… See the full description on the dataset page: https://huggingface.co/datasets/amd/InstructGpt-TriviaQa.UltraChat200K-regenerated
LuminaSFT
LuminaSFT is a synthetic SFT dataset suite specifically designed to improve both general-purpose and task-specific SLMs. LuminaSFT consists of multiple curated splits that target diverse capabilities:
UltraChat200K-DeepSeek - A regenerated base SFT dataset for broad instruction following.
InstructGPT-NaturalQA and InstructGPT-TriviaQA - Factual question answering datasets to strengthen knowledge recall and answer accuracy.
CoT-Drop - A reading comprehension dataset with… See the full description on the dataset page: https://huggingface.co/datasets/amd/UltraChat200K-regenerated.InstructGpt-NaturalQa
LuminaSFT
LuminaSFT is a synthetic SFT dataset suite specifically designed to improve both general-purpose and task-specific SLMs. LuminaSFT consists of multiple curated splits that target diverse capabilities:
UltraChat200K-DeepSeek - A regenerated base SFT dataset for broad instruction following.
InstructGPT-NaturalQA and InstructGPT-TriviaQA - Factual question answering datasets to strengthen knowledge recall and answer accuracy.
CoT-Drop - A reading comprehension dataset with… See the full description on the dataset page: https://huggingface.co/datasets/amd/InstructGpt-NaturalQa.InstructGpt-educational
LuminaSFT
LuminaSFT is a synthetic SFT dataset suite specifically designed to improve both general-purpose and task-specific SLMs. LuminaSFT consists of multiple curated splits that target diverse capabilities:
UltraChat200K-DeepSeek - A regenerated base SFT dataset for broad instruction following.
InstructGPT-NaturalQA and InstructGPT-TriviaQA - Factual question answering datasets to strengthen knowledge recall and answer accuracy.
CoT-Drop - A reading comprehension dataset with… See the full description on the dataset page: https://huggingface.co/datasets/amd/InstructGpt-educational.AM-DeepSeek-R1-Distilled-1.4M-Pureamdpilot-lora-sft-dataset-v5
v5
Frozen release with canonical split and 3-view derivatives.
amd__AMD-Llama-135m-details
Dataset Card for Evaluation run of amd/AMD-Llama-135m
Dataset automatically created during the evaluation run of model amd/AMD-Llama-135m
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/amd__AMD-Llama-135m-details.am-deepseek-r1-distilled-prompts-1.4m
AM DeepSeek R1 Distilled Prompts 1.4M
This dataset contains prompt-only rows extracted from a-m-team/AM-DeepSeek-R1-Distilled-1.4M.
Extraction
For each source JSONL row, every message with role == "user" was emitted as one prompt row. Assistant responses, reasoning traces, answers, and source metadata were not included.
Source files:
am_0.5M.jsonl.zst
am_0.9M.jsonl.zst
Extraction results:
Input rows: 1,400,000
Output prompt rows: 1,400,000
JSON parse errors: 0… See the full description on the dataset page: https://huggingface.co/datasets/tim-gabie/am-deepseek-r1-distilled-prompts-1.4m.NetworkTraffic
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Amdeous/NetworkTraffic.autotrain-data-amdal-mining-llama2-7b-cleanamdpilot-lora-sft-dataset-v5_1
v5.1
Updated release with canonical split (89 train + 3 eval) and 3-view derivatives.
AM-Deepseek-Translate_DistilAMD-eAIrs-SFT-Demo-Data
AMD enterprise AI reference stack Supervised Fine-tuning Demo Data
This is a dataset for demonstration purposes. It is based on crawling the
AMD enterprise AI reference stack documentation and creating prompt-answer
pairs from it. Additional negative refusal examples were also generated.
This can be used to fine-tune a Large Language Model with Supervised
Fine-tuning.
Intended model learning outcome
The model fine-tuned on this data is intended to answer user… See the full description on the dataset page: https://huggingface.co/datasets/SiloAI/AMD-eAIrs-SFT-Demo-Data.hipifyplus
