datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Instella-Long
Instella-Long
The Instella-Long dataset is a collection of pre-training and instruction following data that is used to train Instella-3B-Long-Instruct. The pre-training data is sourced from Prolong. For the SFT data, we use public datasets: Ultrachat 200K, OpenMathinstruct-2, Tülu-3 Instruction Following, and MMLU auxiliary train set. In addition, we generate synthetic long instruction data using documents of the books and arxiv from our pre-training corpus and the dclm subset from… See the full description on the dataset page: https://huggingface.co/datasets/amd/Instella-Long.amd-nrtransformers_daily_citransformers_pr_ciAM-DeepSeek-R1-Distilled-1.4MFor more open-source datasets, models, and methodologies, please visit our GitHub repository.
AM-DeepSeek-R1-Distilled-1.4M is a large-scale general reasoning task dataset composed of
high-quality and challenging reasoning problems. These problems are collected from numerous
open-source datasets, semantically deduplicated, and cleaned to eliminate test set contamination.
All responses in the dataset are distilled from the reasoning model (mostly DeepSeek-R1) and have undergone
rigorous… See the full description on the dataset page: https://huggingface.co/datasets/a-m-team/AM-DeepSeek-R1-Distilled-1.4M.AM-DeepSeek-Distilled-40MFor more open-source datasets, models, and methodologies, please visit our GitHub repository and paper: DeepDistill: Enhancing LLM Reasoning Capabilities via Large-Scale Difficulty-Graded Data Training.
Due to certain constraints, we are only able to open-source a subset of the complete dataset.
Model Training Performance based on our complete dataset
On AIME 2024, our 72B model achieved a score of 79.2 using only supervised fine-tuning (SFT). The 32B model reached 75.8 and… See the full description on the dataset page: https://huggingface.co/datasets/a-m-team/AM-DeepSeek-Distilled-40M.AM-DeepSeek-R1-0528-Distilled
📘 Dataset Summary
This dataset is a high-quality reasoning corpus distilled from DeepSeek-R1-0528, an improved version of the DeepSeek-R1 large language model. Compared to its initial release, DeepSeek-R1-0528 demonstrates significant advances in reasoning, instruction following, and multi-turn dialogue. Motivated by these improvements, we collected and distilled a diverse set of 2.6 million queries across multiple domains, using DeepSeek-R1-0528 as the teacher.
A notable… See the full description on the dataset page: https://huggingface.co/datasets/a-m-team/AM-DeepSeek-R1-0528-Distilled.ReasonLite-Dataset
GitHub |
Dataset |
Blog
ReasonLite is an ultra-lightweight math reasoning model. With only 0.6B parameters, it leverages high-quality data distillation to achieve performance comparable to models over 10× its size, such as Qwen3-8B, reaching 75.2 on AIME24 and extending the scaling law of small models.
🔥 Best-performing 0.6B math reasoning model
🔓 Fully open-source — weights, scripts, datasets, synthesis pipeline⚙️ Distilled in two stages to balance efficiency and high… See the full description on the dataset page: https://huggingface.co/datasets/amd/ReasonLite-Dataset.AM-DeepSeek-R1-Distilled-1.4Mbert-base-uncased-2022_tokenized_datasetCot-Drop
LuminaSFT
LuminaSFT is a synthetic SFT dataset suite specifically designed to improve both general-purpose and task-specific SLMs. LuminaSFT consists of multiple curated splits that target diverse capabilities:
UltraChat200K-DeepSeek - A regenerated base SFT dataset for broad instruction following.
InstructGPT-NaturalQA and InstructGPT-TriviaQA - Factual question answering datasets to strengthen knowledge recall and answer accuracy.
CoT-Drop - A reading comprehension dataset with… See the full description on the dataset page: https://huggingface.co/datasets/amd/Cot-Drop.AIG-Datasets
AMD AIG GPU Kernel Datasets
AMD AIG-Datasets is a collection of GPU-kernel generation, translation, optimization,
and ROCm-library supervision data. It contains PyTorch/CUDA-to-HIP, HIP-to-HIP,
PyTorch-to-Triton, and production-grounded rocBLAS/rocSOLVER entries, together with
metadata, samples, conversion utilities, and reproducible evaluation tools.
The repository is organized into versioned releases. New training and evaluation
workflows should use the unified-schema datasets… See the full description on the dataset page: https://huggingface.co/datasets/amd/AIG-Datasets.SAND-MATH
SAND-Math: A Synthetic Dataset of Difficult Problems to Elevate LLM Math Performance
📃 Paper | 🤗 Dataset
SAND-Math (Synthetic Augmented Novel and Difficult Mathematics) is a high-quality, high-difficulty dataset of mathematics problems and solutions. It is generated using a novel pipeline that addresses the critical bottleneck of scarce, high-difficulty training data for mathematical Large Language Models (LLMs).
Key Features
Novel Problem Generation: Problems are… See the full description on the dataset page: https://huggingface.co/datasets/amd/SAND-MATH.AM-DeepSeek-R1-Distilled-1.4MFor more open-source datasets, models, and methodologies, please visit our GitHub repository.
AM-DeepSeek-R1-Distilled-1.4M is a large-scale general reasoning task dataset composed of
high-quality and challenging reasoning problems. These problems are collected from numerous
open-source datasets, semantically deduplicated, and cleaned to eliminate test set contamination.
All responses in the dataset are distilled from the reasoning model (mostly DeepSeek-R1) and have undergone
rigorous… See the full description on the dataset page: https://huggingface.co/datasets/abcdefj123/AM-DeepSeek-R1-Distilled-1.4M.AudioJailbreak
Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models
AudioJailbreak is a benchmark framework specifically designed for evaluating the security of Audio Language Models (Audio LLMs). This project tests model defenses against malicious requests through various audio perturbation techniques.Note: This project aims to improve the security of audio language models. Researchers should use this tool responsibly.
📋 Table of… See the full description on the dataset page: https://huggingface.co/datasets/amd1234567/AudioJailbreak.hackathon_amd_mission2_black_sortThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 155,
"total_frames": 48397,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:155"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/giacomoran/hackathon_amd_mission2_black_sort.AM-DeepSeek-R1-0528-Distilled-with-SystemInstella-GSM8K-synthetic
Instella-GSM8K-synthetic
The Instella-GSM8K-synthetic dataset was used in the second stage pre-training of Instella-3B model, which was trained on top of the Instella-3B-Stage1 model.
This synthetic dataset was generated using the training set of GSM8k dataset, where we first used Qwen2.5-72B-Instruct to
Abstract numerical values as function parameters and generate a Python program to solve the math question.
Identify and replace numerical values in the existing question with… See the full description on the dataset page: https://huggingface.co/datasets/amd/Instella-GSM8K-synthetic.AM-DeepSeek-R1-Filtered-Math-Code
AM DeepSeek R1 Filtered Math and Code
This repository publishes reproducible training subsets derived from
a-m-team/AM-DeepSeek-R1-Distilled-1.4M
at pinned revision 53531c06634904118a2dcd83961918c4d69d1cdf.
The source dataset and this derived release use CC BY-NC 4.0. Commercial use is
not permitted by that license. Preserve attribution and review the upstream
dataset card before use.
Contents
Config / split
Records
Bytes
SHA-256
math / train
111,657
2… See the full description on the dataset page: https://huggingface.co/datasets/chhao/AM-DeepSeek-R1-Filtered-Math-Code.flores_plusAM-DeepSeek-R1-Distilled-1.4M-Englishmories-caeba-amd64
🚀 Mories-CAEBA & Zenith Production Release Guide (Linux AMD64)
완전 에어갭(Air-Gapped) x86_64(AMD64) Linux 서버용 프로덕션 배포 패키지Mories 인지 지식 그래프, CAEBA 오케스트레이터, Zenith Vue 3 대시보드, NATS JetStream, Keycloak IAM, Redis 및 스탠드얼론 MCP 서버 일체 포함
📦 1. 배포 패키지 구성 요소 (7대 도커 이미지 & 설정)
파일 / 디렉토리
설명
세부 정보
docker-compose.yml
7대 마이크로서비스 오케스트레이션 구성 파일
pull_policy: never, 포트 충돌 방지, AUTH_DISABLED 지원
.env.example
프로덕션 환경변수 템플릿
포트, DB 인증 정보, LLM/Embedding 설정
docker_images/
순수… See the full description on the dataset page: https://huggingface.co/datasets/VanyaJ/mories-caeba-amd64.telephony-amd-dataset
Telephony AMD (Answering Machine Detection) Dataset
Overview
A multilingual 4-class telephony audio classification dataset for training streaming Answering Machine Detection models. Contains real human speech (PolyAI/MINDS14) mixed with TTS-generated audio (Microsoft Neural TTS / edge-tts) across English, French, Spanish, and German.
Key design principle: Voicemail greetings are recorded by real humans and sound acoustically identical to live speech. This dataset includes… See the full description on the dataset page: https://huggingface.co/datasets/AbijahKaj/telephony-amd-dataset.SAND-Post-Training-Dataset
SAND-Post-Training-Dataset: High-Quality Synthetic Reasoning Dataset Built with AMD GPUs
Dataset Summary
We introduce the SAND-Post-Training-Dataset, a high-quality synthetic reasoning dataset for mathematics and science built entirely using a synthetic data pipeline running on the AMD ROCm™ stack and AMD Instinct™ MI325 GPUs.
This dataset prioritizes difficulty and novelty over volume, demonstrating that high-difficulty synthetic data can elevate… See the full description on the dataset page: https://huggingface.co/datasets/amd/SAND-Post-Training-Dataset.AMD-SFT-Mix_3.5M
AMD-SFT-Mix_3.5M
3,556,428 SFT conversations — five AMD instruction-tuning datasets merged
into a single pre-shuffled stream, with per-row provenance so any example can
be traced back to its source.
Nothing was regenerated: this is a normalisation, provenance and shuffling
pass over existing public datasets. All credit for the data belongs to AMD.
Composition
source
rows
share
origin
naturalqa
1,304,792
36.69%
amd/InstructGpt-NaturalQa
triviaqa
1,118… See the full description on the dataset page: https://huggingface.co/datasets/Yxanul/AMD-SFT-Mix_3.5M.xbmu_amdo31
Dataset Card for [XBMU-AMDO31]
Dataset Summary
XBMU-AMDO31 dataset is a speech recognition corpus of Amdo Tibetan dialect. The open source corpus contains 31 hours of speech data and resources related to build speech recognition systems, including transcribed texts and a Tibetan pronunciation dictionary.
Supported Tasks and Leaderboards
automatic-speech-recognition: The dataset can be used to train a model for Amdo Tibetan Automatic Speech Recognition (ASR). It… See the full description on the dataset page: https://huggingface.co/datasets/syzym/xbmu_amdo31.asm_cuda_to_amdAM-DeepSeek-R1-Distilled-1.4M-am_0.5Mamd-test53This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "eyou_ft7_follower",
"total_episodes": 3,
"total_frames": 983,
"total_tasks": 1,
"total_videos": 9,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:3"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/1g0rrr/amd-test53.TTT-Bench
TTT-Bench: A Benchmark for Evaluating Reasoning Ability with Simple and Novel Tic-Tac-Toe-style Games
📃 Paper | 🤗 Dataset | 🌐 Website
We introduce TTT-Bench, a new benchmark specifically created to evaluate the reasoning capability of LRMs through a suite of simple and novel two-player Tic-Tac-Toe-style games.
Although trivial for humans, these games require basic strategic reasoning, including predicting an opponent's intentions and understanding spatial configurations.… See the full description on the dataset page: https://huggingface.co/datasets/amd/TTT-Bench.
