CoolFace
20 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01a-m-team /AM-DeepSeek-R1-Distilled-1.4MFor more open-source datasets, models, and methodologies, please visit our GitHub repository. AM-DeepSeek-R1-Distilled-1.4M is a large-scale general reasoning task dataset composed of high-quality and challenging reasoning problems. These problems are collected from numerous open-source datasets, semantically deduplicated, and cleaned to eliminate test set contamination. All responses in the dataset are distilled from the reasoning model (mostly DeepSeek-R1) and have undergone rigorous… See the full description on the dataset page: https://huggingface.co/datasets/a-m-team/AM-DeepSeek-R1-Distilled-1.4M.text-generation1M<n<10M184 likes2.2k downloads1y agoHugging Face02a-m-team /AM-DeepSeek-Distilled-40MFor more open-source datasets, models, and methodologies, please visit our GitHub repository and paper: DeepDistill: Enhancing LLM Reasoning Capabilities via Large-Scale Difficulty-Graded Data Training. Due to certain constraints, we are only able to open-source a subset of the complete dataset. Model Training Performance based on our complete dataset On AIME 2024, our 72B model achieved a score of 79.2 using only supervised fine-tuning (SFT). The 32B model reached 75.8 and… See the full description on the dataset page: https://huggingface.co/datasets/a-m-team/AM-DeepSeek-Distilled-40M.tabulartext-generation10M<n<100M56 likes2.1k downloads1y agoHugging Face03a-m-team /AM-DeepSeek-R1-0528-Distilled 📘 Dataset Summary This dataset is a high-quality reasoning corpus distilled from DeepSeek-R1-0528, an improved version of the DeepSeek-R1 large language model. Compared to its initial release, DeepSeek-R1-0528 demonstrates significant advances in reasoning, instruction following, and multi-turn dialogue. Motivated by these improvements, we collected and distilled a diverse set of 2.6 million queries across multiple domains, using DeepSeek-R1-0528 as the teacher. A notable… See the full description on the dataset page: https://huggingface.co/datasets/a-m-team/AM-DeepSeek-R1-0528-Distilled.text-generation1M<n<10M102 likes1k downloads1y agoHugging Face04amd /Cot-Drop LuminaSFT LuminaSFT is a synthetic SFT dataset suite specifically designed to improve both general-purpose and task-specific SLMs. LuminaSFT consists of multiple curated splits that target diverse capabilities: UltraChat200K-DeepSeek - A regenerated base SFT dataset for broad instruction following. InstructGPT-NaturalQA and InstructGPT-TriviaQA - Factual question answering datasets to strengthen knowledge recall and answer accuracy. CoT-Drop - A reading comprehension dataset with… See the full description on the dataset page: https://huggingface.co/datasets/amd/Cot-Drop.texttext-generation10K<n<100K1 likes349 downloads7mo agoHugging Face05amd /AIG-Datasets AMD AIG GPU Kernel Datasets AMD AIG-Datasets is a collection of GPU-kernel generation, translation, optimization, and ROCm-library supervision data. It contains PyTorch/CUDA-to-HIP, HIP-to-HIP, PyTorch-to-Triton, and production-grounded rocBLAS/rocSOLVER entries, together with metadata, samples, conversion utilities, and reproducible evaluation tools. The repository is organized into versioned releases. New training and evaluation workflows should use the unified-schema datasets… See the full description on the dataset page: https://huggingface.co/datasets/amd/AIG-Datasets.text-generation4 likes340 downloads29d agoHugging Face06abcdefj123 /AM-DeepSeek-R1-Distilled-1.4MFor more open-source datasets, models, and methodologies, please visit our GitHub repository. AM-DeepSeek-R1-Distilled-1.4M is a large-scale general reasoning task dataset composed of high-quality and challenging reasoning problems. These problems are collected from numerous open-source datasets, semantically deduplicated, and cleaned to eliminate test set contamination. All responses in the dataset are distilled from the reasoning model (mostly DeepSeek-R1) and have undergone rigorous… See the full description on the dataset page: https://huggingface.co/datasets/abcdefj123/AM-DeepSeek-R1-Distilled-1.4M.text-generation1M<n<10M2 likes206 downloads2mo agoHugging Face07chhao /AM-DeepSeek-R1-Filtered-Math-Code AM DeepSeek R1 Filtered Math and Code This repository publishes reproducible training subsets derived from a-m-team/AM-DeepSeek-R1-Distilled-1.4M at pinned revision 53531c06634904118a2dcd83961918c4d69d1cdf. The source dataset and this derived release use CC BY-NC 4.0. Commercial use is not permitted by that license. Preserve attribution and review the upstream dataset card before use. Contents Config / split Records Bytes SHA-256 math / train 111,657 2… See the full description on the dataset page: https://huggingface.co/datasets/chhao/AM-DeepSeek-R1-Filtered-Math-Code.texttext-generation100K<n<1M0 likes169 downloads2mo agoHugging Face08amd /Instella-GSM8K-synthetic Instella-GSM8K-synthetic The Instella-GSM8K-synthetic dataset was used in the second stage pre-training of Instella-3B model, which was trained on top of the Instella-3B-Stage1 model. This synthetic dataset was generated using the training set of GSM8k dataset, where we first used Qwen2.5-72B-Instruct to Abstract numerical values as function parameters and generate a Python program to solve the math question. Identify and replace numerical values in the existing question with… See the full description on the dataset page: https://huggingface.co/datasets/amd/Instella-GSM8K-synthetic.textquestion-answering1M<n<10M7 likes168 downloads10mo agoHugging Face09Yxanul /AMD-SFT-Mix_3.5M AMD-SFT-Mix_3.5M 3,556,428 SFT conversations — five AMD instruction-tuning datasets merged into a single pre-shuffled stream, with per-row provenance so any example can be traced back to its source. Nothing was regenerated: this is a normalisation, provenance and shuffling pass over existing public datasets. All credit for the data belongs to AMD. Composition source rows share origin naturalqa 1,304,792 36.69% amd/InstructGpt-NaturalQa triviaqa 1,118… See the full description on the dataset page: https://huggingface.co/datasets/Yxanul/AMD-SFT-Mix_3.5M.texttext-generation1M<n<10M0 likes111 downloads2mo agoHugging Face10amd /InstructGpt-TriviaQa LuminaSFT LuminaSFT is a synthetic SFT dataset suite specifically designed to improve both general-purpose and task-specific SLMs. LuminaSFT consists of multiple curated splits that target diverse capabilities: UltraChat200K-DeepSeek - A regenerated base SFT dataset for broad instruction following. InstructGPT-NaturalQA and InstructGPT-TriviaQA - Factual question answering datasets to strengthen knowledge recall and answer accuracy. CoT-Drop - A reading comprehension dataset with… See the full description on the dataset page: https://huggingface.co/datasets/amd/InstructGpt-TriviaQa.texttext-generation1M<n<10M0 likes98 downloads7mo agoHugging Face11amd /UltraChat200K-regenerated LuminaSFT LuminaSFT is a synthetic SFT dataset suite specifically designed to improve both general-purpose and task-specific SLMs. LuminaSFT consists of multiple curated splits that target diverse capabilities: UltraChat200K-DeepSeek - A regenerated base SFT dataset for broad instruction following. InstructGPT-NaturalQA and InstructGPT-TriviaQA - Factual question answering datasets to strengthen knowledge recall and answer accuracy. CoT-Drop - A reading comprehension dataset with… See the full description on the dataset page: https://huggingface.co/datasets/amd/UltraChat200K-regenerated.texttext-generation100K<n<1M2 likes91 downloads7mo agoHugging Face12amd /InstructGpt-NaturalQa LuminaSFT LuminaSFT is a synthetic SFT dataset suite specifically designed to improve both general-purpose and task-specific SLMs. LuminaSFT consists of multiple curated splits that target diverse capabilities: UltraChat200K-DeepSeek - A regenerated base SFT dataset for broad instruction following. InstructGPT-NaturalQA and InstructGPT-TriviaQA - Factual question answering datasets to strengthen knowledge recall and answer accuracy. CoT-Drop - A reading comprehension dataset with… See the full description on the dataset page: https://huggingface.co/datasets/amd/InstructGpt-NaturalQa.texttext-generation1M<n<10M1 likes90 downloads7mo agoHugging Face13amd /InstructGpt-educational LuminaSFT LuminaSFT is a synthetic SFT dataset suite specifically designed to improve both general-purpose and task-specific SLMs. LuminaSFT consists of multiple curated splits that target diverse capabilities: UltraChat200K-DeepSeek - A regenerated base SFT dataset for broad instruction following. InstructGPT-NaturalQA and InstructGPT-TriviaQA - Factual question answering datasets to strengthen knowledge recall and answer accuracy. CoT-Drop - A reading comprehension dataset with… See the full description on the dataset page: https://huggingface.co/datasets/amd/InstructGpt-educational.texttext-generation100K<n<1M3 likes84 downloads7mo agoHugging Face14TOAO-Killer /AM-DeepSeek-R1-Distilled-1.4MFor more open-source datasets, models, and methodologies, please visit our GitHub repository. AM-DeepSeek-R1-Distilled-1.4M is a large-scale general reasoning task dataset composed of high-quality and challenging reasoning problems. These problems are collected from numerous open-source datasets, semantically deduplicated, and cleaned to eliminate test set contamination. All responses in the dataset are distilled from the reasoning model (mostly DeepSeek-R1) and have undergone rigorous… See the full description on the dataset page: https://huggingface.co/datasets/TOAO-Killer/AM-DeepSeek-R1-Distilled-1.4M.text-generation1M<n<10M1 likes35 downloads8mo agoHugging Face15lablab-ai-amd-developer-hackathon /OncoAgent-Clinical-266K 🧬 OncoAgent Clinical Dataset — 266K Curated Multi-Source Oncology Training Dataset AMD Developer Hackathon 2026 · Used to fine-tune OncoAgent v1.0 Dataset Description This dataset contains 266,854 clinical oncology training samples curated for fine-tuning large language models on cancer diagnosis, treatment recommendation, and clinical reasoning tasks. Composition Source Samples Description PMC-Patients ~100,000 Real clinical case presentations… See the full description on the dataset page: https://huggingface.co/datasets/lablab-ai-amd-developer-hackathon/OncoAgent-Clinical-266K.text-generation100K<n<1M6 likes21 downloads5mo agoHugging Face16tim-gabie /am-deepseek-r1-distilled-prompts-1.4m AM DeepSeek R1 Distilled Prompts 1.4M This dataset contains prompt-only rows extracted from a-m-team/AM-DeepSeek-R1-Distilled-1.4M. Extraction For each source JSONL row, every message with role == "user" was emitted as one prompt row. Assistant responses, reasoning traces, answers, and source metadata were not included. Source files: am_0.5M.jsonl.zst am_0.9M.jsonl.zst Extraction results: Input rows: 1,400,000 Output prompt rows: 1,400,000 JSON parse errors: 0… See the full description on the dataset page: https://huggingface.co/datasets/tim-gabie/am-deepseek-r1-distilled-prompts-1.4m.texttext-generation1M<n<10M0 likes19 downloads3mo agoHugging Face17cs-giung /am-distilled-mini AM Distilled Mini A compact reasoning dataset derived from a-m-team/AM-DeepSeek-R1-0528-Distilled. Each row contains id, source, question, steps, and answer; steps is a nonempty list of reasoning-step strings. Only verified single-turn conversations whose tagged response agrees with the upstream reasoning and answer metadata are retained. Instances with fewer than 3 or more than 50 reasoning steps are excluded, every step is whitespace-stripped, and examples must satisfy the… See the full description on the dataset page: https://huggingface.co/datasets/cs-giung/am-distilled-mini.texttext-generation100K<n<1M0 likes16 downloads2mo agoHugging Face18kabsis /AM-DeepSeek-R1-0528-Distilled 📘 Dataset Summary This dataset is a high-quality reasoning corpus distilled from DeepSeek-R1-0528, an improved version of the DeepSeek-R1 large language model. Compared to its initial release, DeepSeek-R1-0528 demonstrates significant advances in reasoning, instruction following, and multi-turn dialogue. Motivated by these improvements, we collected and distilled a diverse set of 2.6 million queries across multiple domains, using DeepSeek-R1-0528 as the teacher. A notable… See the full description on the dataset page: https://huggingface.co/datasets/kabsis/AM-DeepSeek-R1-0528-Distilled.text-generation1M<n<10M0 likes11 downloads9mo agoHugging Face19rkarhila-amd /long_context_predictable_dataset Long Context Predictable Dataset A dataset of long-context editing and translation prompts built from Project Gutenberg texts. Description Each example consists of a task instruction (prompt) prepended to a long passage of text (~549,000 words per passage). The tasks are designed to require long output, such as translating, rewriting, or editing the full text. Dataset Structure Each example contains the following fields: Field Type Description prompt… See the full description on the dataset page: https://huggingface.co/datasets/rkarhila-amd/long_context_predictable_dataset.texttext-generationn<1K0 likes9 downloads7mo agoHugging Face20alexei-v-ivanov-amd /wiki Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/alexei-v-ivanov-amd/wiki.audiotext-generationn<1K0 likes5 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.