CoolFace
22 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01a-m-team /AM-DeepSeek-Distilled-40MFor more open-source datasets, models, and methodologies, please visit our GitHub repository and paper: DeepDistill: Enhancing LLM Reasoning Capabilities via Large-Scale Difficulty-Graded Data Training. Due to certain constraints, we are only able to open-source a subset of the complete dataset. Model Training Performance based on our complete dataset On AIME 2024, our 72B model achieved a score of 79.2 using only supervised fine-tuning (SFT). The 32B model reached 75.8 and… See the full description on the dataset page: https://huggingface.co/datasets/a-m-team/AM-DeepSeek-Distilled-40M.tabulartext-generation10M<n<100M56 likes2.1k downloads1y agoHugging Face02amd /Cot-Drop LuminaSFT LuminaSFT is a synthetic SFT dataset suite specifically designed to improve both general-purpose and task-specific SLMs. LuminaSFT consists of multiple curated splits that target diverse capabilities: UltraChat200K-DeepSeek - A regenerated base SFT dataset for broad instruction following. InstructGPT-NaturalQA and InstructGPT-TriviaQA - Factual question answering datasets to strengthen knowledge recall and answer accuracy. CoT-Drop - A reading comprehension dataset with… See the full description on the dataset page: https://huggingface.co/datasets/amd/Cot-Drop.texttext-generation10K<n<100K2 likes349 downloads7mo agoHugging Face03amd /SAND-MATH SAND-Math: A Synthetic Dataset of Difficult Problems to Elevate LLM Math Performance 📃 Paper | 🤗 Dataset SAND-Math (Synthetic Augmented Novel and Difficult Mathematics) is a high-quality, high-difficulty dataset of mathematics problems and solutions. It is generated using a novel pipeline that addresses the critical bottleneck of scarce, high-difficulty training data for mathematical Large Language Models (LLMs). Key Features Novel Problem Generation: Problems are… See the full description on the dataset page: https://huggingface.co/datasets/amd/SAND-MATH.tabularquestion-answering10K<n<100K3 likes255 downloads11mo agoHugging Face04amd /SAND-Post-Training-Dataset SAND-Post-Training-Dataset: High-Quality Synthetic Reasoning Dataset Built with AMD GPUs Dataset Summary We introduce the SAND-Post-Training-Dataset, a high-quality synthetic reasoning dataset for mathematics and science built entirely using a synthetic data pipeline running on the AMD ROCm™ stack and AMD Instinct™ MI325 GPUs. This dataset prioritizes difficulty and novelty over volume, demonstrating that high-difficulty synthetic data can elevate… See the full description on the dataset page: https://huggingface.co/datasets/amd/SAND-Post-Training-Dataset.tabularquestion-answering10K<n<100K3 likes181 downloads10mo agoHugging Face05chhao /AM-DeepSeek-R1-Filtered-Math-Code AM DeepSeek R1 Filtered Math and Code This repository publishes reproducible training subsets derived from a-m-team/AM-DeepSeek-R1-Distilled-1.4M at pinned revision 53531c06634904118a2dcd83961918c4d69d1cdf. The source dataset and this derived release use CC BY-NC 4.0. Commercial use is not permitted by that license. Preserve attribution and review the upstream dataset card before use. Contents Config / split Records Bytes SHA-256 math / train 111,657 2… See the full description on the dataset page: https://huggingface.co/datasets/chhao/AM-DeepSeek-R1-Filtered-Math-Code.texttext-generation100K<n<1M0 likes169 downloads2mo agoHugging Face06JinnP /amdpilot-lora-sft-dataset AMDPilot LoRA SFT Dataset SFT training data for fine-tuning LLMs on AMD GPU debugging, optimization, and kernel engineering tasks. Each example is a multi-turn conversation in OpenAI messages format with tool-use annotations. Usage from datasets import load_dataset # Load a specific version ds = load_dataset("JinnP/amdpilot-lora-sft-dataset", "v5_2") # Load a specific view ds = load_dataset("JinnP/amdpilot-lora-sft-dataset", "v5_2_chunks") # Available configs: v4, v5… See the full description on the dataset page: https://huggingface.co/datasets/JinnP/amdpilot-lora-sft-dataset.1K<n<10K0 likes133 downloads5mo agoHugging Face07Yxanul /AMD-SFT-Mix_3.5M AMD-SFT-Mix_3.5M 3,556,428 SFT conversations — five AMD instruction-tuning datasets merged into a single pre-shuffled stream, with per-row provenance so any example can be traced back to its source. Nothing was regenerated: this is a normalisation, provenance and shuffling pass over existing public datasets. All credit for the data belongs to AMD. Composition source rows share origin naturalqa 1,304,792 36.69% amd/InstructGpt-NaturalQa triviaqa 1,118… See the full description on the dataset page: https://huggingface.co/datasets/Yxanul/AMD-SFT-Mix_3.5M.texttext-generation1M<n<10M0 likes111 downloads2mo agoHugging Face08amd /TTT-Bench TTT-Bench: A Benchmark for Evaluating Reasoning Ability with Simple and Novel Tic-Tac-Toe-style Games 📃 Paper | 🤗 Dataset | 🌐 Website We introduce TTT-Bench, a new benchmark specifically created to evaluate the reasoning capability of LRMs through a suite of simple and novel two-player Tic-Tac-Toe-style games. Although trivial for humans, these games require basic strategic reasoning, including predicting an opponent's intentions and understanding spatial configurations.… See the full description on the dataset page: https://huggingface.co/datasets/amd/TTT-Bench.textquestion-answeringn<1K1 likes110 downloads1y agoHugging Face09amd /InstructGpt-TriviaQa LuminaSFT LuminaSFT is a synthetic SFT dataset suite specifically designed to improve both general-purpose and task-specific SLMs. LuminaSFT consists of multiple curated splits that target diverse capabilities: UltraChat200K-DeepSeek - A regenerated base SFT dataset for broad instruction following. InstructGPT-NaturalQA and InstructGPT-TriviaQA - Factual question answering datasets to strengthen knowledge recall and answer accuracy. CoT-Drop - A reading comprehension dataset with… See the full description on the dataset page: https://huggingface.co/datasets/amd/InstructGpt-TriviaQa.texttext-generation1M<n<10M0 likes98 downloads7mo agoHugging Face10amd /UltraChat200K-regenerated LuminaSFT LuminaSFT is a synthetic SFT dataset suite specifically designed to improve both general-purpose and task-specific SLMs. LuminaSFT consists of multiple curated splits that target diverse capabilities: UltraChat200K-DeepSeek - A regenerated base SFT dataset for broad instruction following. InstructGPT-NaturalQA and InstructGPT-TriviaQA - Factual question answering datasets to strengthen knowledge recall and answer accuracy. CoT-Drop - A reading comprehension dataset with… See the full description on the dataset page: https://huggingface.co/datasets/amd/UltraChat200K-regenerated.texttext-generation100K<n<1M2 likes91 downloads7mo agoHugging Face11amd /InstructGpt-NaturalQa LuminaSFT LuminaSFT is a synthetic SFT dataset suite specifically designed to improve both general-purpose and task-specific SLMs. LuminaSFT consists of multiple curated splits that target diverse capabilities: UltraChat200K-DeepSeek - A regenerated base SFT dataset for broad instruction following. InstructGPT-NaturalQA and InstructGPT-TriviaQA - Factual question answering datasets to strengthen knowledge recall and answer accuracy. CoT-Drop - A reading comprehension dataset with… See the full description on the dataset page: https://huggingface.co/datasets/amd/InstructGpt-NaturalQa.texttext-generation1M<n<10M1 likes90 downloads7mo agoHugging Face12amd /InstructGpt-educational LuminaSFT LuminaSFT is a synthetic SFT dataset suite specifically designed to improve both general-purpose and task-specific SLMs. LuminaSFT consists of multiple curated splits that target diverse capabilities: UltraChat200K-DeepSeek - A regenerated base SFT dataset for broad instruction following. InstructGPT-NaturalQA and InstructGPT-TriviaQA - Factual question answering datasets to strengthen knowledge recall and answer accuracy. CoT-Drop - A reading comprehension dataset with… See the full description on the dataset page: https://huggingface.co/datasets/amd/InstructGpt-educational.texttext-generation100K<n<1M3 likes84 downloads7mo agoHugging Face13zhsh17 /AM-DeepSeek-R1-Distilled-1.4M-Puretext1K<n<10K0 likes68 downloads8mo agoHugging Face14JinnP /amdpilot-lora-sft-dataset-v5 v5 Frozen release with canonical split and 3-view derivatives. n<1K0 likes23 downloads6mo agoHugging Face15open-llm-leaderboard /amd__AMD-Llama-135m-detailsgated Dataset Card for Evaluation run of amd/AMD-Llama-135m Dataset automatically created during the evaluation run of model amd/AMD-Llama-135m The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/amd__AMD-Llama-135m-details.tabular10K<n<100K0 likes19 downloads2y agoHugging Face16tim-gabie /am-deepseek-r1-distilled-prompts-1.4m AM DeepSeek R1 Distilled Prompts 1.4M This dataset contains prompt-only rows extracted from a-m-team/AM-DeepSeek-R1-Distilled-1.4M. Extraction For each source JSONL row, every message with role == "user" was emitted as one prompt row. Assistant responses, reasoning traces, answers, and source metadata were not included. Source files: am_0.5M.jsonl.zst am_0.9M.jsonl.zst Extraction results: Input rows: 1,400,000 Output prompt rows: 1,400,000 JSON parse errors: 0… See the full description on the dataset page: https://huggingface.co/datasets/tim-gabie/am-deepseek-r1-distilled-prompts-1.4m.texttext-generation1M<n<10M0 likes19 downloads3mo agoHugging Face17Amdeous /NetworkTraffic Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Amdeous/NetworkTraffic.text100K<n<1M1 likes14 downloads2y agoHugging Face18RandomOscillations /autotrain-data-amdal-mining-llama2-7b-cleantabularn<1K0 likes12 downloads1y agoHugging Face19JinnP /amdpilot-lora-sft-dataset-v5_1 v5.1 Updated release with canonical split (89 train + 3 eval) and 3-view derivatives. n<1K0 likes11 downloads6mo agoHugging Face20Alphacode-AI /AM-Deepseek-Translate_Distiltext100K<n<1M0 likes10 downloads1y agoHugging Face21SiloAI /AMD-eAIrs-SFT-Demo-Data AMD enterprise AI reference stack Supervised Fine-tuning Demo Data This is a dataset for demonstration purposes. It is based on crawling the AMD enterprise AI reference stack documentation and creating prompt-answer pairs from it. Additional negative refusal examples were also generated. This can be used to fine-tune a Large Language Model with Supervised Fine-tuning. Intended model learning outcome The model fine-tuned on this data is intended to answer user… See the full description on the dataset page: https://huggingface.co/datasets/SiloAI/AMD-eAIrs-SFT-Demo-Data.text1K<n<10K0 likes8 downloads4mo agoHugging Face22andyfriedrich-amd /hipifyplustextn<1K0 likes6 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.