CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01amd /Instella-Long Instella-Long The Instella-Long dataset is a collection of pre-training and instruction following data that is used to train Instella-3B-Long-Instruct. The pre-training data is sourced from Prolong. For the SFT data, we use public datasets: Ultrachat 200K, OpenMathinstruct-2, Tülu-3 Instruction Following, and MMLU auxiliary train set. In addition, we generate synthetic long instruction data using documents of the books and arxiv from our pre-training corpus and the dclm subset from… See the full description on the dataset page: https://huggingface.co/datasets/amd/Instella-Long.0 likes14k downloads10mo agoHugging Face02zmodelerlover /amd-nr0 likes4.7k downloads2d agoHugging Face03optimum-amd /transformers_daily_ci1 likes2.9k downloads18h agoHugging Face04optimum-amd /transformers_pr_ci0 likes2.6k downloads7h agoHugging Face05a-m-team /AM-DeepSeek-R1-Distilled-1.4MFor more open-source datasets, models, and methodologies, please visit our GitHub repository. AM-DeepSeek-R1-Distilled-1.4M is a large-scale general reasoning task dataset composed of high-quality and challenging reasoning problems. These problems are collected from numerous open-source datasets, semantically deduplicated, and cleaned to eliminate test set contamination. All responses in the dataset are distilled from the reasoning model (mostly DeepSeek-R1) and have undergone rigorous… See the full description on the dataset page: https://huggingface.co/datasets/a-m-team/AM-DeepSeek-R1-Distilled-1.4M.text-generation1M<n<10M184 likes2.2k downloads1y agoHugging Face06a-m-team /AM-DeepSeek-Distilled-40MFor more open-source datasets, models, and methodologies, please visit our GitHub repository and paper: DeepDistill: Enhancing LLM Reasoning Capabilities via Large-Scale Difficulty-Graded Data Training. Due to certain constraints, we are only able to open-source a subset of the complete dataset. Model Training Performance based on our complete dataset On AIME 2024, our 72B model achieved a score of 79.2 using only supervised fine-tuning (SFT). The 32B model reached 75.8 and… See the full description on the dataset page: https://huggingface.co/datasets/a-m-team/AM-DeepSeek-Distilled-40M.tabulartext-generation10M<n<100M56 likes2.1k downloads1y agoHugging Face07a-m-team /AM-DeepSeek-R1-0528-Distilled 📘 Dataset Summary This dataset is a high-quality reasoning corpus distilled from DeepSeek-R1-0528, an improved version of the DeepSeek-R1 large language model. Compared to its initial release, DeepSeek-R1-0528 demonstrates significant advances in reasoning, instruction following, and multi-turn dialogue. Motivated by these improvements, we collected and distilled a diverse set of 2.6 million queries across multiple domains, using DeepSeek-R1-0528 as the teacher. A notable… See the full description on the dataset page: https://huggingface.co/datasets/a-m-team/AM-DeepSeek-R1-0528-Distilled.text-generation1M<n<10M102 likes1k downloads1y agoHugging Face08amd /ReasonLite-Dataset GitHub | Dataset | Blog ReasonLite is an ultra-lightweight math reasoning model. With only 0.6B parameters, it leverages high-quality data distillation to achieve performance comparable to models over 10× its size, such as Qwen3-8B, reaching 75.2 on AIME24 and extending the scaling law of small models. 🔥 Best-performing 0.6B math reasoning model 🔓 Fully open-source — weights, scripts, datasets, synthesis pipeline⚙️ Distilled in two stages to balance efficiency and high… See the full description on the dataset page: https://huggingface.co/datasets/amd/ReasonLite-Dataset.text1M<n<10M16 likes631 downloads8mo agoHugging Face09koajoel /AM-DeepSeek-R1-Distilled-1.4Mtext1M<n<10M0 likes435 downloads1y agoHugging Face10amd-nicknick /bert-base-uncased-2022_tokenized_dataset10M<n<100M0 likes344 downloads3y agoHugging Face11amd /Cot-Drop LuminaSFT LuminaSFT is a synthetic SFT dataset suite specifically designed to improve both general-purpose and task-specific SLMs. LuminaSFT consists of multiple curated splits that target diverse capabilities: UltraChat200K-DeepSeek - A regenerated base SFT dataset for broad instruction following. InstructGPT-NaturalQA and InstructGPT-TriviaQA - Factual question answering datasets to strengthen knowledge recall and answer accuracy. CoT-Drop - A reading comprehension dataset with… See the full description on the dataset page: https://huggingface.co/datasets/amd/Cot-Drop.texttext-generation10K<n<100K2 likes332 downloads7mo agoHugging Face12amd /AIG-Datasets AMD AIG GPU Kernel Datasets AMD AIG-Datasets is a collection of GPU-kernel generation, translation, optimization, and ROCm-library supervision data. It contains PyTorch/CUDA-to-HIP, HIP-to-HIP, PyTorch-to-Triton, and production-grounded rocBLAS/rocSOLVER entries, together with metadata, samples, conversion utilities, and reproducible evaluation tools. The repository is organized into versioned releases. New training and evaluation workflows should use the unified-schema datasets… See the full description on the dataset page: https://huggingface.co/datasets/amd/AIG-Datasets.text-generation4 likes272 downloads28d agoHugging Face13amd /SAND-MATH SAND-Math: A Synthetic Dataset of Difficult Problems to Elevate LLM Math Performance 📃 Paper | 🤗 Dataset SAND-Math (Synthetic Augmented Novel and Difficult Mathematics) is a high-quality, high-difficulty dataset of mathematics problems and solutions. It is generated using a novel pipeline that addresses the critical bottleneck of scarce, high-difficulty training data for mathematical Large Language Models (LLMs). Key Features Novel Problem Generation: Problems are… See the full description on the dataset page: https://huggingface.co/datasets/amd/SAND-MATH.tabularquestion-answering10K<n<100K3 likes260 downloads11mo agoHugging Face14abcdefj123 /AM-DeepSeek-R1-Distilled-1.4MFor more open-source datasets, models, and methodologies, please visit our GitHub repository. AM-DeepSeek-R1-Distilled-1.4M is a large-scale general reasoning task dataset composed of high-quality and challenging reasoning problems. These problems are collected from numerous open-source datasets, semantically deduplicated, and cleaned to eliminate test set contamination. All responses in the dataset are distilled from the reasoning model (mostly DeepSeek-R1) and have undergone rigorous… See the full description on the dataset page: https://huggingface.co/datasets/abcdefj123/AM-DeepSeek-R1-Distilled-1.4M.text-generation1M<n<10M2 likes205 downloads2mo agoHugging Face15amd1234567 /AudioJailbreak Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models AudioJailbreak is a benchmark framework specifically designed for evaluating the security of Audio Language Models (Audio LLMs). This project tests model defenses against malicious requests through various audio perturbation techniques.Note: This project aims to improve the security of audio language models. Researchers should use this tool responsibly. 📋 Table of… See the full description on the dataset page: https://huggingface.co/datasets/amd1234567/AudioJailbreak.0 likes201 downloads3mo agoHugging Face16giacomoran /hackathon_amd_mission2_black_sortThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so101_follower", "total_episodes": 155, "total_frames": 48397, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 30, "splits": { "train": "0:155" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/giacomoran/hackathon_amd_mission2_black_sort.tabularrobotics10K<n<100K0 likes177 downloads9mo agoHugging Face17MaziyarPanahi /AM-DeepSeek-R1-0528-Distilled-with-Systemtext1M<n<10M4 likes170 downloads1y agoHugging Face18amd /Instella-GSM8K-synthetic Instella-GSM8K-synthetic The Instella-GSM8K-synthetic dataset was used in the second stage pre-training of Instella-3B model, which was trained on top of the Instella-3B-Stage1 model. This synthetic dataset was generated using the training set of GSM8k dataset, where we first used Qwen2.5-72B-Instruct to Abstract numerical values as function parameters and generate a Python program to solve the math question. Identify and replace numerical values in the existing question with… See the full description on the dataset page: https://huggingface.co/datasets/amd/Instella-GSM8K-synthetic.textquestion-answering1M<n<10M7 likes169 downloads10mo agoHugging Face19chhao /AM-DeepSeek-R1-Filtered-Math-Code AM DeepSeek R1 Filtered Math and Code This repository publishes reproducible training subsets derived from a-m-team/AM-DeepSeek-R1-Distilled-1.4M at pinned revision 53531c06634904118a2dcd83961918c4d69d1cdf. The source dataset and this derived release use CC BY-NC 4.0. Commercial use is not permitted by that license. Preserve attribution and review the upstream dataset card before use. Contents Config / split Records Bytes SHA-256 math / train 111,657 2… See the full description on the dataset page: https://huggingface.co/datasets/chhao/AM-DeepSeek-R1-Filtered-Math-Code.texttext-generation100K<n<1M0 likes169 downloads2mo agoHugging Face20alexei-v-ivanov-amd /flores_plustext100K<n<1M0 likes164 downloads2y agoHugging Face21koajoel /AM-DeepSeek-R1-Distilled-1.4M-Englishtext1M<n<10M0 likes157 downloads1y agoHugging Face22VanyaJ /mories-caeba-amd64 🚀 Mories-CAEBA & Zenith Production Release Guide (Linux AMD64) 완전 에어갭(Air-Gapped) x86_64(AMD64) Linux 서버용 프로덕션 배포 패키지Mories 인지 지식 그래프, CAEBA 오케스트레이터, Zenith Vue 3 대시보드, NATS JetStream, Keycloak IAM, Redis 및 스탠드얼론 MCP 서버 일체 포함 📦 1. 배포 패키지 구성 요소 (7대 도커 이미지 & 설정) 파일 / 디렉토리 설명 세부 정보 docker-compose.yml 7대 마이크로서비스 오케스트레이션 구성 파일 pull_policy: never, 포트 충돌 방지, AUTH_DISABLED 지원 .env.example 프로덕션 환경변수 템플릿 포트, DB 인증 정보, LLM/Embedding 설정 docker_images/ 순수… See the full description on the dataset page: https://huggingface.co/datasets/VanyaJ/mories-caeba-amd64.text0 likes145 downloads23d agoHugging Face23AbijahKaj /telephony-amd-dataset Telephony AMD (Answering Machine Detection) Dataset Overview A multilingual 4-class telephony audio classification dataset for training streaming Answering Machine Detection models. Contains real human speech (PolyAI/MINDS14) mixed with TTS-generated audio (Microsoft Neural TTS / edge-tts) across English, French, Spanish, and German. Key design principle: Voicemail greetings are recorded by real humans and sound acoustically identical to live speech. This dataset includes… See the full description on the dataset page: https://huggingface.co/datasets/AbijahKaj/telephony-amd-dataset.audio1K<n<10K2 likes142 downloads5mo agoHugging Face24amd /SAND-Post-Training-Dataset SAND-Post-Training-Dataset: High-Quality Synthetic Reasoning Dataset Built with AMD GPUs Dataset Summary We introduce the SAND-Post-Training-Dataset, a high-quality synthetic reasoning dataset for mathematics and science built entirely using a synthetic data pipeline running on the AMD ROCm™ stack and AMD Instinct™ MI325 GPUs. This dataset prioritizes difficulty and novelty over volume, demonstrating that high-difficulty synthetic data can elevate… See the full description on the dataset page: https://huggingface.co/datasets/amd/SAND-Post-Training-Dataset.tabularquestion-answering10K<n<100K3 likes140 downloads10mo agoHugging Face25Yxanul /AMD-SFT-Mix_3.5M AMD-SFT-Mix_3.5M 3,556,428 SFT conversations — five AMD instruction-tuning datasets merged into a single pre-shuffled stream, with per-row provenance so any example can be traced back to its source. Nothing was regenerated: this is a normalisation, provenance and shuffling pass over existing public datasets. All credit for the data belongs to AMD. Composition source rows share origin naturalqa 1,304,792 36.69% amd/InstructGpt-NaturalQa triviaqa 1,118… See the full description on the dataset page: https://huggingface.co/datasets/Yxanul/AMD-SFT-Mix_3.5M.texttext-generation1M<n<10M0 likes131 downloads2mo agoHugging Face26syzym /xbmu_amdo31 Dataset Card for [XBMU-AMDO31] Dataset Summary XBMU-AMDO31 dataset is a speech recognition corpus of Amdo Tibetan dialect. The open source corpus contains 31 hours of speech data and resources related to build speech recognition systems, including transcribed texts and a Tibetan pronunciation dictionary. Supported Tasks and Leaderboards automatic-speech-recognition: The dataset can be used to train a model for Amdo Tibetan Automatic Speech Recognition (ASR). It… See the full description on the dataset page: https://huggingface.co/datasets/syzym/xbmu_amdo31.textautomatic-speech-recognition10K<n<100K4 likes128 downloads4y agoHugging Face27ahmedheakl /asm_cuda_to_amdtabular10K<n<100K1 likes126 downloads2y agoHugging Face28mlfoundations-dev /AM-DeepSeek-R1-Distilled-1.4M-am_0.5Mtext100K<n<1M2 likes123 downloads1y agoHugging Face291g0rrr /amd-test53This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "eyou_ft7_follower", "total_episodes": 3, "total_frames": 983, "total_tasks": 1, "total_videos": 9, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:3" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/1g0rrr/amd-test53.tabularrobotics1K<n<10K0 likes116 downloads11mo agoHugging Face30amd /TTT-Bench TTT-Bench: A Benchmark for Evaluating Reasoning Ability with Simple and Novel Tic-Tac-Toe-style Games 📃 Paper | 🤗 Dataset | 🌐 Website We introduce TTT-Bench, a new benchmark specifically created to evaluate the reasoning capability of LRMs through a suite of simple and novel two-player Tic-Tac-Toe-style games. Although trivial for humans, these games require basic strategic reasoning, including predicting an opponent's intentions and understanding spatial configurations.… See the full description on the dataset page: https://huggingface.co/datasets/amd/TTT-Bench.textquestion-answeringn<1K1 likes111 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.