CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01a-m-team /AM-DeepSeek-R1-Distilled-1.4MFor more open-source datasets, models, and methodologies, please visit our GitHub repository. AM-DeepSeek-R1-Distilled-1.4M is a large-scale general reasoning task dataset composed of high-quality and challenging reasoning problems. These problems are collected from numerous open-source datasets, semantically deduplicated, and cleaned to eliminate test set contamination. All responses in the dataset are distilled from the reasoning model (mostly DeepSeek-R1) and have undergone rigorous… See the full description on the dataset page: https://huggingface.co/datasets/a-m-team/AM-DeepSeek-R1-Distilled-1.4M.text-generation1M<n<10M184 likes2.2k downloads1y agoHugging Face02a-m-team /AM-DeepSeek-Distilled-40MFor more open-source datasets, models, and methodologies, please visit our GitHub repository and paper: DeepDistill: Enhancing LLM Reasoning Capabilities via Large-Scale Difficulty-Graded Data Training. Due to certain constraints, we are only able to open-source a subset of the complete dataset. Model Training Performance based on our complete dataset On AIME 2024, our 72B model achieved a score of 79.2 using only supervised fine-tuning (SFT). The 32B model reached 75.8 and… See the full description on the dataset page: https://huggingface.co/datasets/a-m-team/AM-DeepSeek-Distilled-40M.tabulartext-generation10M<n<100M56 likes2.1k downloads1y agoHugging Face03a-m-team /AM-Thinking-v1-Distilled 📘 Dataset Summary AM-Thinking-v1 and Qwen3-235B-A22B are two reasoning datasets distilled from state-of-the-art teacher models. Each dataset contains high-quality, automatically verified responses generated from a shared set of 1.89 million queries spanning a wide range of reasoning domains. The datasets share the same format and verification pipeline, allowing for direct comparison and seamless integration into downstream tasks. They are intended to support the development of… See the full description on the dataset page: https://huggingface.co/datasets/a-m-team/AM-Thinking-v1-Distilled.text-generation1M<n<10M64 likes1.1k downloads1y agoHugging Face04a-m-team /AM-DeepSeek-R1-0528-Distilled 📘 Dataset Summary This dataset is a high-quality reasoning corpus distilled from DeepSeek-R1-0528, an improved version of the DeepSeek-R1 large language model. Compared to its initial release, DeepSeek-R1-0528 demonstrates significant advances in reasoning, instruction following, and multi-turn dialogue. Motivated by these improvements, we collected and distilled a diverse set of 2.6 million queries across multiple domains, using DeepSeek-R1-0528 as the teacher. A notable… See the full description on the dataset page: https://huggingface.co/datasets/a-m-team/AM-DeepSeek-R1-0528-Distilled.text-generation1M<n<10M102 likes1k downloads1y agoHugging Face05a-m-team /AM-Qwen3-Distilled 📘 Dataset Summary AM-Thinking-v1 and Qwen3-235B-A22B are two reasoning datasets distilled from state-of-the-art teacher models. Each dataset contains high-quality, automatically verified responses generated from a shared set of 1.89 million queries spanning a wide range of reasoning domains. The datasets share the same format and verification pipeline, allowing for direct comparison and seamless integration into downstream tasks. They are intended to support the development of… See the full description on the dataset page: https://huggingface.co/datasets/a-m-team/AM-Qwen3-Distilled.text-generation1M<n<10M25 likes494 downloads1y agoHugging Face06a-m-team /AM-Math-Difficulty-RLFor more open-source datasets, models, and methodologies, please visit our GitHub repository. We believe that the selection of training data for reinforcement learning is crucial. To validate this, we conducted several experiments exploring how data difficulty influences training performance. Our data sources originate from numerous excellent open-source projects, and we sincerely appreciate their contributions, without which our current achievements would not have been possible.… See the full description on the dataset page: https://huggingface.co/datasets/a-m-team/AM-Math-Difficulty-RL.texttext-generation100K<n<1M16 likes171 downloads1y agoHugging Face07ahmed02mk /amthal-hassaniya 🇲🇷 الأمثال الحسانية — Dataset جاهز لـ LoRA Fine-tuning مجموعة بيانات تضم 319 مثلاً حسانياً بصيغة Alpaca القياسية، مستخرجة من كتاب موسوعة الأمثال الحسانية لبكار ولد احمدو. الصيغة صيغة Alpaca — الأكثر توافقاً مع مكتبات LoRA مثل trl, unsloth, axolotl: { "instruction": "أنت خبير في التراث الحساني الموريتاني. اشرح المثل الحساني التالي وبيّن معناه وفي أي سياق يُستخدم.", "input": "ألْبَلْ تبرك على أكبارها", "output": "يضرب لأهمية الكبار في مجتمعهم وحتمية التبعية لهم" }… See the full description on the dataset page: https://huggingface.co/datasets/ahmed02mk/amthal-hassaniya.texttext-generationn<1K1 likes31 downloads6mo agoHugging Face08Remixonwin /amtp-chunks-v1 Amtp Chunks V1 Amtp Chunks V1 is a high-quality document dataset generated by DocParserEngine. Dataset Summary Documents Processed: 1 Total Records: 794 Schema Format: chunks Extraction Features: Structural detection, image extraction, AI-powered captioning, and OCR. Supported Tasks OCR & Text Extraction: High-accuracy text extraction from complex document layouts. Image Captioning & Categorization: Vision-based descriptions and classification of extracted… See the full description on the dataset page: https://huggingface.co/datasets/Remixonwin/amtp-chunks-v1.tabulartext-generationn<1K0 likes13 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.