datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AM-DeepSeek-R1-Distilled-1.4MFor more open-source datasets, models, and methodologies, please visit our GitHub repository.
AM-DeepSeek-R1-Distilled-1.4M is a large-scale general reasoning task dataset composed of
high-quality and challenging reasoning problems. These problems are collected from numerous
open-source datasets, semantically deduplicated, and cleaned to eliminate test set contamination.
All responses in the dataset are distilled from the reasoning model (mostly DeepSeek-R1) and have undergone
rigorous… See the full description on the dataset page: https://huggingface.co/datasets/a-m-team/AM-DeepSeek-R1-Distilled-1.4M.AM-DeepSeek-Distilled-40MFor more open-source datasets, models, and methodologies, please visit our GitHub repository and paper: DeepDistill: Enhancing LLM Reasoning Capabilities via Large-Scale Difficulty-Graded Data Training.
Due to certain constraints, we are only able to open-source a subset of the complete dataset.
Model Training Performance based on our complete dataset
On AIME 2024, our 72B model achieved a score of 79.2 using only supervised fine-tuning (SFT). The 32B model reached 75.8 and… See the full description on the dataset page: https://huggingface.co/datasets/a-m-team/AM-DeepSeek-Distilled-40M.AM-Thinking-v1-Distilled
📘 Dataset Summary
AM-Thinking-v1 and Qwen3-235B-A22B are two reasoning datasets distilled from state-of-the-art teacher models. Each dataset contains high-quality, automatically verified responses generated from a shared set of 1.89 million queries spanning a wide range of reasoning domains.
The datasets share the same format and verification pipeline, allowing for direct comparison and seamless integration into downstream tasks. They are intended to support the development of… See the full description on the dataset page: https://huggingface.co/datasets/a-m-team/AM-Thinking-v1-Distilled.AM-DeepSeek-R1-0528-Distilled
📘 Dataset Summary
This dataset is a high-quality reasoning corpus distilled from DeepSeek-R1-0528, an improved version of the DeepSeek-R1 large language model. Compared to its initial release, DeepSeek-R1-0528 demonstrates significant advances in reasoning, instruction following, and multi-turn dialogue. Motivated by these improvements, we collected and distilled a diverse set of 2.6 million queries across multiple domains, using DeepSeek-R1-0528 as the teacher.
A notable… See the full description on the dataset page: https://huggingface.co/datasets/a-m-team/AM-DeepSeek-R1-0528-Distilled.AM-Qwen3-Distilled
📘 Dataset Summary
AM-Thinking-v1 and Qwen3-235B-A22B are two reasoning datasets distilled from state-of-the-art teacher models. Each dataset contains high-quality, automatically verified responses generated from a shared set of 1.89 million queries spanning a wide range of reasoning domains.
The datasets share the same format and verification pipeline, allowing for direct comparison and seamless integration into downstream tasks. They are intended to support the development of… See the full description on the dataset page: https://huggingface.co/datasets/a-m-team/AM-Qwen3-Distilled.AM-Math-Difficulty-RLFor more open-source datasets, models, and methodologies, please visit our GitHub repository.
We believe that the selection of training data for reinforcement learning is crucial.
To validate this, we conducted several experiments exploring how data difficulty influences training performance.
Our data sources originate from numerous excellent open-source projects, and we sincerely appreciate their contributions, without which our current achievements would not have been possible.… See the full description on the dataset page: https://huggingface.co/datasets/a-m-team/AM-Math-Difficulty-RL.amthal-hassaniya
🇲🇷 الأمثال الحسانية — Dataset جاهز لـ LoRA Fine-tuning
مجموعة بيانات تضم 319 مثلاً حسانياً بصيغة Alpaca القياسية، مستخرجة من كتاب موسوعة الأمثال الحسانية لبكار ولد احمدو.
الصيغة
صيغة Alpaca — الأكثر توافقاً مع مكتبات LoRA مثل trl, unsloth, axolotl:
{
"instruction": "أنت خبير في التراث الحساني الموريتاني. اشرح المثل الحساني التالي وبيّن معناه وفي أي سياق يُستخدم.",
"input": "ألْبَلْ تبرك على أكبارها",
"output": "يضرب لأهمية الكبار في مجتمعهم وحتمية التبعية لهم"
}… See the full description on the dataset page: https://huggingface.co/datasets/ahmed02mk/amthal-hassaniya.amtp-chunks-v1
Amtp Chunks V1
Amtp Chunks V1 is a high-quality document dataset generated by DocParserEngine.
Dataset Summary
Documents Processed: 1
Total Records: 794
Schema Format: chunks
Extraction Features: Structural detection, image extraction, AI-powered captioning, and OCR.
Supported Tasks
OCR & Text Extraction: High-accuracy text extraction from complex document layouts.
Image Captioning & Categorization: Vision-based descriptions and classification of extracted… See the full description on the dataset page: https://huggingface.co/datasets/Remixonwin/amtp-chunks-v1.
