datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GLM-5.1-Reasoning-1M-Cleaned
GLM-5.1-Reasoning-1M-Cleaned
GLM-5.1-Reasoning-1M-Cleaned is a cleaned and reformatted derivative of Kassadin88/GLM-5.1-1000000x. It preserves the original four-subset layout (main, PHD-Science, Multilingual-STEM, Math) while converting every example into a unified SFT-ready schema with explicit conversations, input, output, domain, and meta fields.
This release was prepared from the original dataset published by Kassadin88.
Summary
Teacher model in the data: GLM-5.1… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/GLM-5.1-Reasoning-1M-Cleaned.Kimi-K2.5-Reasoning-1M-Cleaned
🪐 Kimi-K2.5-Reasoning-1M-Cleaned
Kimi-K2.5-Reasoning-1M-Cleaned is a cleaned derivative of ianncity/KIMI-K2.5-1000000x. It preserves the original four-config layout from the source dataset and rewrites each record into a unified reasoning-SFT schema with id, conversations, input, output, domain, and meta.
Summary
Source dataset: ianncity/KIMI-K2.5-1000000x
Source author: ianncity
Teacher model recorded in meta.teacher_model: KIMI-K2.5
Token lengths computed with… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Kimi-K2.5-Reasoning-1M-Cleaned.milu-cleaned
MILU: A Multi-task Indic Language Understanding Benchmark
Overview
MILU (Multi-task Indic Language Understanding Benchmark) is a comprehensive evaluation dataset designed to assess the performance of Large Language Models (LLMs) across 11 Indic languages. It spans 8 domains and 41 subjects, reflecting both general and culturally specific knowledge from India.
Key Features
11 Indian Languages: Bengali, Gujarati, Hindi, Kannada, Malayalam… See the full description on the dataset page: https://huggingface.co/datasets/murthyrudra/milu-cleaned.orca-agentinstruct-1M-v1-cleaned
🐋 Orca-AgentInstruct-1M-v1-cleaned
This is a cleaned version of the microsoft/orca-agentinstruct-1M-v1 dataset released by Microsoft.
orca-agentinstruct-1M-v1 is a fully synthetic dataset using only raw text publicly available on the web as seed data. It is a subset of the full AgentInstruct dataset (~25M samples) that created Orca-3-Mistral. Compared to Mistral 7B Instruct, the authors claim 40% improvement on AGIEval, 19% improvement on MMLU, 54% improvement on GSM8K, 38%… See the full description on the dataset page: https://huggingface.co/datasets/mlabonne/orca-agentinstruct-1M-v1-cleaned.tool-reasoning-sft-TOOLS-hermes_reasoning_tool_use-data-cleaned-rectified
Hermes Reasoning Tool Use — Cleaned & Rectified
👥 Follow the Author
Aman Priyanshu
Overview
This dataset is a cleaned and restructured version of interstellarninja/hermes_reasoning_tool_use. The original dataset uses the Hermes/NousResearch multi-turn format with from/value fields and embedded <think> + <tool_call> tags inside single gpt turns. This version converts it into a strict multi-turn conversation structure with validated role transitions.… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-TOOLS-hermes_reasoning_tool_use-data-cleaned-rectified.Kimi-K2.5-Reasoning-1M-Cleaned
🪐 Kimi-K2.5-Reasoning-1M-Cleaned
Kimi-K2.5-Reasoning-1M-Cleaned is a cleaned derivative of ianncity/KIMI-K2.5-1000000x. It preserves the original four-config layout from the source dataset and rewrites each record into a unified reasoning-SFT schema with id, conversations, input, output, domain, and meta.
Summary
Source dataset: ianncity/KIMI-K2.5-1000000x
Source author: ianncity
Teacher model recorded in meta.teacher_model: KIMI-K2.5
Token lengths computed with… See the full description on the dataset page: https://huggingface.co/datasets/rAVEUK/Kimi-K2.5-Reasoning-1M-Cleaned.GLM-5.1-Reasoning-1M-Cleaned
GLM-5.1-Reasoning-1M-Cleaned
GLM-5.1-Reasoning-1M-Cleaned is a cleaned and reformatted derivative of Kassadin88/GLM-5.1-1000000x. It preserves the original four-subset layout (main, PHD-Science, Multilingual-STEM, Math) while converting every example into a unified SFT-ready schema with explicit conversations, input, output, domain, and meta fields.
This release was prepared from the original dataset published by Kassadin88.
Summary
Teacher model in the data: GLM-5.1… See the full description on the dataset page: https://huggingface.co/datasets/EngMuhammadAtef/GLM-5.1-Reasoning-1M-Cleaned.alpaca-cleaned-pt
Data Description
This HF data repository contains the Portuguese Alpaca dataset used in our study of monolingual versus multilingual instruction tuning.
GitHub
Paper
Creation
Machine-translated from yahma/alpaca-cleaned into Portuguese.
Usage
This data is intended to be used for Portuguese instruction tuning.
The dataset has roughly 52K instances in the JSON format.
Each instance has an instruction, an output, and an optional input. An example is shown… See the full description on the dataset page: https://huggingface.co/datasets/pinzhenchen/alpaca-cleaned-pt.OpenR1-CleanedThis dataset is used in the paper Value-Guided Search for Efficient Chain-of-Thought Reasoning. It contains data for training and evaluating value models for improved long-context reasoning.
GitHub Repository: https://github.com/kaiwenw/value-guided-search
Related resources:
Dataset (OpenR1-Cleaned): https://huggingface.co/datasets/VGS-AI/OpenR1-Cleaned
Dataset (OpenR1-VM): https://huggingface.co/datasets/VGS-AI/OpenR1-VM
Value Model (DeepSeek-VM-1.5B):… See the full description on the dataset page: https://huggingface.co/datasets/VGS-AI/OpenR1-Cleaned.alpaca-cleaned-italian
Dataset Card for Alpaca-Cleaned-Italian
About the translation and the original data
The translation was done with X-ALMA, a 13-billion-parameter model that surpasses state-of-the-art open-source multilingual LLMs (as of Q1 2025, paper here).
The original alpaca-cleaned dataset is also kept here so that there is parallel data for Italian and English.
Additional notes on the translation
Despite the good quality of the translation, errors, though rare, are… See the full description on the dataset page: https://huggingface.co/datasets/DanielSc4/alpaca-cleaned-italian.alpaca-cleaned-52k-th
Summary
This is a Thai 🇹🇭-instructed dataset translated from cleaned version of the original Alpaca Dataset released by Stanford using Google Cloud Translation, contain 52,000 instructions and demonstrations generated by OpenAI's text-davinci-003 engine.
This instruction data can be used to conduct instruction-tuning for language models and make the language model follow instruction better.
The following issues have been identified in the original release and fixed in this… See the full description on the dataset page: https://huggingface.co/datasets/Thaweewat/alpaca-cleaned-52k-th.tool-reasoning-sft-TOOLS-toolace-sft-tool-use-agent-data-cleaned-rectified
ToolACE - Tool-Use Agent Data Cleaned & Rectified
👥 Follow the Author
Aman Priyanshu
Overview
This dataset is a cleaned and restructured version of the Team-ACE/ToolACE dataset. ToolACE is a high-quality conversational tool-use dataset containing 11,300+ examples of natural language interactions requiring function calling across diverse domains. This version converts the original OpenAI function-call format into a standardized multi-turn tool-use… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-TOOLS-toolace-sft-tool-use-agent-data-cleaned-rectified.10k_rows_cleaned_prompts
10K Rows Cleaned Prompts Dataset
Created by Aipresso LIMITED, London, UK
⚠️ IMPORTANT: By using this dataset, you agree to our Terms of Use
You must provide attribution when using this data in publications, research, or commercial products.
Dataset Overview
A chunked collection of 2.7 million cleaned English prompts, organized into 200 files of 10,000 rows each for easy processing and distributed training of language models.
📊 Dataset Statistics
Metric… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/10k_rows_cleaned_prompts.glm-5.1-reasoning-1m-cleaned
GLM-5.1-Reasoning-1M-Cleaned
GLM-5.1-Reasoning-1M-Cleaned is a cleaned and reformatted derivative of Kassadin88/GLM-5.1-1000000x. It preserves the original four-subset layout (main, PHD-Science, Multilingual-STEM, Math) while converting every example into a unified SFT-ready schema with explicit conversations, input, output, domain, and meta fields.
This release was prepared from the original dataset published by Kassadin88.
Summary
Teacher model in the data:… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/glm-5.1-reasoning-1m-cleaned.Kimi-K2.5-Reasoning-1M-Cleaned
🪐 Kimi-K2.5-Reasoning-1M-Cleaned
Kimi-K2.5-Reasoning-1M-Cleaned is a cleaned derivative of ianncity/KIMI-K2.5-1000000x. It preserves the original four-config layout from the source dataset and rewrites each record into a unified reasoning-SFT schema with id, conversations, input, output, domain, and meta.
Summary
Source dataset: ianncity/KIMI-K2.5-1000000x
Source author: ianncity
Teacher model recorded in meta.teacher_model: KIMI-K2.5
Token lengths computed… See the full description on the dataset page: https://huggingface.co/datasets/JBrightmanAI/Kimi-K2.5-Reasoning-1M-Cleaned.ZamAI-Pashto-Dataset-Cleaned
ZamAI Pashto Dataset Cleaned
Languages: psLicense: apache-2.0Task categories: text-classification, text-generation, question-answeringSize categories: 10K<n<100K
Summary
This dataset is part of the ZamAI Pashto data collection. It is intended for text-classification, text-generation, question-answering tasks in Pashto.
How to use
from datasets import load_dataset
dataset = load_dataset("tasal9/ZamAI-Pashto-Dataset-Cleaned")
print(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/ZamAI-Pashto-Dataset-Cleaned.tool-reasoning-sft-RESEARCH-grill-lab-browsecomp-plus-runs-data-cleaned-rectified
Tool-Reasoning SFT — BrowseComp-Plus Runs (Cleaned & Rectified)
Multi-turn tool-use reasoning trajectories derived from grill-lab/browsecomp-plus-runs, converted to a structured SFT format following the interstellarninja/hermes_reasoning_tool_use convention.
Source
Based on the execution trajectories from "Revisiting Text Ranking in Deep Research" (arXiv:2602.21456):
Original data: grill-lab/browsecomp-plus-runs (MIT)
Format
Each row contains a messages… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-RESEARCH-grill-lab-browsecomp-plus-runs-data-cleaned-rectified.GLM-5.1-Reasoning-1M-Cleaned
GLM-5.1-Reasoning-1M-Cleaned
GLM-5.1-Reasoning-1M-Cleaned is a cleaned and reformatted derivative of Kassadin88/GLM-5.1-1000000x. It preserves the original four-subset layout (main, PHD-Science, Multilingual-STEM, Math) while converting every example into a unified SFT-ready schema with explicit conversations, input, output, domain, and meta fields.
This release was prepared from the original dataset published by Kassadin88.
Summary
Teacher model in the data: GLM-5.1… See the full description on the dataset page: https://huggingface.co/datasets/zhangbo2008/GLM-5.1-Reasoning-1M-Cleaned.alpaca-cleaned-es
Data Description
This HF data repository contains the Spanish Alpaca dataset used in our study of monolingual versus multilingual instruction tuning.
GitHub
Paper
Creation
Machine-translated from yahma/alpaca-cleaned into Spanish.
Usage
This data is intended to be used for Spanish instruction tuning.
The dataset has roughly 52K instances in the JSON format.
Each instance has an instruction, an output, and an optional input. An example is shown below:
{… See the full description on the dataset page: https://huggingface.co/datasets/pinzhenchen/alpaca-cleaned-es.GLM-5.1-Reasoning-1M-Cleaned
GLM-5.1-Reasoning-1M-Cleaned
GLM-5.1-Reasoning-1M-Cleaned is a cleaned and reformatted derivative of Kassadin88/GLM-5.1-1000000x. It preserves the original four-subset layout (main, PHD-Science, Multilingual-STEM, Math) while converting every example into a unified SFT-ready schema with explicit conversations, input, output, domain, and meta fields.
This release was prepared from the original dataset published by Kassadin88.
Summary
Teacher model in the data: GLM-5.1… See the full description on the dataset page: https://huggingface.co/datasets/rlandismd/GLM-5.1-Reasoning-1M-Cleaned.alpaca-cleaned-dutch
Dataset Card for Alpaca Cleaned Dutch
Dataset Summary
This dataset contains 51,712 conversations between een AI assistant and a (fake) "Human" (generated) in Dutch. They are translations of Alpaca Cleaned Dataset.
☕ Want to help me out? Translating the data with the OpenAI API, and prompt testing, cost me 💸$57.99💸. If you like this dataset, please consider buying me a coffee to offset a portion of this cost, I appreciate it a lot! ☕
If you use this dataset or refer to… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/alpaca-cleaned-dutch.alpaca-cleaned-cs
Data Description
This HF data repository contains the Czech Alpaca dataset used in our study of monolingual versus multilingual instruction tuning.
GitHub
Paper
Creation
Machine-translated from yahma/alpaca-cleaned into Czech.
Usage
This data is intended to be used for Czech instruction tuning.
The dataset has roughly 52K instances in the JSON format.
Each instance has an instruction, an output, and an optional input. An example is shown below:
{… See the full description on the dataset page: https://huggingface.co/datasets/pinzhenchen/alpaca-cleaned-cs.persian-gk-cleanedThis is a cleaned and validated version of the original mshojaei77/persian-gk dataset.
The purpose of this version is to ensure robust compatibility with modern fine-tuning workflows that rely on strict chat templates (e.g., tokenizer.apply_chat_template). The cleaning process resolves structural errors in the original dataset that could cause TemplateError or other silent failures during training with models like Gemma 3N, Llama 3, and others.
Cleaning and Validation Process… See the full description on the dataset page: https://huggingface.co/datasets/mshojaei77/persian-gk-cleaned.alpaca-cleaned-fr
Data Description
This HF data repository contains the French Alpaca dataset used in our study of monolingual versus multilingual instruction tuning.
GitHub
Paper
Creation
Machine-translated from yahma/alpaca-cleaned into French.
Usage
This data is intended to be used for French instruction tuning.
The dataset has roughly 52K instances in the JSON format.
Each instance has an instruction, an output, and an optional input. An example is shown below:
{… See the full description on the dataset page: https://huggingface.co/datasets/pinzhenchen/alpaca-cleaned-fr.GLM-5.1-Reasoning-1M-Cleaned
GLM-5.1-Reasoning-1M-Cleaned
GLM-5.1-Reasoning-1M-Cleaned is a cleaned and reformatted derivative of Kassadin88/GLM-5.1-1000000x. It preserves the original four-subset layout (main, PHD-Science, Multilingual-STEM, Math) while converting every example into a unified SFT-ready schema with explicit conversations, input, output, domain, and meta fields.
This release was prepared from the original dataset published by Kassadin88.
Summary
Teacher model in the data: GLM-5.1… See the full description on the dataset page: https://huggingface.co/datasets/txchmechanicus/GLM-5.1-Reasoning-1M-Cleaned.SlimOrca-Dedup-Uzbek-cleanedThis is an Uzbek translated and cleaned version of https://huggingface.co/datasets/Open-Orca/SlimOrca-Dedup.
Specifically, these replaced/removed records that had 'Uzbek translation|Uzbekcha tarjima|Uzbek tarjima|impossible to translate|not possible to translate|cannot fulfill your request|text is in|tilida yozilgan|Uzbek|o'zbek|ozbek|I am sorry'.
You can use this dataset for chat fine-tuning of LLMs.
This dataset has around 100M tokens (500M*0.8/4 = 100M assuming 4 chars are one token).… See the full description on the dataset page: https://huggingface.co/datasets/MLDataScientist/SlimOrca-Dedup-Uzbek-cleaned.GLM-5.1-Reasoning-1M-Cleaned
GLM-5.1-Reasoning-1M-Cleaned
GLM-5.1-Reasoning-1M-Cleaned is a cleaned and reformatted derivative of Kassadin88/GLM-5.1-1000000x. It preserves the original four-subset layout (main, PHD-Science, Multilingual-STEM, Math) while converting every example into a unified SFT-ready schema with explicit conversations, input, output, domain, and meta fields.
This release was prepared from the original dataset published by Kassadin88.
Summary
Teacher model in the data: GLM-5.1… See the full description on the dataset page: https://huggingface.co/datasets/Bas95/GLM-5.1-Reasoning-1M-Cleaned.GLM-5.1-Reasoning-1M-Cleaned
GLM-5.1-Reasoning-1M-Cleaned
GLM-5.1-Reasoning-1M-Cleaned is a cleaned and reformatted derivative of Kassadin88/GLM-5.1-1000000x. It preserves the original four-subset layout (main, PHD-Science, Multilingual-STEM, Math) while converting every example into a unified SFT-ready schema with explicit conversations, input, output, domain, and meta fields.
This release was prepared from the original dataset published by Kassadin88.
Summary
Teacher model in the data:… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/GLM-5.1-Reasoning-1M-Cleaned.alpaca-cleaned-ru
Data Description
This HF data repository contains the Russian Alpaca dataset used in our study of monolingual versus multilingual instruction tuning.
GitHub
Paper
Creation
Machine-translated from yahma/alpaca-cleaned into Russian.
Usage
This data is intended to be used for Russian instruction tuning.
The dataset has roughly 52K instances in the JSON format.
Each instance has an instruction, an output, and an optional input. An example is shown below:
{… See the full description on the dataset page: https://huggingface.co/datasets/pinzhenchen/alpaca-cleaned-ru.korean-quality-cleaned
Korean Quality Dataset (Cleaned)
고품질 한국어 Instruction 데이터셋 (정제 버전)
English
Dataset Description
This is a cleaned and standardized Korean instruction dataset, combining multiple high-quality open-source Korean datasets with unified formatting and quality filtering.
Key Features
✅ Unified Format: Standardized messages format (OpenAI-compatible)
✅ Quality Filtering: Length, special characters, repetition filtering
✅ Clean Structure: Removed redundant… See the full description on the dataset page: https://huggingface.co/datasets/MyeongHo0621/korean-quality-cleaned.
