CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Jackrong /GLM-5.1-Reasoning-1M-Cleaned GLM-5.1-Reasoning-1M-Cleaned GLM-5.1-Reasoning-1M-Cleaned is a cleaned and reformatted derivative of Kassadin88/GLM-5.1-1000000x. It preserves the original four-subset layout (main, PHD-Science, Multilingual-STEM, Math) while converting every example into a unified SFT-ready schema with explicit conversations, input, output, domain, and meta fields. This release was prepared from the original dataset published by Kassadin88. Summary Teacher model in the data: GLM-5.1… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/GLM-5.1-Reasoning-1M-Cleaned.texttext-generation100K<n<1M297 likes1k downloads5mo agoHugging Face02Jackrong /Kimi-K2.5-Reasoning-1M-Cleaned 🪐 Kimi-K2.5-Reasoning-1M-Cleaned Kimi-K2.5-Reasoning-1M-Cleaned is a cleaned derivative of ianncity/KIMI-K2.5-1000000x. It preserves the original four-config layout from the source dataset and rewrites each record into a unified reasoning-SFT schema with id, conversations, input, output, domain, and meta. Summary Source dataset: ianncity/KIMI-K2.5-1000000x Source author: ianncity Teacher model recorded in meta.teacher_model: KIMI-K2.5 Token lengths computed with… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Kimi-K2.5-Reasoning-1M-Cleaned.texttext-generation100K<n<1M36 likes604 downloads5mo agoHugging Face03murthyrudra /milu-cleaned MILU: A Multi-task Indic Language Understanding Benchmark Overview MILU (Multi-task Indic Language Understanding Benchmark) is a comprehensive evaluation dataset designed to assess the performance of Large Language Models (LLMs) across 11 Indic languages. It spans 8 domains and 41 subjects, reflecting both general and culturally specific knowledge from India. Key Features 11 Indian Languages: Bengali, Gujarati, Hindi, Kannada, Malayalam… See the full description on the dataset page: https://huggingface.co/datasets/murthyrudra/milu-cleaned.textmultiple-choice10K<n<100K1 likes569 downloads1y agoHugging Face04mlabonne /orca-agentinstruct-1M-v1-cleaned 🐋 Orca-AgentInstruct-1M-v1-cleaned This is a cleaned version of the microsoft/orca-agentinstruct-1M-v1 dataset released by Microsoft. orca-agentinstruct-1M-v1 is a fully synthetic dataset using only raw text publicly available on the web as seed data. It is a subset of the full AgentInstruct dataset (~25M samples) that created Orca-3-Mistral. Compared to Mistral 7B Instruct, the authors claim 40% improvement on AGIEval, 19% improvement on MMLU, 54% improvement on GSM8K, 38%… See the full description on the dataset page: https://huggingface.co/datasets/mlabonne/orca-agentinstruct-1M-v1-cleaned.textquestion-answering1M<n<10M69 likes254 downloads2y agoHugging Face05AmanPriyanshu /tool-reasoning-sft-TOOLS-hermes_reasoning_tool_use-data-cleaned-rectified Hermes Reasoning Tool Use — Cleaned & Rectified 👥 Follow the Author Aman Priyanshu Overview This dataset is a cleaned and restructured version of interstellarninja/hermes_reasoning_tool_use. The original dataset uses the Hermes/NousResearch multi-turn format with from/value fields and embedded <think> + <tool_call> tags inside single gpt turns. This version converts it into a strict multi-turn conversation structure with validated role transitions.… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-TOOLS-hermes_reasoning_tool_use-data-cleaned-rectified.texttext-generation10K<n<100K2 likes176 downloads7mo agoHugging Face06rAVEUK /Kimi-K2.5-Reasoning-1M-Cleaned 🪐 Kimi-K2.5-Reasoning-1M-Cleaned Kimi-K2.5-Reasoning-1M-Cleaned is a cleaned derivative of ianncity/KIMI-K2.5-1000000x. It preserves the original four-config layout from the source dataset and rewrites each record into a unified reasoning-SFT schema with id, conversations, input, output, domain, and meta. Summary Source dataset: ianncity/KIMI-K2.5-1000000x Source author: ianncity Teacher model recorded in meta.teacher_model: KIMI-K2.5 Token lengths computed with… See the full description on the dataset page: https://huggingface.co/datasets/rAVEUK/Kimi-K2.5-Reasoning-1M-Cleaned.texttext-generation100K<n<1M2 likes163 downloads5mo agoHugging Face07EngMuhammadAtef /GLM-5.1-Reasoning-1M-Cleaned GLM-5.1-Reasoning-1M-Cleaned GLM-5.1-Reasoning-1M-Cleaned is a cleaned and reformatted derivative of Kassadin88/GLM-5.1-1000000x. It preserves the original four-subset layout (main, PHD-Science, Multilingual-STEM, Math) while converting every example into a unified SFT-ready schema with explicit conversations, input, output, domain, and meta fields. This release was prepared from the original dataset published by Kassadin88. Summary Teacher model in the data: GLM-5.1… See the full description on the dataset page: https://huggingface.co/datasets/EngMuhammadAtef/GLM-5.1-Reasoning-1M-Cleaned.texttext-generation100K<n<1M1 likes162 downloads5mo agoHugging Face08pinzhenchen /alpaca-cleaned-pt Data Description This HF data repository contains the Portuguese Alpaca dataset used in our study of monolingual versus multilingual instruction tuning. GitHub Paper Creation Machine-translated from yahma/alpaca-cleaned into Portuguese. Usage This data is intended to be used for Portuguese instruction tuning. The dataset has roughly 52K instances in the JSON format. Each instance has an instruction, an output, and an optional input. An example is shown… See the full description on the dataset page: https://huggingface.co/datasets/pinzhenchen/alpaca-cleaned-pt.texttext-generation10K<n<100K5 likes155 downloads3y agoHugging Face09VGS-AI /OpenR1-CleanedThis dataset is used in the paper Value-Guided Search for Efficient Chain-of-Thought Reasoning. It contains data for training and evaluating value models for improved long-context reasoning. GitHub Repository: https://github.com/kaiwenw/value-guided-search Related resources: Dataset (OpenR1-Cleaned): https://huggingface.co/datasets/VGS-AI/OpenR1-Cleaned Dataset (OpenR1-VM): https://huggingface.co/datasets/VGS-AI/OpenR1-VM Value Model (DeepSeek-VM-1.5B):… See the full description on the dataset page: https://huggingface.co/datasets/VGS-AI/OpenR1-Cleaned.textquestion-answering10K<n<100K0 likes153 downloads1y agoHugging Face10DanielSc4 /alpaca-cleaned-italian Dataset Card for Alpaca-Cleaned-Italian About the translation and the original data The translation was done with X-ALMA, a 13-billion-parameter model that surpasses state-of-the-art open-source multilingual LLMs (as of Q1 2025, paper here). The original alpaca-cleaned dataset is also kept here so that there is parallel data for Italian and English. Additional notes on the translation Despite the good quality of the translation, errors, though rare, are… See the full description on the dataset page: https://huggingface.co/datasets/DanielSc4/alpaca-cleaned-italian.texttext-generation100K<n<1M7 likes148 downloads2y agoHugging Face11Thaweewat /alpaca-cleaned-52k-th Summary This is a Thai 🇹🇭-instructed dataset translated from cleaned version of the original Alpaca Dataset released by Stanford using Google Cloud Translation, contain 52,000 instructions and demonstrations generated by OpenAI's text-davinci-003 engine. This instruction data can be used to conduct instruction-tuning for language models and make the language model follow instruction better. The following issues have been identified in the original release and fixed in this… See the full description on the dataset page: https://huggingface.co/datasets/Thaweewat/alpaca-cleaned-52k-th.textquestion-answering10K<n<100K17 likes144 downloads3y agoHugging Face12AmanPriyanshu /tool-reasoning-sft-TOOLS-toolace-sft-tool-use-agent-data-cleaned-rectified ToolACE - Tool-Use Agent Data Cleaned & Rectified 👥 Follow the Author Aman Priyanshu Overview This dataset is a cleaned and restructured version of the Team-ACE/ToolACE dataset. ToolACE is a high-quality conversational tool-use dataset containing 11,300+ examples of natural language interactions requiring function calling across diverse domains. This version converts the original OpenAI function-call format into a standardized multi-turn tool-use… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-TOOLS-toolace-sft-tool-use-agent-data-cleaned-rectified.tabulartext-generation10K<n<100K0 likes132 downloads7mo agoHugging Face13Aipresso /10k_rows_cleaned_prompts 10K Rows Cleaned Prompts Dataset Created by Aipresso LIMITED, London, UK ⚠️ IMPORTANT: By using this dataset, you agree to our Terms of Use You must provide attribution when using this data in publications, research, or commercial products. Dataset Overview A chunked collection of 2.7 million cleaned English prompts, organized into 200 files of 10,000 rows each for easy processing and distributed training of language models. 📊 Dataset Statistics Metric… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/10k_rows_cleaned_prompts.texttext-generation1M<n<10M0 likes127 downloads11mo agoHugging Face14ansulev /glm-5.1-reasoning-1m-cleaned GLM-5.1-Reasoning-1M-Cleaned GLM-5.1-Reasoning-1M-Cleaned is a cleaned and reformatted derivative of Kassadin88/GLM-5.1-1000000x. It preserves the original four-subset layout (main, PHD-Science, Multilingual-STEM, Math) while converting every example into a unified SFT-ready schema with explicit conversations, input, output, domain, and meta fields. This release was prepared from the original dataset published by Kassadin88. Summary Teacher model in the data:… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/glm-5.1-reasoning-1m-cleaned.texttext-generation100K<n<1M1 likes118 downloads5mo agoHugging Face15JBrightmanAI /Kimi-K2.5-Reasoning-1M-Cleaned 🪐 Kimi-K2.5-Reasoning-1M-Cleaned Kimi-K2.5-Reasoning-1M-Cleaned is a cleaned derivative of ianncity/KIMI-K2.5-1000000x. It preserves the original four-config layout from the source dataset and rewrites each record into a unified reasoning-SFT schema with id, conversations, input, output, domain, and meta. Summary Source dataset: ianncity/KIMI-K2.5-1000000x Source author: ianncity Teacher model recorded in meta.teacher_model: KIMI-K2.5 Token lengths computed… See the full description on the dataset page: https://huggingface.co/datasets/JBrightmanAI/Kimi-K2.5-Reasoning-1M-Cleaned.texttext-generation100K<n<1M1 likes117 downloads2mo agoHugging Face16tasal9 /ZamAI-Pashto-Dataset-Cleaned ZamAI Pashto Dataset Cleaned Languages: psLicense: apache-2.0Task categories: text-classification, text-generation, question-answeringSize categories: 10K<n<100K Summary This dataset is part of the ZamAI Pashto data collection. It is intended for text-classification, text-generation, question-answering tasks in Pashto. How to use from datasets import load_dataset dataset = load_dataset("tasal9/ZamAI-Pashto-Dataset-Cleaned") print(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/ZamAI-Pashto-Dataset-Cleaned.texttext-classification10K<n<100K0 likes107 downloads2mo agoHugging Face17AmanPriyanshu /tool-reasoning-sft-RESEARCH-grill-lab-browsecomp-plus-runs-data-cleaned-rectified Tool-Reasoning SFT — BrowseComp-Plus Runs (Cleaned & Rectified) Multi-turn tool-use reasoning trajectories derived from grill-lab/browsecomp-plus-runs, converted to a structured SFT format following the interstellarninja/hermes_reasoning_tool_use convention. Source Based on the execution trajectories from "Revisiting Text Ranking in Deep Research" (arXiv:2602.21456): Original data: grill-lab/browsecomp-plus-runs (MIT) Format Each row contains a messages… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-RESEARCH-grill-lab-browsecomp-plus-runs-data-cleaned-rectified.texttext-generation10K<n<100K0 likes97 downloads7mo agoHugging Face18zhangbo2008 /GLM-5.1-Reasoning-1M-Cleaned GLM-5.1-Reasoning-1M-Cleaned GLM-5.1-Reasoning-1M-Cleaned is a cleaned and reformatted derivative of Kassadin88/GLM-5.1-1000000x. It preserves the original four-subset layout (main, PHD-Science, Multilingual-STEM, Math) while converting every example into a unified SFT-ready schema with explicit conversations, input, output, domain, and meta fields. This release was prepared from the original dataset published by Kassadin88. Summary Teacher model in the data: GLM-5.1… See the full description on the dataset page: https://huggingface.co/datasets/zhangbo2008/GLM-5.1-Reasoning-1M-Cleaned.texttext-generation100K<n<1M0 likes91 downloads5mo agoHugging Face19pinzhenchen /alpaca-cleaned-es Data Description This HF data repository contains the Spanish Alpaca dataset used in our study of monolingual versus multilingual instruction tuning. GitHub Paper Creation Machine-translated from yahma/alpaca-cleaned into Spanish. Usage This data is intended to be used for Spanish instruction tuning. The dataset has roughly 52K instances in the JSON format. Each instance has an instruction, an output, and an optional input. An example is shown below: {… See the full description on the dataset page: https://huggingface.co/datasets/pinzhenchen/alpaca-cleaned-es.texttext-generation10K<n<100K4 likes90 downloads3y agoHugging Face20rlandismd /GLM-5.1-Reasoning-1M-Cleaned GLM-5.1-Reasoning-1M-Cleaned GLM-5.1-Reasoning-1M-Cleaned is a cleaned and reformatted derivative of Kassadin88/GLM-5.1-1000000x. It preserves the original four-subset layout (main, PHD-Science, Multilingual-STEM, Math) while converting every example into a unified SFT-ready schema with explicit conversations, input, output, domain, and meta fields. This release was prepared from the original dataset published by Kassadin88. Summary Teacher model in the data: GLM-5.1… See the full description on the dataset page: https://huggingface.co/datasets/rlandismd/GLM-5.1-Reasoning-1M-Cleaned.texttext-generation100K<n<1M0 likes77 downloads5mo agoHugging Face21BramVanroy /alpaca-cleaned-dutch Dataset Card for Alpaca Cleaned Dutch Dataset Summary This dataset contains 51,712 conversations between een AI assistant and a (fake) "Human" (generated) in Dutch. They are translations of Alpaca Cleaned Dataset. ☕ Want to help me out? Translating the data with the OpenAI API, and prompt testing, cost me 💸$57.99💸. If you like this dataset, please consider buying me a coffee to offset a portion of this cost, I appreciate it a lot! ☕ If you use this dataset or refer to… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/alpaca-cleaned-dutch.textquestion-answering10K<n<100K10 likes73 downloads3y agoHugging Face22pinzhenchen /alpaca-cleaned-cs Data Description This HF data repository contains the Czech Alpaca dataset used in our study of monolingual versus multilingual instruction tuning. GitHub Paper Creation Machine-translated from yahma/alpaca-cleaned into Czech. Usage This data is intended to be used for Czech instruction tuning. The dataset has roughly 52K instances in the JSON format. Each instance has an instruction, an output, and an optional input. An example is shown below: {… See the full description on the dataset page: https://huggingface.co/datasets/pinzhenchen/alpaca-cleaned-cs.texttext-generation10K<n<100K0 likes69 downloads3y agoHugging Face23mshojaei77 /persian-gk-cleanedThis is a cleaned and validated version of the original mshojaei77/persian-gk dataset. The purpose of this version is to ensure robust compatibility with modern fine-tuning workflows that rely on strict chat templates (e.g., tokenizer.apply_chat_template). The cleaning process resolves structural errors in the original dataset that could cause TemplateError or other silent failures during training with models like Gemma 3N, Llama 3, and others. Cleaning and Validation Process… See the full description on the dataset page: https://huggingface.co/datasets/mshojaei77/persian-gk-cleaned.textquestion-answering1K<n<10K1 likes69 downloads1y agoHugging Face24pinzhenchen /alpaca-cleaned-fr Data Description This HF data repository contains the French Alpaca dataset used in our study of monolingual versus multilingual instruction tuning. GitHub Paper Creation Machine-translated from yahma/alpaca-cleaned into French. Usage This data is intended to be used for French instruction tuning. The dataset has roughly 52K instances in the JSON format. Each instance has an instruction, an output, and an optional input. An example is shown below: {… See the full description on the dataset page: https://huggingface.co/datasets/pinzhenchen/alpaca-cleaned-fr.texttext-generation10K<n<100K1 likes65 downloads3y agoHugging Face25txchmechanicus /GLM-5.1-Reasoning-1M-Cleaned GLM-5.1-Reasoning-1M-Cleaned GLM-5.1-Reasoning-1M-Cleaned is a cleaned and reformatted derivative of Kassadin88/GLM-5.1-1000000x. It preserves the original four-subset layout (main, PHD-Science, Multilingual-STEM, Math) while converting every example into a unified SFT-ready schema with explicit conversations, input, output, domain, and meta fields. This release was prepared from the original dataset published by Kassadin88. Summary Teacher model in the data: GLM-5.1… See the full description on the dataset page: https://huggingface.co/datasets/txchmechanicus/GLM-5.1-Reasoning-1M-Cleaned.texttext-generation100K<n<1M0 likes65 downloads5mo agoHugging Face26MLDataScientist /SlimOrca-Dedup-Uzbek-cleanedThis is an Uzbek translated and cleaned version of https://huggingface.co/datasets/Open-Orca/SlimOrca-Dedup. Specifically, these replaced/removed records that had 'Uzbek translation|Uzbekcha tarjima|Uzbek tarjima|impossible to translate|not possible to translate|cannot fulfill your request|text is in|tilida yozilgan|Uzbek|o'zbek|ozbek|I am sorry'. You can use this dataset for chat fine-tuning of LLMs. This dataset has around 100M tokens (500M*0.8/4 = 100M assuming 4 chars are one token).… See the full description on the dataset page: https://huggingface.co/datasets/MLDataScientist/SlimOrca-Dedup-Uzbek-cleaned.texttext-generation100K<n<1M1 likes58 downloads2y agoHugging Face27Bas95 /GLM-5.1-Reasoning-1M-Cleaned GLM-5.1-Reasoning-1M-Cleaned GLM-5.1-Reasoning-1M-Cleaned is a cleaned and reformatted derivative of Kassadin88/GLM-5.1-1000000x. It preserves the original four-subset layout (main, PHD-Science, Multilingual-STEM, Math) while converting every example into a unified SFT-ready schema with explicit conversations, input, output, domain, and meta fields. This release was prepared from the original dataset published by Kassadin88. Summary Teacher model in the data: GLM-5.1… See the full description on the dataset page: https://huggingface.co/datasets/Bas95/GLM-5.1-Reasoning-1M-Cleaned.texttext-generation100K<n<1M0 likes56 downloads5mo agoHugging Face28kshitijthakkar /GLM-5.1-Reasoning-1M-Cleaned GLM-5.1-Reasoning-1M-Cleaned GLM-5.1-Reasoning-1M-Cleaned is a cleaned and reformatted derivative of Kassadin88/GLM-5.1-1000000x. It preserves the original four-subset layout (main, PHD-Science, Multilingual-STEM, Math) while converting every example into a unified SFT-ready schema with explicit conversations, input, output, domain, and meta fields. This release was prepared from the original dataset published by Kassadin88. Summary Teacher model in the data:… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/GLM-5.1-Reasoning-1M-Cleaned.text-generation100K<n<1M0 likes55 downloads5mo agoHugging Face29pinzhenchen /alpaca-cleaned-ru Data Description This HF data repository contains the Russian Alpaca dataset used in our study of monolingual versus multilingual instruction tuning. GitHub Paper Creation Machine-translated from yahma/alpaca-cleaned into Russian. Usage This data is intended to be used for Russian instruction tuning. The dataset has roughly 52K instances in the JSON format. Each instance has an instruction, an output, and an optional input. An example is shown below: {… See the full description on the dataset page: https://huggingface.co/datasets/pinzhenchen/alpaca-cleaned-ru.texttext-generation10K<n<100K1 likes54 downloads3y agoHugging Face30MyeongHo0621 /korean-quality-cleaned Korean Quality Dataset (Cleaned) 고품질 한국어 Instruction 데이터셋 (정제 버전) English Dataset Description This is a cleaned and standardized Korean instruction dataset, combining multiple high-quality open-source Korean datasets with unified formatting and quality filtering. Key Features ✅ Unified Format: Standardized messages format (OpenAI-compatible) ✅ Quality Filtering: Length, special characters, repetition filtering ✅ Clean Structure: Removed redundant… See the full description on the dataset page: https://huggingface.co/datasets/MyeongHo0621/korean-quality-cleaned.texttext-generation10K<n<100K0 likes53 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.