CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01chendelong /MEP-3Mtext1M<n<10M5 likes5.4k downloads3y agoHugging Face02Yxanul /Mephisto-Knowledge_538k Mephisto-Knowledge_538k 538,861 English knowledge SFT examples generated by Qwen/Qwen3.5-4B in non-thinking (Instruct) mode on the Knowledge prompts of openbmb/UltraData-SFT-2605. Responses contain no chain-of-thought — thinking was disabled at generation time, so each assistant turn is a direct answer, usually with a short justification. Companion dataset: Mephisto-IF_172k (instruction-following, same teacher and pipeline). Read this before training: ref_agrees… See the full description on the dataset page: https://huggingface.co/datasets/Yxanul/Mephisto-Knowledge_538k.textquestion-answering100K<n<1M2 likes291 downloads2mo agoHugging Face03Yxanul /Mephisto-IF_172k Mephisto-IF_172k 172,761 English instruction-following SFT examples, generated by Qwen/Qwen3.5-4B in non-thinking (Instruct) mode on the instruction-following prompts of openbmb/UltraData-SFT-2605. Responses contain no chain-of-thought — thinking was disabled at generation time, so every assistant turn is a direct answer. Format One JSON object per line: { "messages": [ {"role": "user", "content": "..."}, {"role": "assistant", "content": "..."} ]… See the full description on the dataset page: https://huggingface.co/datasets/Yxanul/Mephisto-IF_172k.texttext-generation100K<n<1M0 likes122 downloads2mo agoHugging Face04zixiaozhu /MePO MePO Prompt Optimization Dataset This dataset is designed for research in prompt optimization, particularly for training and evaluating MePO — a lightweight, locally deployable prompt optimization model. 📂 File: MePO.jsonl (40,151 entries) Each JSONL record includes: rejectedThe original prompt from BPO or Alpaca, used as the rejected example. chosenThe optimized prompt generated by MePO, used as the chosen example. sliver_responseThe response produced from the… See the full description on the dataset page: https://huggingface.co/datasets/zixiaozhu/MePO.text10K<n<100K1 likes54 downloads1y agoHugging Face05Mephesto1 /Diabetic-trackertextn<1K1 likes53 downloads3mo agoHugging Face06zixiaozhu /MePO_BPO MePO Prompt Optimization Dataset (BPO version) This dataset is designed for research in prompt optimization, particularly for training and evaluating MePO — a lightweight, locally deployable prompt optimization model. Each JSONL record includes: rejectedThe original prompt from BP, used as the rejected example. chosenThe optimized prompt generated by MePO, used as the chosen example. sliver_responseThe response produced from the BPO prompt (baseline response). golden_responseThe… See the full description on the dataset page: https://huggingface.co/datasets/zixiaozhu/MePO_BPO.texttext-generation10K<n<100K1 likes49 downloads1y agoHugging Face07zixiaozhu /MePO_Alpaca 📦 MePO Prompt Optimization Dataset (Alpaca Version) The MePO Prompt Optimization Dataset is designed to support research in prompt optimization, especially for training and evaluating MePO — a lightweight and locally deployable prompt optimization model. 📁 Dataset Structure Each .jsonl record contains the following fields: rejectedThe original prompt from the Alpaca dataset, serving as the rejected example. chosenThe optimized prompt generated by MePO, serving as the… See the full description on the dataset page: https://huggingface.co/datasets/zixiaozhu/MePO_Alpaca.text10K<n<100K0 likes39 downloads1y agoHugging Face08mepartha /Agentic-Chain-of-Thought-Coding-SFT-Dataset-v1.1 🤖 Agentic Coding CoT Dataset v1.1 A high-quality supervised fine-tuning (SFT) dataset for training agentic coding assistants with Chain-of-Thought reasoning capabilities. 📋 Dataset Description This dataset was created by processing and distilling ~20GB of GitHub crawl data using Minimax-M2 & MiniMax M2.1 to generate structured, reasoning-rich coding examples. Each sample demonstrates systematic problem-solving with explicit tool usage patterns. 🏗️ Assistant… See the full description on the dataset page: https://huggingface.co/datasets/mepartha/Agentic-Chain-of-Thought-Coding-SFT-Dataset-v1.1.texttext-generation1K<n<10K3 likes21 downloads5mo agoHugging Face09mepartha /openscad-vision-sfttexttext-generation1K<n<10K0 likes13 downloads1y agoHugging Face10misclassified /meps_speeches_with_translation.csvtabular10K<n<100K0 likes12 downloads3y agoHugging Face11misclassified /meps_speechesThis dataset contains nearly 18,000 European Member of Parliament (meps) speeches beween 2019 and 2023. The speeches are from Italian, German, French and Belgium meps. All the speeches were gently scraped for the european parliament website using this code: https://github.com/misclassified/meps-text-mining text10K<n<100K1 likes10 downloads3y agoHugging Face120xZee /list_meptextn<1K0 likes8 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.