CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01KodCode /KodCode-V1-SFT-R1 🐱 KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding KodCode is the largest fully-synthetic open-source dataset providing verifiable solutions and tests for coding tasks. It contains 12 distinct subsets spanning various domains (from algorithmic to package-specific knowledge) and difficulty levels (from basic coding exercises to interview and competitive programming challenges). KodCode is designed for both supervised fine-tuning (SFT) and RL tuning. 🕸️… See the full description on the dataset page: https://huggingface.co/datasets/KodCode/KodCode-V1-SFT-R1.tabularquestion-answering100K<n<1M40 likes15k downloads2y agoHugging Face02SHSLab /Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection 🧬 Omni-Frontier Collection Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible. 📖 Jump to What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/SHSLab/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.tabulartext-generation10M<n<100M3 likes4.6k downloads27d agoHugging Face03KodCode /KodCode-V1-SFT-4o 🐱 KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding KodCode is the largest fully-synthetic open-source dataset providing verifiable solutions and tests for coding tasks. It contains 12 distinct subsets spanning various domains (from algorithmic to package-specific knowledge) and difficulty levels (from basic coding exercises to interview and competitive programming challenges). KodCode is designed for both supervised fine-tuning (SFT) and RL tuning. 🕸️… See the full description on the dataset page: https://huggingface.co/datasets/KodCode/KodCode-V1-SFT-4o.tabularquestion-answering100K<n<1M10 likes2.6k downloads2y agoHugging Face04soketlabs /bhasha-sft Bhasha SFT Bhasha SFT is a massive collection of multiple open sourced Supervised Fine-Tuning datasets for training Multilingual Large Language Models. The dataset contains collation of over 13 million instances of instruction-response data for 3 Indian languages (Hindi, Gujarati, Bengali) and English having both human annotated and synthetic data. Curated by: Soket AI Labs Language(s) (NLP): [English, Hindi, Bengali, Gujarati] License: [cc-by-4.0, apache-2.0, mit]… See the full description on the dataset page: https://huggingface.co/datasets/soketlabs/bhasha-sft.tabularquestion-answering10M<n<100M4 likes625 downloads2y agoHugging Face05Manusagents /Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2 🧬 Omni-Frontier Collection Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible. 📖 Jump to What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2.tabulartext-generation10M<n<100M0 likes562 downloads26d agoHugging Face06Congliu /Chinese-DeepSeek-R1-Distill-data-110k-SFT 中文基于满血DeepSeek-R1蒸馏数据集(Chinese-Data-Distill-From-R1) 🤗 Hugging Face   |   🤖 ModelScope    |   🚀 Github    |   📑 Blog 注意:该版本为,可以直接SFT使用的版本,将原始数据中的思考和答案整合成output字段,大部分SFT代码框架均可直接直接加载训练。 本数据集为中文开源蒸馏满血R1的数据集,数据集中不仅包含math数据,还包括大量的通用类型数据,总数量为110K。 为什么开源这个数据? R1的效果十分强大,并且基于R1蒸馏数据SFT的小模型也展现出了强大的效果,但检索发现,大部分开源的R1蒸馏数据集均为英文数据集。 同时,R1的报告中展示,蒸馏模型中同时也使用了部分通用场景数据集。 为了帮助大家更好地复现R1蒸馏模型的效果,特此开源中文数据集。该中文数据集中的数据分布如下: Math:共计36568个样本, Exam:共计2432个样本, STEM:共计12648个样本,… See the full description on the dataset page: https://huggingface.co/datasets/Congliu/Chinese-DeepSeek-R1-Distill-data-110k-SFT.tabulartext-generation100K<n<1M225 likes419 downloads2y agoHugging Face07MercanAI /turkce-sft-qa-3.7m 🇹🇷 Turkish SFT/QA — Birleştirilmiş ve Tekrarsız Veri Seti 3,723,264 örnek. 24 açık Türkçe SFT/QA veri setinin, satır düzeyinde tekrar temizliği ve kalite kontrolünden geçirilmiş birleşimi. Her satır hangi veri setinden geldiğini taşır. English: A merged, row-level deduplicated and quality-filtered collection of 24 open Turkish SFT/QA datasets (3,723,264 examples). Every row carries its source dataset, source URL and original license. 🙏 Teşekkür /… See the full description on the dataset page: https://huggingface.co/datasets/MercanAI/turkce-sft-qa-3.7m.tabulartext-generation1M<n<10M0 likes347 downloads2mo agoHugging Face08zachnorton03 /synthetic-pre1930-sftTL;DR A vintage finetuning dataset (~416k rows, eleven task routes). Sourced by taking excerpts from pre-1930's texts, turning these into verbatim answers, and then using deepseek-chat to generate period-appropriate questions of those answers. Any model tuned on this dataset should, theoretically, never update its weights on anachronistic text, since questions are masked in the finetuning stages. Features composition, verse, narrative, reasoning, multiturn dialogue, and calibrated uncertainty… See the full description on the dataset page: https://huggingface.co/datasets/zachnorton03/synthetic-pre1930-sft.tabularquestion-answering100K<n<1M2 likes313 downloads1mo agoHugging Face09gyung /korean-bar-exam-hard-current-law-precedent-sft-1000 Korean Current-Law Bar Exam Hard SFT 1000 대한민국 현행 법령을 기준으로 만든 변호사시험 선택형 고난도 스타일 SFT 데이터 1,000문항입니다. 초기 직접 조문확인형 생성본은 실제 제14ㆍ15회 변호사시험보다 쉬워서, 이 버전은 다음 기준으로 다시 만들었습니다. ㄱ/ㄴ/ㄷ/ㄹ 복합정오형 중심 甲/乙/丙, 검사ㆍ사법경찰관ㆍ행정청ㆍ회사ㆍ소송당사자 등이 등장하는 사례형 비중 확대 단순 근거 조문 선택형 제거 정답뿐 아니라 각 지문별 O/X 이유와 참고 법령 조문 제공 제15회 변호사시험 data/questions.csv와 높은 유사도 문항 제외 Files data/questions.csv: Hugging Face preview용 메인 CSV입니다. sft/train.jsonl: messages 형식 SFT용 JSONL입니다. metadata/qa_report.json: 생성 수량, 난도 관련… See the full description on the dataset page: https://huggingface.co/datasets/gyung/korean-bar-exam-hard-current-law-precedent-sft-1000.tabularquestion-answering1K<n<10K0 likes306 downloads3mo agoHugging Face10lm-provers /FineProofs-SFT FineProofs SFT Dataset Description FineProofs SFT is a high-quality supervised fine-tuning dataset containing mathematical Olympiad problems paired with chain-of-thought reasoning and formal proofs distilled from DeepSeek-Math-V2. The dataset comprises 7,777 samples (4,300 unique problems) sourced from international Olympiad competitions and Art of Problem Solving (AoPS), each annotated with: Detailed reasoning traces (thinking content) generated by… See the full description on the dataset page: https://huggingface.co/datasets/lm-provers/FineProofs-SFT.tabulartext-generation10K<n<100K43 likes280 downloads7mo agoHugging Face11TypeSafeAI /Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection 🧬 Omni-Frontier Collection Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible. 📖 Jump to What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/TypeSafeAI/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.tabulartext-generation10M<n<100M1 likes170 downloads2d agoHugging Face12noah248 /chinese-legal-sft Chinese Legal SFT Dataset(中文法律 SFT 数据集) 面向大模型监督微调(SFT)的中文法律问答数据集,共 19,332 条问答对, 每条附带 LLM 质量评分。覆盖数据采集 → 清洗 → 去重 → 质量过滤 → 格式化 → 质量打分的完整数据工程流程。 配套代码与完整流水线:https://github.com/noah-white-python/legal-sft-dataset 数据构建流程 冷启动:基于开源数据集 DISC-Law-SFT 整理。 清洗:NFKC 全角半角统一、去控制字符、去空白、缺失过滤。 去重:精确去重(MD5)+ MinHash + LSH 近似去重(阈值 0.8)。 质量过滤:长度、中文字符占比等启发式规则,有效率 96.7%(20,000 → 19,332)。 格式化:输出标准 Alpaca 指令格式。 质量打分:用 LLM-as-judge 对全部数据从复杂度、清晰度、信息量三维度打分(1-5 分)。 字段说明… See the full description on the dataset page: https://huggingface.co/datasets/noah248/chinese-legal-sft.tabularquestion-answering10K<n<100K0 likes145 downloads3mo agoHugging Face13matonski /toy-models-of-sft-data Toy Models of SFT Data This is a public-clean candidate data package for the Toy Models of SFT project. It is built for researcher inspection first. The package answers two questions: What were the models trained on? How did the models actually behave under evaluation? The package includes training data, eval inputs, model rollouts, judge scores, parsed GPQA outputs, aggregate tables, paper figures, frozen plot data, and provenance records. It deliberately includes some… See the full description on the dataset page: https://huggingface.co/datasets/matonski/toy-models-of-sft-data.tabulartext-generation10K<n<100K0 likes140 downloads2mo agoHugging Face14OliveiraJLT /gigaverbo-v2-rec-sft GigaVerbo-v2 REC SFT A model should not merely know how to reason; it should learn when reasoning is worth the cost. Dataset repository: OliveiraJLT/gigaverbo-v2-rec-sftBase dataset: Polygl0t/gigaverbo-v2-sftAnswer-generation model: openai/gpt-oss-20bQuality classifier: Polygl0t/portuguese-qwen3-4b-instruct-quality-classifierReasoning translation model and token accounting tokenizer: Qwen/Qwen3.5-9B Dataset Summary GigaVerbo-v2 REC SFT — short for GigaVerbo-v2… See the full description on the dataset page: https://huggingface.co/datasets/OliveiraJLT/gigaverbo-v2-rec-sft.tabulartext-generation100K<n<1M0 likes138 downloads4mo agoHugging Face15xmanii /Maux-Persian-SFT-30k Maux-Persian-SFT-30k Dataset Description This dataset contains 30,000 high-quality Persian (Farsi) conversations for supervised fine-tuning (SFT) of conversational AI models. The dataset combines multiple sources to provide diverse, natural Persian conversations covering various topics and interaction patterns. Dataset Structure Each entry contains: messages: List of conversation messages with role (user/assistant/system) and content source: Source dataset… See the full description on the dataset page: https://huggingface.co/datasets/xmanii/Maux-Persian-SFT-30k.tabularquestion-answering10K<n<100K4 likes136 downloads1y agoHugging Face16AmanPriyanshu /tool-reasoning-sft-TOOLS-toolace-sft-tool-use-agent-data-cleaned-rectified ToolACE - Tool-Use Agent Data Cleaned & Rectified 👥 Follow the Author Aman Priyanshu Overview This dataset is a cleaned and restructured version of the Team-ACE/ToolACE dataset. ToolACE is a high-quality conversational tool-use dataset containing 11,300+ examples of natural language interactions requiring function calling across diverse domains. This version converts the original OpenAI function-call format into a standardized multi-turn tool-use… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-TOOLS-toolace-sft-tool-use-agent-data-cleaned-rectified.tabulartext-generation10K<n<100K0 likes125 downloads7mo agoHugging Face17Jackrong /LogosForge-scored-sft-v1 LogosForge-scored-sft-v1 This dataset is a scored Supervised Fine-Tuning (SFT) distillation corpus built on top of the Natural Reasoning question set.It is constructed in two stages: First, a large-scale reasoning-oriented teacher model (gpt-oss-120B-high) is used to generate distilled student responses, including explicit chain-of-thought reasoning, for natural reasoning questions. Second, these distilled responses are evaluated by a separate instruction-following model… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/LogosForge-scored-sft-v1.tabulartext-generation10K<n<100K2 likes107 downloads8mo agoHugging Face18Training-Datasmith /k3-sft-cc0-flan Dataset Card for K3 SFT CC0 FLAN 844-row Kimi K3 synthetic instruction-tuning shard built from DPI-traced CC0/public-domain FLAN prompts in the Tülu mix. Four overlapping Hub configs expose different cohort views; adaptive is the recommended default for quality-conscious SFT mixing. Dataset Details Curated by: Training Datasmith Teacher: kimi-k3 via deltafin (local inference) Languages: English prompts; translation pairs include German, Spanish, Czech, Igbo… See the full description on the dataset page: https://huggingface.co/datasets/Training-Datasmith/k3-sft-cc0-flan.tabulartext-classification1K<n<10K0 likes103 downloads6d agoHugging Face19TTP01 /Vietverse-SFT-1K-Gold 🇻🇳 Vietverse-SFT (1K Gold Edition) The "Less is More" Alignment Paradigm for Native Vietnamese Large Language Models Bộ Dữ Liệu SFT Tiếng Việt Bản Xứ 1.000 Mẫu Gold Tinh Hoa — Chuẩn Mực Căn Chỉnh Mô Hình Ngôn Ngữ 🇻🇳 [Đọc Báo Cáo Kỹ Thuật Tiếng Việt] &nbsp;•&nbsp; 🇬🇧 [Read English Technical Card] 🤗 Hugging Face Dataset • ⚡ Hướng Dẫn Huấn Luyện / Quickstart 🌐 Ngôn Ngữ / Language 📌 Chuyển Hướng Nhanh / Quick Jump… See the full description on the dataset page: https://huggingface.co/datasets/TTP01/Vietverse-SFT-1K-Gold.tabulartext-generation1K<n<10K4 likes95 downloads1mo agoHugging Face20sxiong /MLR_sft_data MLR SFT Data MLR SFT Data is a teacher-generated supervised fine-tuning dataset for training Multi-Level Reasoning (MLR) models in the paper Enhancing Language Model Reasoning with Structured Multi-Level Modeling (ICLR 2026). It decomposes complete reasoning trajectories into two types of step-level examples: Planner: plans the next reasoning goal and task based on the problem and reasoning history. Executor: executes the Planner's instruction and updates the reasoning state.… See the full description on the dataset page: https://huggingface.co/datasets/sxiong/MLR_sft_data.tabularquestion-answering100K<n<1M1 likes93 downloads12d agoHugging Face21trillionlabs /SimScholar-SFT S3 SFT Trajectories Complete ReAct trajectories for scientific-literature search. Code · S3 collection · Source corpus This dataset contains 14,633 single- and two-hop tool-use trajectories. In each trajectory, a policy searches and reads a fixed scientific corpus through nine tools, then submits an answer with a correctness label. The messages column uses OpenAI tool-calling chat format. At a glance Question type Rows Correct Incorrect Single-hop… See the full description on the dataset page: https://huggingface.co/datasets/trillionlabs/SimScholar-SFT.tabularquestion-answering10K<n<100K0 likes85 downloads2mo agoHugging Face22Jackrong /LogicMind-Chat-Reasoning-SFT-300K Nemotron-Post-Training-Dataset-v2-chat Dataset Card Overview 📌 This dataset contains 296,168 chat-style instruction/response samples generated by qwen-3-32b. Each record provides a user prompt, an explicit reasoning trace, and a final answer, plus precomputed length fields. The data is packaged as JSONL (one JSON object per line). Highlights Scale: 296,168 samples Category: chat (100%) Generator: qwen-3-32b (100%) Structure: problem → qwen3-reasoning → qwen3-solution… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/LogicMind-Chat-Reasoning-SFT-300K.tabularquestion-answering100K<n<1M10 likes80 downloads8mo agoHugging Face23malr07 /opc-sft-stage2-dense-extracted OpenCoder Dataset Dense Region Extracted This dataset is a post-processed version of the OpenCoder SFT Stage2 dataset (opc-sft-stage2). We use gpt-4o API to extract the information dense regions from each sample and logged them in the dense_snippets column.Detailed information about the data can be found in our paper. OpenCoder's sft-stage2 summary The original version of this dataset is used in OpenCoder's Stage 2 and consists of four parts: educational_instruct:… See the full description on the dataset page: https://huggingface.co/datasets/malr07/opc-sft-stage2-dense-extracted.tabulartext-generation100K<n<1M0 likes72 downloads6mo agoHugging Face24somasekhar-dev /nexttoken-pmkisan-domain-sft-data NextToken pmkisan domain SFT data (v1) Grounded multilingual QA dataset for fine-tuning somasekhar-dev/NextToken-model-1 on the Indian government-schemes / banking-financial domain. Generated by a pipeline (chunk source docs -> generate questions -> generate grounded answers -> validate/assemble) using a local LLM generator, from ~57 scheme/product source documents (PM-KISAN, Ayushman Bharat, MGNREGA, banking products, insurance, savings instruments, etc.). Files… See the full description on the dataset page: https://huggingface.co/datasets/somasekhar-dev/nexttoken-pmkisan-domain-sft-data.tabularquestion-answering1K<n<10K0 likes55 downloads8d agoHugging Face25forseasons /sft-data-1 sft-data-1 This dataset contains 100 SFT trajectories for DABench-style data-agent training. Contents train_sft.jsonl: JSONL SFT records. train_sft.parquet: Parquet version of the same records. summary.json: generation and filtering summary. Generation Expert model metadata: gpt-5.5 Tool/action format: qwen3_xml Model judge: glm-5.2 Model-judge completed: 266 Model-judge accepted: 143 Selected SFT rows: 100 Assistant actions are serialized as… See the full description on the dataset page: https://huggingface.co/datasets/forseasons/sft-data-1.tabularquestion-answeringn<1K0 likes42 downloads3mo agoHugging Face26TTP01 /Vietverse-SFT-Preview 🇻🇳 Vietverse-SFT v1.0 (Preview Edition) A High-Density, Native Vietnamese Foundation Dataset for Supervised Fine-Tuning Bộ Dữ Liệu Nền Tảng SFT Tiếng Việt Bản Xứ Chuẩn Mực Cho Huấn Luyện Mô Hình Ngôn Ngữ Lớn 🇻🇳 [Đọc Bản Tiếng Việt] &nbsp;•&nbsp; 🇬🇧 [Read English Version] 🤗 Hugging Face Dataset • ⚡ Hướng Dẫn Sử Dụng / Quickstart 🌐 Ngôn Ngữ / Language 📌 Chuyển Hướng / Quick Navigation 🇻🇳 Bản Tiếng Việt Đầy Đủ… See the full description on the dataset page: https://huggingface.co/datasets/TTP01/Vietverse-SFT-Preview.tabulartext-generationn<1K3 likes36 downloads1mo agoHugging Face27hozifa1 /faqih_sft_dataset 💎 MAFQA: Perfected Multi-Hop Arabic Fatwa QA Dataset (388 Samples) مجموعة بيانات الاستدلال الفقهي المركب ومتعدد الخطوات (جامعة الملك سعود / MDPI 2026) 100% Curated & Unabridged MAFQA Multi-Hop Dataset تم تنقيح وتدقيق البيانات بالكامل: 1. إزالة جميع التقطيعات النصية وإيراد النصوص والأدلة كاملة دون بتر. 2. تفعيل خطوة التركيب والترجيح النهائي (الخطوة 4) بربط استدلالي حقيقي بين المسائل الفرعية. 3. تصحيح الأخطاء المطبعية في دلالات الحل والحرمة.… See the full description on the dataset page: https://huggingface.co/datasets/hozifa1/faqih_sft_dataset.tabulartext-generation1K<n<10K0 likes34 downloads1mo agoHugging Face28mai-ll /ipw-sft-trajectories IPW SFT Trajectories Supervised fine-tuning trajectories for training tool-use orchestrator models. Each trajectory contains a multi-turn conversation where a model solves a task by selecting and using tools (calculator, code interpreter, think). Dataset Details Stat Value Total trajectories (deduped, correct only) 44,301 Total trajectories (with all opus) 44,866 Categories 2,122 Avg turns per trajectory 1.2 Source dataset GeneralThought… See the full description on the dataset page: https://huggingface.co/datasets/mai-ll/ipw-sft-trajectories.tabulartext-generation10K<n<100K0 likes31 downloads6mo agoHugging Face29Finnish-NLP /belebele-fi-filtered-sft Dataset Card for Finnish-NLP/benebele Creation process Finnish subset loaded from facebook/belebele tabulartext-generationn<1K0 likes27 downloads3y agoHugging Face30drguolai /distill_r1_110k_sft_modifiedBorrowed from https://huggingface.co/datasets/Congliu/Chinese-DeepSeek-R1-Distill-data-110k-SFT Fix the <image> placeholder issue, which will cause error during training: raise ValueError(f"The number of images does not match the number of {IMAGE_PLACEHOLDER} tokens.") tabularquestion-answering100K<n<1M0 likes20 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.