CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01KodCode /KodCode-V1-SFT-R1 🐱 KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding KodCode is the largest fully-synthetic open-source dataset providing verifiable solutions and tests for coding tasks. It contains 12 distinct subsets spanning various domains (from algorithmic to package-specific knowledge) and difficulty levels (from basic coding exercises to interview and competitive programming challenges). KodCode is designed for both supervised fine-tuning (SFT) and RL tuning. 🕸️… See the full description on the dataset page: https://huggingface.co/datasets/KodCode/KodCode-V1-SFT-R1.tabularquestion-answering100K<n<1M40 likes15k downloads2y agoHugging Face02geodesic-research /pa-warm-start-sft-heavy-25b-mix geodesic-research/pa-warm-start-sft-heavy-25b-mix Auto-generated by dataset-builder. Each config below is a separate dataset produced from a versioned YAML build config. Load with: from datasets import load_dataset ds = load_dataset("geodesic-research/pa-warm-start-sft-heavy-25b-mix", "<config_name>", revision="<commit-sha>") Pin revision= to the specific commit SHA you want; without it, you get the current HEAD of the dataset repo, which may change when the builder re-pushes.… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/pa-warm-start-sft-heavy-25b-mix.tabular10M<n<100M0 likes8.5k downloads21d agoHugging Face03mlfoundations-dev /Eurus-2-7B-SFT_eval_2e29 mlfoundations-dev/Eurus-2-7B-SFT_eval_2e29 Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces AIME25 HLE LiveCodeBenchv5 Accuracy 2.3 21.0 30.6 11.0 11.4 10.4 6.8 1.5 2.1 1.3 4.1 4.4 AIME24 Average Accuracy: 2.33% ± 0.67% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 0.00% 0 30 2 3.33% 1… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/Eurus-2-7B-SFT_eval_2e29.tabular1K<n<10K0 likes5.5k downloads1y agoHugging Face04SHSLab /Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection 🧬 Omni-Frontier Collection Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible. 📖 Jump to What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/SHSLab/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.tabulartext-generation10M<n<100M3 likes4k downloads24d agoHugging Face05geodesic-research /pa-warm-start-sft-xl-50b-mix geodesic-research/pa-warm-start-sft-xl-50b-mix Auto-generated by dataset-builder. Each config below is a separate dataset produced from a versioned YAML build config. Load with: from datasets import load_dataset ds = load_dataset("geodesic-research/pa-warm-start-sft-xl-50b-mix", "<config_name>", revision="<commit-sha>") Pin revision= to the specific commit SHA you want; without it, you get the current HEAD of the dataset repo, which may change when the builder re-pushes.… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/pa-warm-start-sft-xl-50b-mix.tabular10M<n<100M0 likes3.9k downloads9d agoHugging Face06joshycodes /sorrel-sft-voicetabular100K<n<1M0 likes2.5k downloads5d agoHugging Face07KodCode /KodCode-V1-SFT-4o 🐱 KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding KodCode is the largest fully-synthetic open-source dataset providing verifiable solutions and tests for coding tasks. It contains 12 distinct subsets spanning various domains (from algorithmic to package-specific knowledge) and difficulty levels (from basic coding exercises to interview and competitive programming challenges). KodCode is designed for both supervised fine-tuning (SFT) and RL tuning. 🕸️… See the full description on the dataset page: https://huggingface.co/datasets/KodCode/KodCode-V1-SFT-4o.tabularquestion-answering100K<n<1M10 likes2.4k downloads2y agoHugging Face08geodesic-research /pa-warm-start-sft-xl-smoketabular10K<n<100K0 likes1.3k downloads10d agoHugging Face09OpenMOSS-Team /moss-002-sft-data Dataset Card for "moss-002-sft-data" Dataset Summary An open-source conversational dataset that was used to train MOSS-002. The user prompts are extended based on a small set of human-written seed prompts in a way similar to Self-Instruct. The AI responses are generated using text-davinci-003. The user prompts of en_harmlessness are from Anthropic red teaming data. Data Splits name # samples en_helpfulness.json 419049 en_honesty.json 112580… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/moss-002-sft-data.tabulartext-generation1M<n<10M96 likes1.2k downloads3y agoHugging Face10toroe /ReasonXL-SFT ReasonXL: A Multilingual Cross-Domain Reasoning Corpus ReasonXL is a large-scale multilingual reasoning corpus spanning five languages, with 2,538,450 positionally aligned examples per language (12,692,250 rows total). It is designed to support supervised fine-tuning of reasoning models with in-language chain-of-thought traces across diverse technical domains. Data Generation English source samples were drawn from 10 existing reasoning datasets, filtered and… See the full description on the dataset page: https://huggingface.co/datasets/toroe/ReasonXL-SFT.tabular10M<n<100M0 likes1.2k downloads2mo agoHugging Face11kelexine /fable-5-sft-traces Fable-5 SFT Traces Author / maintainer: kelexine (github.com/kelexine) A cleaned, anonymised, schema-normalised derivative of Kelexine/Fable-5-traces — agentic traces from Fable-5 (claude-fable-5), the model now publicly known as Claude Mythos — Anthropic's top-of-family frontier model at time of collection. The dataset supports three fine-tuning shapes off a single JSONL with no preprocessing required: Mode Fields used Full SFT (thinking + response) messages or… See the full description on the dataset page: https://huggingface.co/datasets/kelexine/fable-5-sft-traces.tabulartext-generation1K<n<10K14 likes1.2k downloads3mo agoHugging Face12chankhavu /smolmo-sft-v2-seqlen64k smolmo-sft-v2-seqlen64k A supervised fine-tuning (SFT) dataset of math problems with full chain-of-thought solutions, formatted for the Olmo 3 "Thinking" models. 2,813,055 examples · ~37.9 B tokens. Three task families: proofs, numeric-answer problems, and tool-augmented (Python) problems. Every assistant turn carries an explicit <think> … </think> reasoning trace before the answer. Olmo 3 native chat + function-calling format; every example fits within a 64k-token context.… See the full description on the dataset page: https://huggingface.co/datasets/chankhavu/smolmo-sft-v2-seqlen64k.tabulartext-generation1M<n<10M0 likes1.1k downloads4mo agoHugging Face13ByteDance /VR-X-SFT-RL VR-X: Visual Reasoning Benchmark for UniVR VR-X contains three independent data blocks: SFT data organized by capability. VR-X-RL data for visual-reasoning reinforcement learning. VR-X-Eval held-out evaluation data. VR-X-RL and VR-X-Eval are independent from SFT and must be loaded separately. Public repository paths use anonymous source codes; no source-to-code mapping is published. Repository layout . ├── Robot Manipulation/ # SFT only │ └── RM-###/ │… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance/VR-X-SFT-RL.tabularvisual-question-answering100K<n<1M0 likes978 downloads2mo agoHugging Face14geodesic-research /pa-warm-start-sft-medium-5b-mix geodesic-research/pa-warm-start-sft-medium-5b-mix Auto-generated by dataset-builder. Each config below is a separate dataset produced from a versioned YAML build config. Load with: from datasets import load_dataset ds = load_dataset("geodesic-research/pa-warm-start-sft-medium-5b-mix", "<config_name>", revision="<commit-sha>") Pin revision= to the specific commit SHA you want; without it, you get the current HEAD of the dataset repo, which may change when the builder re-pushes.… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/pa-warm-start-sft-medium-5b-mix.tabular1M<n<10M0 likes798 downloads1mo agoHugging Face15argo11 /0399-tv-valid-clean-sft-tokenized-llmjp4-8btabular1M<n<10M0 likes796 downloads3mo agoHugging Face16laion /tts-realspeech-sft-en-de LAION TTS Real-Speech SFT — English + German, emotion-balanced 1,947,272 real recorded utterances — no synthetic voices — selected from freely-licensed corpora and balanced across 40 emotions x 2 languages. 6,996 hours, 313,844,544 MOSS frames (3,766,134,528 audio tokens), 79,337,527 aligned words. Each row is a self-contained TTS example: a corrected procedural caption, the transcript with word-level timestamps, the original audio, and the target MOSS-Audio-Tokenizer-v2 codes.… See the full description on the dataset page: https://huggingface.co/datasets/laion/tts-realspeech-sft-en-de.tabulartext-to-speech1M<n<10M0 likes785 downloads13d agoHugging Face17AdaMLLab /smolkalam-arabic-conversational-sft SmolKalam SmolKalam is a quality-filtered Arabic SFT dataset of 1,790,478 examples (~2.45B tokens), built as an ensemble translation of SmolTalk2. It covers multi-turn dialogue (23% of rows), reasoning traces (19% carry <think>), tool and function calling (4.4%), and long context, categories that are underrepresented in existing Arabic post-training data. The SmolTalk2 source mixtures are kept as subsets. Released with the paper SmolKalam: Ensemble Quality-Filtered Translation… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/smolkalam-arabic-conversational-sft.tabulartext-generation1M<n<10M3 likes770 downloads1mo agoHugging Face18Polygl0t /gigaverbo-v2-sft GigaVerbo-v2 SFT: A Large-Scale Portuguese Instruction-Tuning Dataset Dataset Summary GigaVerbo-v2 SFT is a large-scale instruction-tuning dataset designed for supervised fine-tuning of language models in Portuguese. The dataset comprises approximately 2.1 billion tokens (~4.4 GB) across 4 million instruction-following examples, organized into 12 distinct task categories. It is entirely composed of high-quality, LLM-generated data that has been carefully curated and… See the full description on the dataset page: https://huggingface.co/datasets/Polygl0t/gigaverbo-v2-sft.imagetext-generation1M<n<10M3 likes767 downloads7mo agoHugging Face19onnoboru /maniskill3-sft-rgb-lerobot-1200epThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "maniskill3_sim", "total_episodes": 1200, "total_frames": 171480, "total_tasks": 6, "total_videos": 0, "total_chunks": 2, "chunks_size": 1000, "fps": 20, "splits": { "train": "0:1200" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path": null, "features":… See the full description on the dataset page: https://huggingface.co/datasets/onnoboru/maniskill3-sft-rgb-lerobot-1200ep.tabularrobotics100K<n<1M0 likes736 downloads3mo agoHugging Face20yikeee /rubrichub-sft-judgment-gentabular100K<n<1M0 likes718 downloads6mo agoHugging Face21nyu-dice-lab /lm-eval-results-princeton-nlp-Llama-3-Base-8B-SFT-RDPO-private Dataset Card for Evaluation run of princeton-nlp/Llama-3-Base-8B-SFT-RDPO Dataset automatically created during the evaluation run of model princeton-nlp/Llama-3-Base-8B-SFT-RDPO The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 7 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-princeton-nlp-Llama-3-Base-8B-SFT-RDPO-private.tabular100K<n<1M0 likes697 downloads2y agoHugging Face22geodesic-research /pa-warm-start-sft-xl-calibrationtabular100K<n<1M0 likes697 downloads10d agoHugging Face23openeurollm /Dolci-Think-SFT-translated Dolci-Think-SFT-translated Machine translations of the Dolci-Think-SFT-32B dataset, produced with gemma-4-31B-it. The samples selected for translation are those where content_quality == "excellent" according to the propella annotations. Columns Each row is a translated conversation plus the result of a post-translation quality filter: id — source record id. messages — the translated conversation (list of {content, role}). filter_pass — true if the row passed… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/Dolci-Think-SFT-translated.tabulartext-generation1M<n<10M0 likes687 downloads8d agoHugging Face24cnut1648 /LLM-fingerprinted-SFTtabularn<1K0 likes641 downloads3y agoHugging Face25soketlabs /bhasha-sft Bhasha SFT Bhasha SFT is a massive collection of multiple open sourced Supervised Fine-Tuning datasets for training Multilingual Large Language Models. The dataset contains collation of over 13 million instances of instruction-response data for 3 Indian languages (Hindi, Gujarati, Bengali) and English having both human annotated and synthetic data. Curated by: Soket AI Labs Language(s) (NLP): [English, Hindi, Bengali, Gujarati] License: [cc-by-4.0, apache-2.0, mit]… See the full description on the dataset page: https://huggingface.co/datasets/soketlabs/bhasha-sft.tabularquestion-answering10M<n<100M4 likes636 downloads2y agoHugging Face26toroe /Soofi-Think-SFT-10B-multilingual ReasonXL: A Multilingual Cross-Domain Reasoning Corpus ReasonXL is a large-scale multilingual reasoning corpus spanning 5 languages and ~44B tokens in total. It is designed to support supervised fine-tuning of reasoning models with in-language chain-of-thought traces across diverse technical domains. Data Generation English source samples were drawn from 10 existing reasoning datasets, filtered and quality-annotated using ellamind/propella-1-4b, and then translated into… See the full description on the dataset page: https://huggingface.co/datasets/toroe/Soofi-Think-SFT-10B-multilingual.tabular10M<n<100M0 likes537 downloads6mo agoHugging Face27Manusagents /Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2 🧬 Omni-Frontier Collection Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible. 📖 Jump to What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2.tabulartext-generation10M<n<100M0 likes517 downloads24d agoHugging Face28geodesic-research /pa-warm-start-sft-heavy-25b-mix-longtabular1M<n<10M0 likes508 downloads16d agoHugging Face29geodesic-research /pa-warm-start-sft-xl-1b-smoketabular1K<n<10K0 likes491 downloads11d agoHugging Face30tts-sft /round2-oss-matched round2-oss-matched — 第二轮 4 组实验数据(每组 10 节点,共 40) 代码:repo 分支 claude/round2-matched-compute(先 git fetch origin && git merge origin/claude/round2-matched-compute)。 目录: exp0_20b/node00..04/pool.jsonl # 实验 0:shard-05 修复重跑(20B) exp0_120b/node00..04/pool.jsonl # 实验 0:同上(120B) exp1_20b/node00..09/{seeds,budgets,pool}.jsonl # 实验 1:20B token 对齐独立采样 exp2_120b/node00..09/{seeds,budgets,pool}.jsonl # 实验 2:120B 同上 exp3_120b/node00..09/{ck_nonsat/,nonsat_seeds,budgets,pool… See the full description on the dataset page: https://huggingface.co/datasets/tts-sft/round2-oss-matched.tabularn<1K0 likes455 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.