CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01UCB-team /unclickbait-synthetic-27b-trajectories Unclickbait Synthetic 27B Trajectories Synthetic trajectory dataset generated by using rich structured JSON prompts and validated by two-stage judging pipeline. Contents : Full generated trajectories (current snapshot: 48,623 records out of 152,369 pristine event candidates). : 30 benchmark test samples audited end-to-end through the 122B two-stage judge (Stage 1 integrity gate + Stage 2 4D scoring). texttext-generationn<1K0 likes2.7k downloads12d agoHugging Face02MachineLearningLM /machinelearninglm-scm-synthetic-tabularml MachineLearningLM Pretraining Corpus This repository contains the pretraining corpus for MachineLearningLM, a framework designed to equip large language models (LLMs) with robust in-context machine learning (ML) capabilities. The dataset consists of ML tasks synthesized from millions of structural causal models (SCMs), spanning various shot counts up to 1,024. It is designed to enable LLMs to learn from many in-context examples on standard ML tasks purely via in-context learning… See the full description on the dataset page: https://huggingface.co/datasets/MachineLearningLM/machinelearninglm-scm-synthetic-tabularml.texttext-generation1M<n<10M4 likes814 downloads10mo agoHugging Face03devanshamin /synthetic-pii-function-calling Dataset Summary A function calling dataset created by filtering the urchade/synthetic-pii-ner-mistral-v1 dataset. texttext-generation1K<n<10K0 likes531 downloads2y agoHugging Face04aaaaliou /pi-synthetic Coding agent session traces for aaaaliou/pi-synthetic This dataset contains redacted coding agent session traces collected while working on git@github.com:aliou/pi-synthetic.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review. Data description Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line is a structured… See the full description on the dataset page: https://huggingface.co/datasets/aaaaliou/pi-synthetic.tabulartext-generationn<1K0 likes498 downloads5mo agoHugging Face05zachnorton03 /synthetic-pre1930-sftTL;DR A vintage finetuning dataset (~416k rows, eleven task routes). Sourced by taking excerpts from pre-1930's texts, turning these into verbatim answers, and then using deepseek-chat to generate period-appropriate questions of those answers. Any model tuned on this dataset should, theoretically, never update its weights on anachronistic text, since questions are masked in the finetuning stages. Features composition, verse, narrative, reasoning, multiturn dialogue, and calibrated uncertainty… See the full description on the dataset page: https://huggingface.co/datasets/zachnorton03/synthetic-pre1930-sft.tabularquestion-answering100K<n<1M2 likes313 downloads1mo agoHugging Face06ContextReq /Synthetic-Dataset-Childrens-Stories**Status: released 13-09-2026, repacked 14-09-2026.** The 14-09-2026 repack replaced 58 items after the acceptance gates were strengthened (prompt-instruction leaks, markdown bullet lists and blockquotes); the other 29,942 are unchanged. Development stopped, pipeline released 17/09/26. SAMPLE RELEASE: 30,000 synthetic children's short stories for early-reader language modelling. Metrics Value genres 26 stories per genre 1.153-1.154K stories total characters 38… See the full description on the dataset page: https://huggingface.co/datasets/ContextReq/Synthetic-Dataset-Childrens-Stories.texttext-generation10K<n<100K1 likes214 downloads10d agoHugging Face07CohereLabs /fusion-synth-data-geofactx Offline Synthetic Data (GeoFactX) for: Making, not taking, the Best-of-N Content This data contains completions for the GeoFactX training split prompts from 5 different teacher models and 2 aggregations: Teachers: We sample one completion from each of the following models at temperature T=0.3. For kimik2, qwen3, and deepseek-v3 we use TogetherAI, for gemma3-27b and command-a we use locally hosted images. gemma3-27b: GEMMA3-27B-IT kimik2: KIMI-K2-INSTRUCT qwen3:… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/fusion-synth-data-geofactx.texttext-generation1K<n<10K1 likes171 downloads1y agoHugging Face08CohereLabs /fusion-synth-data-s1kx Offline Synthetic Data (s1K-X) for: Making, not taking, the Best-of-N Content This data contains completions for the s1K-X training split prompts from 5 different teacher models and 2 aggregations: Teachers: We sample one completion from each of the following models at temperature T=0.3. For kimik2, qwen3, and deepseek-v3 we use TogetherAI, for gemma3-27b and command-a we use locally hosted images. gemma3-27b: GEMMA3-27B-IT kimik2: KIMI-K2-INSTRUCT qwen3: QWEN3-235B… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/fusion-synth-data-s1kx.texttext-generation10K<n<100K1 likes170 downloads1y agoHugging Face09CohereLabs /fusion-synth-data-ufb Offline Synthetic Data (UFB) for: Making, not taking, the Best-of-N Content This data contains completions for a 10,000 subset of the UFB prompts (translated into 9 languages) from 5 different teacher models and 2 aggregations: Teachers: We sample one completion from each of the following models at temperature T=0.3. For kimik2, qwen3, and deepseek-v3 we use TogetherAI, for gemma3-27b and command-a we use locally hosted images. gemma3-27b: GEMMA3-27B-IT kimik2:… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/fusion-synth-data-ufb.texttext-generation10K<n<100K2 likes164 downloads1y agoHugging Face10amazon /Turnstile-Synthetic-Domains Data Turnstile — Synthetic Domains A large-scale synthetic dataset of function-calling interactions with chain-of-thought reasoning traces, designed for training small language models on tool-use tasks. Dataset Summary Metric Value Interactions 100,262 Unique APIs 1,025 Distractors per interaction 5 Template types 17 Avg roles per interaction ~10 Avg tokens per interaction ~972 Language English Generator model Qwen2.5-32B-Instruct… See the full description on the dataset page: https://huggingface.co/datasets/amazon/Turnstile-Synthetic-Domains.texttext-generation100K<n<1M0 likes143 downloads2mo agoHugging Face11greghavens /fable-5-coding-and-debugging-traces-synthetic-corrections Model Synthetic Corrections 1 TRAJECTORIES · 2 TRAINING ROWS · 16 kB Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Synthetic Corrections companion dataset. The original dataset is greghavens/fable-5-coding-and-debugging-traces. These are narrowly, synthetically corrected, independently re-judged traces that never passed in the original dataset. Behavior-preserving instruction-following… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/fable-5-coding-and-debugging-traces-synthetic-corrections.tabulartext-generationn<1K0 likes142 downloads2mo agoHugging Face12Hodfa71 /pstu-synthetic-secrets PSTU Synthetic Secrets Dataset Synthetic secrets benchmark for evaluating LLM memorization and unlearning, from the paper: Not All Secrets Are Equal: Type-Aware Unlearning for Language Model Secret Removal Hoda Fakhar — ECML PKDD 2026 Dataset Description 175 synthetic secrets across 25 types, each paired with 100 structurally similar decoys for computing the Carlini exposure metric. All data is synthetically generated. No real credentials, PII, or sensitive information… See the full description on the dataset page: https://huggingface.co/datasets/Hodfa71/pstu-synthetic-secrets.texttext-generationn<1K0 likes141 downloads6mo agoHugging Face13dendriteholdings /Dendrite-Synth-Multi-Domain Dendrite Synth Multi-Domain A verified, difficulty-filtered, style-amplified synthetic corpus of question / reasoning / answer triples spanning the 14 MMLU-Pro categories - mathematics, computer science, natural sciences, chemistry, physics, engineering, health, law, business, economics, psychology, philosophy, history and others (expanded to 122 fine-grained categories and 695 subcategories). Problems are written by a pool of generator models, solved with explicit reasoning by… See the full description on the dataset page: https://huggingface.co/datasets/dendriteholdings/Dendrite-Synth-Multi-Domain.texttext-generation100K<n<1M0 likes132 downloads24d agoHugging Face14SINAI /ALIA-es-biomedical-synthetic-instructions Dataset Introduction The ALIA Spanish Biomedical Synthetic Instructions Corpus is a synthetic instruction-tuning resource in Spanish created under the ALIA project using the Magpie methodology. It was designed to train and evaluate language models in biomedical and healthcare tasks with controlled formats and large-scale supervision. It contains: 639,456 instances 961,073,205 tokens 14 task modalities (clinical diagnosis, patient education, ethical reasoning, document… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-biomedical-synthetic-instructions.texttext-generation100K<n<1M0 likes93 downloads4mo agoHugging Face15samwell /synthea-ncd-instructions Synthea NCD Instructions Synthetic EHR-based instruction-tuning dataset for training LLMs to predict non-communicable disease (NCD) risk, specifically Type 2 Diabetes and Hypertension. Quick Start from datasets import load_dataset dataset = load_dataset("samwell/synthea-ncd-instructions") # View a sample print(dataset["train"][0]) Dataset Description This dataset contains instruction-tuning examples derived from synthetic patient records generated using… See the full description on the dataset page: https://huggingface.co/datasets/samwell/synthea-ncd-instructions.texttext-generation10K<n<100K0 likes90 downloads6mo agoHugging Face16himanshunakrani9 /mimo-coding-synthetic-5k MiMo Coding Synthetic 5.4K MiMo Coding Synthetic 5.4K is a purely synthetic coding instruction dataset generated with Xiaomi MiMo mimo-v2.5-pro. It contains 5,411 validated examples across programming languages, coding task types, and difficulty levels. The dataset is provided in two formats: A canonical rich JSONL format with metadata and labels. An OpenAI chat messages JSONL format for supervised fine-tuning pipelines. The generation run used 20 parallel workers for roughly… See the full description on the dataset page: https://huggingface.co/datasets/himanshunakrani9/mimo-coding-synthetic-5k.texttext-generation10K<n<100K0 likes87 downloads3mo agoHugging Face17xlr8harder /synthid-qwen3-4b-instruct-2507-wildchat Qwen3-4B SynthID three-arm corpus This export contains aligned unwatermarked, SynthID key-A, and SynthID key-B responses from Qwen/Qwen3-4B-Instruct-2507. Matched splits share prompts and request seeds across configurations; unmatched splits use mutually disjoint prompt pools. Export complete for its source work queue: true. Generation profile Model revision: cdbee75f17c01a7cc42f958dc650907174af0554 Native model dtype: bfloat16 Maximum generated tokens: 4096… See the full description on the dataset page: https://huggingface.co/datasets/xlr8harder/synthid-qwen3-4b-instruct-2507-wildchat.tabulartext-generation100K<n<1M0 likes87 downloads1mo agoHugging Face18SINAI /ALIA-es-legal-administrative-synthetic-instructions Dataset Introduction The ALIA Spanish Legal and Administrative Synthetic Instructions Corpus is a synthetic instruction-tuning resource in Spanish created under the ALIA project using the Magpie methodology. It was designed to train and evaluate language models in legal and administrative tasks with controlled formats and large-scale supervision. It contains: 763,804 instances 534,112,398 tokens 16 task modalities (questions, instructions, multiple-choice, true/false; with and… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-legal-administrative-synthetic-instructions.texttext-generation100K<n<1M1 likes85 downloads3mo agoHugging Face19hllzmz /synthetic-mental-health-convos Synthetic Mental Health SFT Dataset Dataset Summary This dataset contains high-fidelity, synthetic patient-therapist dialogues designed for Supervised Fine-Tuning (SFT) of Large Language Models (LLMs) in the domain of mental health. The primary goal of this dataset is to train AI assistants to transition from "general knowledge" models to empathetic, supportive, and safety-conscious mental health companions. The dialogues cover a wide spectrum of mental health conditions… See the full description on the dataset page: https://huggingface.co/datasets/hllzmz/synthetic-mental-health-convos.texttext-generation1K<n<10K1 likes82 downloads10mo agoHugging Face20bluecolor777 /Dendrite-Synth-Multi-Domain Dendrite Synth Multi-Domain A verified, difficulty-filtered, style-amplified synthetic corpus of question / reasoning / answer triples spanning the 14 MMLU-Pro categories - mathematics, computer science, natural sciences, chemistry, physics, engineering, health, law, business, economics, psychology, philosophy, history and others (expanded to 122 fine-grained categories and 695 subcategories). Problems are written by a pool of generator models, solved with explicit reasoning by… See the full description on the dataset page: https://huggingface.co/datasets/bluecolor777/Dendrite-Synth-Multi-Domain.texttext-generation100K<n<1M0 likes82 downloads16d agoHugging Face21Heydaritoday /Persian-Synthetic-Instruct Persian Synthetic Instruct High-quality Persian instruction-following dataset generated using LLMs. 4,000+ instruction-response pairs 51 domains Generated by gpt-4.1-mini and gpt-4.1-nano texttext-generation1K<n<10K0 likes81 downloads4mo agoHugging Face22cturan /turkish-synthetic-corpus Turkish Synthetic Corpus A synthetic Turkish text corpus with 1,871,131 documents, designed for Turkish language model training. About Inspired by HuggingFaceTB/smollm-corpus. Questions and prompts were sourced from the SmolLM Corpus pipeline; a language model then generated localized Turkish responses and documents around them. All credit for the original corpus design and methodology goes to the HuggingFace SmolLM team. The resulting dataset covers a wide range of… See the full description on the dataset page: https://huggingface.co/datasets/cturan/turkish-synthetic-corpus.texttext-generation1M<n<10M0 likes78 downloads6mo agoHugging Face23BertilBraun /voice-light-tool-use-synthetic Voice Light Teacher-Led Tool-Use Synthetic This repository contains the current canonical synthetic source dataset for Voice Light's conversational tool-use fine-tuning. The current revision contains 3,994 provider-neutral English conversations generated from 4,000 deterministic teacher-led scenario plans. Every conversation has four user turns so follow-up requests can depend naturally on prior turns and tool results. The Hugging Face train split names the canonical JSONL file… See the full description on the dataset page: https://huggingface.co/datasets/BertilBraun/voice-light-tool-use-synthetic.texttext-generation1K<n<10K0 likes78 downloads2mo agoHugging Face24SINAI /ALIA-es-cultural-heritage-synthetic-instructions Dataset Introduction The ALIA Spanish Cultural and Heritage Synthetic Instructions Corpus is a synthetic instruction-tuning resource in Spanish created under the ALIA project using the Magpie methodology. It was designed to train and evaluate language models in cultural heritage, digital humanities, and historical knowledge tasks with natural linguistic variation and large-scale supervision. It contains: 748,480 instances 629,682,398 tokens 25 task modalities (heritage QA… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-cultural-heritage-synthetic-instructions.texttext-generation100K<n<1M0 likes77 downloads4mo agoHugging Face25foxycuter /column-arithmetic-ru-synthetic Column Arithmetic RU Dataset Синтетический датасет для обучения модели сложению и вычитанию в столбик. Splits train.jsonl: основное обучение eval.jsonl: holdout-оценка hard.jsonl: трудные случаи с длинными переносами и займами Hard cases included 9999+1 10000+9999 9090+1010 55555+55555 10999+2 1234+8766 1000-7 10000-9999 50005-49999 8000-1 10101-909 100000-1 99009+991 12000-3456 700000+300001 1002003-998877 Current release status… See the full description on the dataset page: https://huggingface.co/datasets/foxycuter/column-arithmetic-ru-synthetic.tabulartext-generation1K<n<10K1 likes75 downloads5mo agoHugging Face26codelion /gsm8k-synth GSM8K-Synth 117,955 grade-school math word problems in the style of GSM8K, LLM-generated (Claude and Gemini) as training data for small math-word-problem models. Every problem is round-trip validated (its program re-executes to the stated answer) and decontaminated against the GSM8K test set — 0% 8-gram overlap. Built for and used by codelion/sprog-9m, a 9.37M-parameter LLM-free GSM8K solver. Schema field type description question string the word… See the full description on the dataset page: https://huggingface.co/datasets/codelion/gsm8k-synth.textquestion-answering100K<n<1M2 likes74 downloads4mo agoHugging Face2711-47 /fable-5-coding-and-debugging-traces-synthetic Model Synthetic Corrections 1 TRAJECTORIES · 2 TRAINING ROWS · 16 kB Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Synthetic Corrections companion dataset. The original dataset is greghavens/fable-5-coding-and-debugging-traces. These are narrowly, synthetically corrected, independently re-judged traces that never passed in the original dataset. Behavior-preserving instruction-following… See the full description on the dataset page: https://huggingface.co/datasets/11-47/fable-5-coding-and-debugging-traces-synthetic.tabulartext-generationn<1K0 likes73 downloads10d agoHugging Face28nkthebass /math-synth-400k math-synth — 400k arithmetic problems with exact step-by-step scratchpads Synthetic math SFT data where every answer is provably correct, because nothing was written by a language model — the problems and their worked solutions are generated programmatically in Python, so the label is the computation. Most synthetic math datasets are distilled from an LLM teacher, which means some fraction of the answers are silently wrong and get baked into the student. This set has no teacher… See the full description on the dataset page: https://huggingface.co/datasets/nkthebass/math-synth-400k.texttext-generation100K<n<1M0 likes71 downloads29d agoHugging Face29sabber /snh-loan-adjudication-synthetic SNH Synthetic Loan Adjudication Dialogues This dataset was created for a technical coding challenge about conversational personal-loan adjudication. Every record contains supplied policy rules, a synthetic dialogue, deterministic ground-truth fields, a decision, failed rule IDs, and a customer-facing explanation. No record represents a real person, and the dataset contains no real applicant information. Splits File Records Purpose train.jsonl 4,000… See the full description on the dataset page: https://huggingface.co/datasets/sabber/snh-loan-adjudication-synthetic.texttext-generation1K<n<10K0 likes71 downloads23d agoHugging Face30sxiong /synthetic-math Synthetic MATH Dataset Dataset Summary This dataset contains symthetic math problems generated with GPT-4o and verified with DeepSeek-R1, intended to augment the MATH dataset (Hendrycks et al., 2021) with additional training/evaluation examples. Only problems where R1's final answer matched the reference answer given by GPT-4o are included. Each row bundles the problem, the reference solution, and R1's full reasoning trajectory used for verification. Note: This is… See the full description on the dataset page: https://huggingface.co/datasets/sxiong/synthetic-math.texttext-generation10K<n<100K1 likes70 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.