CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01UCB-team /unclickbait-synthetic-27b-trajectories Unclickbait Synthetic 27B Trajectories Synthetic trajectory dataset generated by using rich structured JSON prompts and validated by two-stage judging pipeline. Contents : Full generated trajectories (current snapshot: 48,623 records out of 152,369 pristine event candidates). : 30 benchmark test samples audited end-to-end through the 122B two-stage judge (Stage 1 integrity gate + Stage 2 4D scoring). texttext-generationn<1K0 likes2.4k downloads10d agoHugging Face02MachineLearningLM /machinelearninglm-scm-synthetic-tabularml MachineLearningLM Pretraining Corpus This repository contains the pretraining corpus for MachineLearningLM, a framework designed to equip large language models (LLMs) with robust in-context machine learning (ML) capabilities. The dataset consists of ML tasks synthesized from millions of structural causal models (SCMs), spanning various shot counts up to 1,024. It is designed to enable LLMs to learn from many in-context examples on standard ML tasks purely via in-context learning… See the full description on the dataset page: https://huggingface.co/datasets/MachineLearningLM/machinelearninglm-scm-synthetic-tabularml.texttext-generation1M<n<10M4 likes857 downloads10mo agoHugging Face03devanshamin /synthetic-pii-function-calling Dataset Summary A function calling dataset created by filtering the urchade/synthetic-pii-ner-mistral-v1 dataset. texttext-generation1K<n<10K0 likes546 downloads2y agoHugging Face04aaaaliou /pi-synthetic Coding agent session traces for aaaaliou/pi-synthetic This dataset contains redacted coding agent session traces collected while working on git@github.com:aliou/pi-synthetic.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review. Data description Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line is a structured… See the full description on the dataset page: https://huggingface.co/datasets/aaaaliou/pi-synthetic.tabulartext-generationn<1K0 likes498 downloads5mo agoHugging Face05zachnorton03 /synthetic-pre1930-sftTL;DR A vintage finetuning dataset (~416k rows, eleven task routes). Sourced by taking excerpts from pre-1930's texts, turning these into verbatim answers, and then using deepseek-chat to generate period-appropriate questions of those answers. Any model tuned on this dataset should, theoretically, never update its weights on anachronistic text, since questions are masked in the finetuning stages. Features composition, verse, narrative, reasoning, multiturn dialogue, and calibrated uncertainty… See the full description on the dataset page: https://huggingface.co/datasets/zachnorton03/synthetic-pre1930-sft.tabularquestion-answering100K<n<1M2 likes333 downloads1mo agoHugging Face06ContextReq /Synthetic-Dataset-Childrens-Stories**Status: released 13-09-2026, repacked 14-09-2026.** The 14-09-2026 repack replaced 58 items after the acceptance gates were strengthened (prompt-instruction leaks, markdown bullet lists and blockquotes); the other 29,942 are unchanged. Development stopped, pipeline released 17/09/26. SAMPLE RELEASE: 30,000 synthetic children's short stories for early-reader language modelling. Metrics Value genres 26 stories per genre 1.153-1.154K stories total characters 38… See the full description on the dataset page: https://huggingface.co/datasets/ContextReq/Synthetic-Dataset-Childrens-Stories.texttext-generation10K<n<100K1 likes210 downloads8d agoHugging Face07Hodfa71 /pstu-synthetic-secrets PSTU Synthetic Secrets Dataset Synthetic secrets benchmark for evaluating LLM memorization and unlearning, from the paper: Not All Secrets Are Equal: Type-Aware Unlearning for Language Model Secret Removal Hoda Fakhar — ECML PKDD 2026 Dataset Description 175 synthetic secrets across 25 types, each paired with 100 structurally similar decoys for computing the Carlini exposure metric. All data is synthetically generated. No real credentials, PII, or sensitive information… See the full description on the dataset page: https://huggingface.co/datasets/Hodfa71/pstu-synthetic-secrets.texttext-generationn<1K0 likes201 downloads6mo agoHugging Face08amazon /Turnstile-Synthetic-Domains Data Turnstile — Synthetic Domains A large-scale synthetic dataset of function-calling interactions with chain-of-thought reasoning traces, designed for training small language models on tool-use tasks. Dataset Summary Metric Value Interactions 100,262 Unique APIs 1,025 Distractors per interaction 5 Template types 17 Avg roles per interaction ~10 Avg tokens per interaction ~972 Language English Generator model Qwen2.5-32B-Instruct… See the full description on the dataset page: https://huggingface.co/datasets/amazon/Turnstile-Synthetic-Domains.texttext-generation100K<n<1M0 likes158 downloads2mo agoHugging Face09greghavens /fable-5-coding-and-debugging-traces-synthetic-corrections Model Synthetic Corrections 1 TRAJECTORIES · 2 TRAINING ROWS · 16 kB Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Synthetic Corrections companion dataset. The original dataset is greghavens/fable-5-coding-and-debugging-traces. These are narrowly, synthetically corrected, independently re-judged traces that never passed in the original dataset. Behavior-preserving instruction-following… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/fable-5-coding-and-debugging-traces-synthetic-corrections.tabulartext-generationn<1K0 likes138 downloads2mo agoHugging Face10Heydaritoday /Persian-Synthetic-Instruct Persian Synthetic Instruct High-quality Persian instruction-following dataset generated using LLMs. 4,000+ instruction-response pairs 51 domains Generated by gpt-4.1-mini and gpt-4.1-nano texttext-generation1K<n<10K0 likes105 downloads3mo agoHugging Face11liodon-ai /math-dow-mod-synthetic-v1 Math + Cyclic-Time Synthetic Dataset Synthetic dataset for training a small (~10M-100M param), task-specialized LLM on arithmetic (addition, multiplication), cyclic time arithmetic (days-of-week, months, 12-hour and 24-hour clock), and mod-k remainder probes — generalization-focused rather than memorization, following on from the T2/T5/T10 modular-circuit discussion. Days-of-week, months, and hours are all instances of the same underlying cyclic/modular-addition structure… See the full description on the dataset page: https://huggingface.co/datasets/liodon-ai/math-dow-mod-synthetic-v1.texttext-generation10K<n<100K0 likes90 downloads2mo agoHugging Face12himanshunakrani9 /mimo-coding-synthetic-5k MiMo Coding Synthetic 5.4K MiMo Coding Synthetic 5.4K is a purely synthetic coding instruction dataset generated with Xiaomi MiMo mimo-v2.5-pro. It contains 5,411 validated examples across programming languages, coding task types, and difficulty levels. The dataset is provided in two formats: A canonical rich JSONL format with metadata and labels. An OpenAI chat messages JSONL format for supervised fine-tuning pipelines. The generation run used 20 parallel workers for roughly… See the full description on the dataset page: https://huggingface.co/datasets/himanshunakrani9/mimo-coding-synthetic-5k.texttext-generation10K<n<100K0 likes88 downloads3mo agoHugging Face13SINAI /ALIA-es-legal-administrative-synthetic-instructions Dataset Introduction The ALIA Spanish Legal and Administrative Synthetic Instructions Corpus is a synthetic instruction-tuning resource in Spanish created under the ALIA project using the Magpie methodology. It was designed to train and evaluate language models in legal and administrative tasks with controlled formats and large-scale supervision. It contains: 763,804 instances 534,112,398 tokens 16 task modalities (questions, instructions, multiple-choice, true/false; with and… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-legal-administrative-synthetic-instructions.texttext-generation100K<n<1M1 likes86 downloads3mo agoHugging Face14SINAI /ALIA-es-biomedical-synthetic-instructions Dataset Introduction The ALIA Spanish Biomedical Synthetic Instructions Corpus is a synthetic instruction-tuning resource in Spanish created under the ALIA project using the Magpie methodology. It was designed to train and evaluate language models in biomedical and healthcare tasks with controlled formats and large-scale supervision. It contains: 639,456 instances 961,073,205 tokens 14 task modalities (clinical diagnosis, patient education, ethical reasoning, document… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-biomedical-synthetic-instructions.texttext-generation100K<n<1M0 likes84 downloads4mo agoHugging Face15liodon-ai /math-dow-mod-synthetic-v1-base6 Math + Cyclic-Time Synthetic Dataset Synthetic dataset for training a small (~10M-100M param), task-specialized LLM on arithmetic (addition, multiplication), cyclic time arithmetic (days-of-week, months, 12-hour and 24-hour clock), and mod-k remainder probes — generalization-focused rather than memorization, following on from the T2/T5/T10 modular-circuit discussion. Days-of-week, months, and hours are all instances of the same underlying cyclic/modular-addition structure… See the full description on the dataset page: https://huggingface.co/datasets/liodon-ai/math-dow-mod-synthetic-v1-base6.texttext-generation10K<n<100K0 likes81 downloads2mo agoHugging Face16liodon-ai /math-dow-mod-synthetic-v1-base7 Math + Cyclic-Time Synthetic Dataset Synthetic dataset for training a small (~10M-100M param), task-specialized LLM on arithmetic (addition, multiplication), cyclic time arithmetic (days-of-week, months, 12-hour and 24-hour clock), and mod-k remainder probes — generalization-focused rather than memorization, following on from the T2/T5/T10 modular-circuit discussion. Days-of-week, months, and hours are all instances of the same underlying cyclic/modular-addition structure… See the full description on the dataset page: https://huggingface.co/datasets/liodon-ai/math-dow-mod-synthetic-v1-base7.texttext-generation10K<n<100K0 likes80 downloads2mo agoHugging Face17hllzmz /synthetic-mental-health-convos Synthetic Mental Health SFT Dataset Dataset Summary This dataset contains high-fidelity, synthetic patient-therapist dialogues designed for Supervised Fine-Tuning (SFT) of Large Language Models (LLMs) in the domain of mental health. The primary goal of this dataset is to train AI assistants to transition from "general knowledge" models to empathetic, supportive, and safety-conscious mental health companions. The dialogues cover a wide spectrum of mental health conditions… See the full description on the dataset page: https://huggingface.co/datasets/hllzmz/synthetic-mental-health-convos.texttext-generation1K<n<10K1 likes79 downloads10mo agoHugging Face18SINAI /ALIA-es-cultural-heritage-synthetic-instructions Dataset Introduction The ALIA Spanish Cultural and Heritage Synthetic Instructions Corpus is a synthetic instruction-tuning resource in Spanish created under the ALIA project using the Magpie methodology. It was designed to train and evaluate language models in cultural heritage, digital humanities, and historical knowledge tasks with natural linguistic variation and large-scale supervision. It contains: 748,480 instances 629,682,398 tokens 25 task modalities (heritage QA… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-cultural-heritage-synthetic-instructions.texttext-generation100K<n<1M0 likes79 downloads4mo agoHugging Face19BertilBraun /voice-light-tool-use-synthetic Voice Light Teacher-Led Tool-Use Synthetic This repository contains the current canonical synthetic source dataset for Voice Light's conversational tool-use fine-tuning. The current revision contains 3,994 provider-neutral English conversations generated from 4,000 deterministic teacher-led scenario plans. Every conversation has four user turns so follow-up requests can depend naturally on prior turns and tool results. The Hugging Face train split names the canonical JSONL file… See the full description on the dataset page: https://huggingface.co/datasets/BertilBraun/voice-light-tool-use-synthetic.texttext-generation1K<n<10K0 likes77 downloads2mo agoHugging Face20pierjoe /function-calling-synthetic-2000 Synthetic Multi-Turn Function-Calling Conversations Synthetic, multi-turn function-calling (tool-use) conversations for fine-tuning and evaluating LLMs. Generated and validated with synthfc — the open-source pipeline (sampler, prompt builder, validator, post-processor, web viewer) lives in that GitHub repo. A strong teacher LLM (Qwen/Qwen3.6-35B-A3B) produces each conversation from controlled, sampled parameters, so the dataset is diverse along many axes (call type, languages… See the full description on the dataset page: https://huggingface.co/datasets/pierjoe/function-calling-synthetic-2000.texttext-generation1K<n<10K0 likes75 downloads4mo agoHugging Face21foxycuter /column-arithmetic-ru-synthetic Column Arithmetic RU Dataset Синтетический датасет для обучения модели сложению и вычитанию в столбик. Splits train.jsonl: основное обучение eval.jsonl: holdout-оценка hard.jsonl: трудные случаи с длинными переносами и займами Hard cases included 9999+1 10000+9999 9090+1010 55555+55555 10999+2 1234+8766 1000-7 10000-9999 50005-49999 8000-1 10101-909 100000-1 99009+991 12000-3456 700000+300001 1002003-998877 Current release status… See the full description on the dataset page: https://huggingface.co/datasets/foxycuter/column-arithmetic-ru-synthetic.tabulartext-generation1K<n<10K1 likes73 downloads5mo agoHugging Face22sxiong /synthetic-math Synthetic MATH Dataset Dataset Summary This dataset contains symthetic math problems generated with GPT-4o and verified with DeepSeek-R1, intended to augment the MATH dataset (Hendrycks et al., 2021) with additional training/evaluation examples. Only problems where R1's final answer matched the reference answer given by GPT-4o are included. Each row bundles the problem, the reference solution, and R1's full reasoning trajectory used for verification. Note: This is… See the full description on the dataset page: https://huggingface.co/datasets/sxiong/synthetic-math.texttext-generation10K<n<100K1 likes71 downloads3mo agoHugging Face23cturan /turkish-synthetic-corpus Turkish Synthetic Corpus A synthetic Turkish text corpus with 1,871,131 documents, designed for Turkish language model training. About Inspired by HuggingFaceTB/smollm-corpus. Questions and prompts were sourced from the SmolLM Corpus pipeline; a language model then generated localized Turkish responses and documents around them. All credit for the original corpus design and methodology goes to the HuggingFace SmolLM team. The resulting dataset covers a wide range of… See the full description on the dataset page: https://huggingface.co/datasets/cturan/turkish-synthetic-corpus.texttext-generation1M<n<10M0 likes69 downloads6mo agoHugging Face24sabber /snh-loan-adjudication-synthetic SNH Synthetic Loan Adjudication Dialogues This dataset was created for a technical coding challenge about conversational personal-loan adjudication. Every record contains supplied policy rules, a synthetic dialogue, deterministic ground-truth fields, a decision, failed rule IDs, and a customer-facing explanation. No record represents a real person, and the dataset contains no real applicant information. Splits File Records Purpose train.jsonl 4,000… See the full description on the dataset page: https://huggingface.co/datasets/sabber/snh-loan-adjudication-synthetic.texttext-generation1K<n<10K0 likes69 downloads21d agoHugging Face25SP4ND4N /Apple-Synthetic Dataset Summary A synthetic question–answer dataset grounded in a curated set of seed documents that reflect Apple's historical design philosophy and cultural principles. Questions are interpretive and scenario-based; answers are required to be derivable from the provided source text [Apple-legacy-corpus](https://huggingface.co/datasets/SP4ND4N/Apple-legacy-corpus. Source type: synthetic, generated from internal seed docs (not scraped Apple manuals) Format: JSON Lines (JSONL) with… See the full description on the dataset page: https://huggingface.co/datasets/SP4ND4N/Apple-Synthetic.texttext-generation10K<n<100K0 likes68 downloads1y agoHugging Face26Aratako /Synthetic-Japanese-Roleplay-NSFW-Claude-3.5s-15.3k-formatted Synthetic-Japanese-Roleplay-NSFW-Claude-3.5s-15.3k-formatted 20240907 データ増量(約10500件→約15300件) 概要 Claude 3.5 Sonnetを用いて作成した日本語ロールプレイデータセットであるAratako/Synthetic-Japanese-Roleplay-NSFW-Claude-3.5s-15.3kにsystem messageを追加して整形したデータセットです。 データの詳細については元データセットのREADMEを参照してください。 ライセンス CC-BY-NC-SA 4.0の元配布します。 また、Anthropicの利用規約に記載のある通り、このデータを使ってAnthropicのサービスやモデルと競合するようなモデルを開発することは禁止されています。 texttext-generation10K<n<100K18 likes65 downloads2y agoHugging Face27parinzee /seed-free-synthetic-instruct-thai-v1 Seed-Free Synthetic Instruct Thai v1 (F+C+D+) This dataset is part of the research paper "Seed-Free Synthetic Data Generation Framework for Instruction-Tuning LLMs: A Case Study in Thai" submitted to ACL SRW 2024. It represents the best-performing synthetic dataset (F+C+D+) generated using our novel seed-free framework for low-resource languages, specifically Thai. Dataset Details Size: 5,000 instructions Language: Thai Task: Instruction-tuning for Large Language Models… See the full description on the dataset page: https://huggingface.co/datasets/parinzee/seed-free-synthetic-instruct-thai-v1.texttext-generation1K<n<10K3 likes64 downloads1y agoHugging Face28mi-obra /miobra-synthetic-construction-material-extraction-v1 Mi Obra Synthetic Construction Material Extraction v1.0 English Dataset Description This dataset contains 10,000 synthetic Spanish construction-material titles paired with structured entity-extraction targets. It is designed for Argentine construction terminology and controlled experiments in fine-tuning, teaching, information extraction, and structured generation. Language: Spanish (es) Regional context: Argentina Rows: 10,000 License: CC BY 4.0… See the full description on the dataset page: https://huggingface.co/datasets/mi-obra/miobra-synthetic-construction-material-extraction-v1.texttext-generation10K<n<100K0 likes64 downloads27d agoHugging Face29croqaz /Synthetic-archive The Synthetic Archive Synthetic Archive is a large synthetic English-text dataset generated from OCR-derived historical and period-style passages. Knowledge cutoff is year 1900. Each source passage was divided into chunks and processed through several generation tasks, including: generating continuations of unfinished passages; creating question-and-answer pairs; extracting and reformulating factual knowledge; rewriting material as a narrative; transforming source material into… See the full description on the dataset page: https://huggingface.co/datasets/croqaz/Synthetic-archive.texttext-generation10M<n<100M0 likes63 downloads1mo agoHugging Face30marccgrau /agentic_synthetic_aggressive_conversations_en Simulated Aggressive Customer Service Conversations Dataset Overview This dataset contains aggressive customer service conversations generated by an agentic simulation system. Each record is stored in JSON Lines (JSONL) format and includes: Scenario Metadata: Selected bank, customer, agent profiles, and task details. Conversation Messages: Full message history between the customer and service agent. Summary: A German summary of the conversation. Cost Metrics: API cost… See the full description on the dataset page: https://huggingface.co/datasets/marccgrau/agentic_synthetic_aggressive_conversations_en.texttext-generationn<1K0 likes61 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.