CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Azzindani /ID_Legal_QA_SynThink 🧠 Indonesian Legal QA Synthetic Think Dataset (ID_Legal_QA_SynThink) This repository features an advanced Synthetic Question-and-Answer dataset for the Indonesian legal domain, distinguished by the inclusion of an explicit Thinking Phase (Chain-of-Thought). 🏛️ 💡 The Concept: Transparent Legal Reasoning Standard QA datasets often provide just the "final answer." This dataset goes deeper by capturing the internal reasoning process of the model before it arrives at a… See the full description on the dataset page: https://huggingface.co/datasets/Azzindani/ID_Legal_QA_SynThink.tabulartext-generation1K<n<10K1 likes2.7k downloads7mo agoHugging Face02richardyoung /synthea-575k-patients Synthea Synthetic Patient Records (575K Patients) A comprehensive synthetic healthcare dataset containing 575,415 patients with complete medical histories, generated using Synthea — the gold standard for synthetic EHR data. No real patient data. Fully synthetic, HIPAA-safe, and ready for ML research and education. Why This Dataset? 575K patients with realistic demographics, conditions, medications, and encounters Privacy-safe: No real PHI — use freely in research… See the full description on the dataset page: https://huggingface.co/datasets/richardyoung/synthea-575k-patients.tabulartext-generation1B<n<10B12 likes678 downloads6mo agoHugging Face03open-athena /recursive-task-synthesis-glm-5.3-rollouts GLM 5.3 agentic rollouts on Recursive-Task-Synthesis This dataset catalogs the full collection made from the pinned Recursive-Task-Synthesis dataset revision be44f96808d5a9b599d5cb024341ff00091adeb7. The repository includes approximately 260.5 GiB of trajectory payload tar shards. Contents at a glance Item Count Source tasks considered 37,284 Source candidates inspected 19,368 Converted tasks after source filters 18,600 Tasks passing gold… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/recursive-task-synthesis-glm-5.3-rollouts.tabulartext-generation100K<n<1M0 likes509 downloads6d agoHugging Face04aaaaliou /pi-synthetic Coding agent session traces for aaaaliou/pi-synthetic This dataset contains redacted coding agent session traces collected while working on git@github.com:aliou/pi-synthetic.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review. Data description Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line is a structured… See the full description on the dataset page: https://huggingface.co/datasets/aaaaliou/pi-synthetic.tabulartext-generationn<1K0 likes498 downloads5mo agoHugging Face05NuBerea /synthesisgated NuBerea/synthesis A cross-corpus synthesis layer for the study of early Jewish and Christian literature. Each config joins pericope-level text units from one corpus — the canonical Bible (Old and New Testament), Second Temple Pseudepigrapha, the Aramaic Targumim, the Nag Hammadi corpus, or Greek and Latin patristic authors — with rhetorical claims extracted from those units and with links into a shared concept vocabulary. The result is a set of per-corpus tables that let a… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/synthesis.tabulartext-generation10K<n<100K0 likes433 downloads2mo agoHugging Face06zachnorton03 /synthetic-pre1930-sftTL;DR A vintage finetuning dataset (~416k rows, eleven task routes). Sourced by taking excerpts from pre-1930's texts, turning these into verbatim answers, and then using deepseek-chat to generate period-appropriate questions of those answers. Any model tuned on this dataset should, theoretically, never update its weights on anachronistic text, since questions are masked in the finetuning stages. Features composition, verse, narrative, reasoning, multiturn dialogue, and calibrated uncertainty… See the full description on the dataset page: https://huggingface.co/datasets/zachnorton03/synthetic-pre1930-sft.tabularquestion-answering100K<n<1M2 likes313 downloads1mo agoHugging Face07Stereotypes-in-LLMs /hiring-bias-mitigation-synthetic-data Hiring-bias mitigation — synthetic training data Semi-synthetic data for training LLMs to make hiring decisions that do not depend on a protected attribute (military status, gender, religion), in English and Ukrainian. Real inputs, synthetic labels. CVs and job descriptions are real, anonymised postings from the Djinni Recruitment Dataset (MIT). Decisions and rationales were written by the teacher model Qwen/Qwen3.5-122B-A10B-GPTQ-Int4. Code and results:… See the full description on the dataset page: https://huggingface.co/datasets/Stereotypes-in-LLMs/hiring-bias-mitigation-synthetic-data.tabulartext-generation100K<n<1M0 likes291 downloads3d agoHugging Face08nlm-dir /MedFact-SynthTo train Med-V1, we construct MedFact-Synth, a large-scale synthetic training set including 1.5 million instances. Each instance contains: a synthetic claim to be verified, a source article serving as evidence, a rationale explaining the verification, and a 5-point Likert-scale verdict, ranging from strong contradiction (-2) and partial contradiction (-1) to neutral (0), partial agreement (+1), and strong agreement (+2). To build this dataset, we begin by sampling one million articles from… See the full description on the dataset page: https://huggingface.co/datasets/nlm-dir/MedFact-Synth.tabulartext-generation1M<n<10M3 likes247 downloads7mo agoHugging Face09KeisukeMiyamoto /SyntheticTalk-jp LambdaTalk-v2 LambdaTalk-v2 is a Japanese synthetic multi-turn conversation dataset generated with Gemma 4 31B. It contains conversations based on seed questions collected from 36 source datasets. Each conversation contains three user-assistant turns. The first user message is the original seed question. The remaining five messages were generated by Gemma 4 31B. Purpose The main purpose of this dataset is supervised fine-tuning of Japanese conversational language… See the full description on the dataset page: https://huggingface.co/datasets/KeisukeMiyamoto/SyntheticTalk-jp.tabulartext-generation1M<n<10M0 likes227 downloads1mo agoHugging Face10fuvty /tau-bench-synthetic tau-bench-synthetic Synthetic tool-use training data for tau-bench, generated using a GT-first task construction pipeline with GLM-5 (via Fireworks API) as the trajectory generator. Overview This dataset was built to train small LLMs (e.g., Qwen3-1.7B) on multi-turn tool-use tasks without using the original tau-bench evaluation set. The pipeline follows a GT-first approach: ground-truth actions are constructed programmatically from the database, then an LLM generates… See the full description on the dataset page: https://huggingface.co/datasets/fuvty/tau-bench-synthetic.tabulartext-generation1K<n<10K3 likes198 downloads6mo agoHugging Face11ibnsina-llm /synthetic-persian-v1 IbnSina Synthetic Persian Corpus v1 Sina Meraji · ORCID 0009-0002-8028-1932 · github.com/sinameraji پیکرهٔ مصنوعی فارسی ابن‌سینا (نسخهٔ ۱) — ۲٫۰۷۵ میلیارد توکن متن آموزشیِ فارسی که از ابتدا به فارسی تولید شده است، نه ترجمه از انگلیسی. این پیکره برای پوشش حوزه‌هایی ساخته شده که وبِ فارسی در آن‌ها کم‌مایه است: توضیح مفاهیم علمی و مهندسی، مسئله‌های حل‌شدهٔ ریاضی و فیزیک، زنجیره‌های استدلال، و متن‌های علمی-پزشکی. هر سند را یک داور خودکار با معیارهای سخت‌گیرانه (درستیِ محاسبه‌ها،… See the full description on the dataset page: https://huggingface.co/datasets/ibnsina-llm/synthetic-persian-v1.tabulartext-generation100K<n<1M2 likes178 downloads24d agoHugging Face12Aratako /Synthetic-JP-EN-Coding-Dataset-801k Synthetic-JP-EN-Coding-Dataset-801k Magpieによって作成したコードSFTデータセットであるAratako/Synthetic-JP-EN-Coding-Dataset-Magpie-69kを元に、Evol-Instructのような手法を用いて複数のinstructionとresonseを生成し拡張して作成した、日英混合801262件のコードSFT用合成データセットです。 日本語: 173849件 英語: 627413件 元のinstructionの作成に利用したモデルは以下の通りです。modelキーに該当レコードの作成に利用したモデル情報があります。 nvidia/Nemotron-4-340B-Instruct microsoft/Phi-3-medium-4k-instruct mistralai/Mixtral-8x22B-Instruct-v0.1… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Synthetic-JP-EN-Coding-Dataset-801k.tabulartext-generation100K<n<1M17 likes168 downloads2y agoHugging Face13KeisukeMiyamoto /SyntheticTextbook-jp SyntheticTextbook-jp SyntheticTextbook-jp is a Japanese synthetic text dataset generated with Gemma 4 26B and Gemma 4 31B. The dataset was created by rewriting noisy source text into textbook-style Japanese for elementary school, junior high school, and high school levels. The rewritten text keeps only general knowledge from the source text. Purpose The main purpose of this dataset is to help LLMs learn natural Japanese text flow. This dataset is designed around… See the full description on the dataset page: https://huggingface.co/datasets/KeisukeMiyamoto/SyntheticTextbook-jp.tabulartext-generation1M<n<10M0 likes152 downloads1mo agoHugging Face14greghavens /fable-5-coding-and-debugging-traces-synthetic-corrections Model Synthetic Corrections 1 TRAJECTORIES · 2 TRAINING ROWS · 16 kB Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Synthetic Corrections companion dataset. The original dataset is greghavens/fable-5-coding-and-debugging-traces. These are narrowly, synthetically corrected, independently re-judged traces that never passed in the original dataset. Behavior-preserving instruction-following… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/fable-5-coding-and-debugging-traces-synthetic-corrections.tabulartext-generationn<1K0 likes142 downloads2mo agoHugging Face15cds-jb /synthweb-gemma4-26b-a4b Gemma-4-26B-A4B FineWeb Rollouts (~580k docs) Open-ended continuations of FineWeb (sample-10BT) document prefixes, generated by google/gemma-4-26b-a4b (the base, non-it Gemma-4 26B-A4B mixture-of-experts model), then mode-collapse filtered. This is the Gemma-4 analogue of cds-jb/qwen3-8b-fineweb-rollouts-100k: a "synthweb" corpus of natural model-generated documents, intended as the substrate for activation-oracle / interpretability probing (extract a base model's residual… See the full description on the dataset page: https://huggingface.co/datasets/cds-jb/synthweb-gemma4-26b-a4b.tabulartext-generation1M<n<10M0 likes141 downloads3mo agoHugging Face16gratex /GNOTHEIA-synthetic-insurance-dataset GNOTHEIA Synthetic Insurance Dataset Published by: Gratex International a.s.Project: InnovAIte — InnovAIte Slovakia License: Apache 2.0Version: 1.0.0Contact: info@gratex.com A synthetic insurance claims dataset designed for AI systems that evaluate insurance claims using OMG SBVR business rules, structured claim polycontexts and synthetic claim-related documents. The dataset main goal is to support: LLM fine-tuning pipeline SBVR reasoning benchmarks insurance claim AI… See the full description on the dataset page: https://huggingface.co/datasets/gratex/GNOTHEIA-synthetic-insurance-dataset.tabulartext-classification1K<n<10K0 likes122 downloads5d agoHugging Face17gretelai /synthetic_multilingual_llm_prompts Image generated by DALL-E. See prompt for more details 📝🌐 Synthetic Multilingual LLM Prompts Welcome to the "Synthetic Multilingual LLM Prompts" dataset! This comprehensive collection features 1,250 synthetic LLM prompts generated using Gretel Navigator, available in seven different languages. To ensure accuracy and diversity in prompts, and translation quality and consistency across the different languages, we employed Gretel Navigator both as a generation tool and as an… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_multilingual_llm_prompts.tabulartext-generation1K<n<10K11 likes115 downloads2y agoHugging Face18mkurman /tulu-3-sft-personas-if-synth Tulu3 Personas Synth Reasoning Synthetic reasoning traces for allenai/tulu-3-sft-personas-instruction-following. Each record contains a persona-based instruction with SYNTH-style reasoning and a generated answer. Dataset Summary 24,747 records (18 dupes + 4,281 incomplete/truncated removed from 29,046 source) 24,747 reasoning turns (99.9% format compliance) Average 1,398 chars per reasoning trace Incomplete records (cut off mid-sentence due to early 1024… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/tulu-3-sft-personas-if-synth.tabulartext-generation10K<n<100K1 likes90 downloads2mo agoHugging Face19xlr8harder /synthid-qwen3-4b-instruct-2507-wildchat Qwen3-4B SynthID three-arm corpus This export contains aligned unwatermarked, SynthID key-A, and SynthID key-B responses from Qwen/Qwen3-4B-Instruct-2507. Matched splits share prompts and request seeds across configurations; unmatched splits use mutually disjoint prompt pools. Export complete for its source work queue: true. Generation profile Model revision: cdbee75f17c01a7cc42f958dc650907174af0554 Native model dtype: bfloat16 Maximum generated tokens: 4096… See the full description on the dataset page: https://huggingface.co/datasets/xlr8harder/synthid-qwen3-4b-instruct-2507-wildchat.tabulartext-generation100K<n<1M0 likes87 downloads1mo agoHugging Face20open-athena /synthetic-misconceptions-conversations Synthetic Misconceptions Conversations All data in this dataset is synthetic. No conversation here was had by a real person. The only human-authored source material is Wikipedia text: the corrections in List of common misconceptions about science, technology, and mathematics (260 entries), plus entries from List of conspiracy theories and Category:Health-related conspiracy theories (85 entries, filtered — see below). All of it is CC BY-SA licensed on Wikipedia. Everything… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/synthetic-misconceptions-conversations.tabulartext-generation1K<n<10K0 likes81 downloads8d agoHugging Face21foxycuter /column-arithmetic-ru-synthetic Column Arithmetic RU Dataset Синтетический датасет для обучения модели сложению и вычитанию в столбик. Splits train.jsonl: основное обучение eval.jsonl: holdout-оценка hard.jsonl: трудные случаи с длинными переносами и займами Hard cases included 9999+1 10000+9999 9090+1010 55555+55555 10999+2 1234+8766 1000-7 10000-9999 50005-49999 8000-1 10101-909 100000-1 99009+991 12000-3456 700000+300001 1002003-998877 Current release status… See the full description on the dataset page: https://huggingface.co/datasets/foxycuter/column-arithmetic-ru-synthetic.tabulartext-generation1K<n<10K1 likes75 downloads5mo agoHugging Face22MaatAI /african-history-sft-synthetic African History SFT (Chat) A deduplicated, chat-format supervised fine-tuning (SFT) dataset about African history and culture, assembled from five source datasets and prepared as a ready-to-train train/test split. Each row is a multi-turn conversation in the standard messages format (system / user / assistant), making it directly usable with tokenizer.apply_chat_template and TRL's SFTTrainer. Dataset at a glance Split Rows train 25,552 test 1,345… See the full description on the dataset page: https://huggingface.co/datasets/MaatAI/african-history-sft-synthetic.tabulartext-generation10K<n<100K1 likes75 downloads1mo agoHugging Face2311-47 /fable-5-coding-and-debugging-traces-synthetic Model Synthetic Corrections 1 TRAJECTORIES · 2 TRAINING ROWS · 16 kB Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Synthetic Corrections companion dataset. The original dataset is greghavens/fable-5-coding-and-debugging-traces. These are narrowly, synthetically corrected, independently re-judged traces that never passed in the original dataset. Behavior-preserving instruction-following… See the full description on the dataset page: https://huggingface.co/datasets/11-47/fable-5-coding-and-debugging-traces-synthetic.tabulartext-generationn<1K0 likes73 downloads10d agoHugging Face24cds-jb /synthcot synthcot — carrier-paired chains-of-thought for CODI latent decoding Decoding CODI latent thoughts into chains-of-thought while controlling for the prompt-inversion pathway (a decoder that cheats by reconstructing the problem text from activations and re-solving). Analog of cds-jb/synthcognition2 for math CoT: every card pairs one computational skeleton (the "cognition") with 5 GSM8k-style carriers — word problems that all force exactly that computation — and the decode label is… See the full description on the dataset page: https://huggingface.co/datasets/cds-jb/synthcot.tabulartext-generation100K<n<1M0 likes70 downloads2mo agoHugging Face25robworks-software /database-query-logs-synthetic Database Query Logs (synthetic) 3,995 database query-log entries spanning 10 engines - MySQL, PostgreSQL, MongoDB, SQL Server, Oracle, MariaDB, SQLite, Cassandra, Redis, and Elasticsearch - with query text, type, complexity, execution timing, and row-count metadata. These queries are synthetic The queries were programmatically generated, not captured from production systems. They were produced by templating a set of query shapes across industry-flavored schema… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/database-query-logs-synthetic.tabulartext-classification1K<n<10K0 likes63 downloads2mo agoHugging Face26squeezebits /synthesized-coding-assistant-dataset Synthesized Coding Assistant Dataset Overview Coding assistants are increasingly used for real-world software engineering workflows. However, there are relatively few datasets that closely resemble how such assistants operate in practice. Many existing coding datasets are based on single-turn or single-iteration tasks, where a model receives one coding request and directly produces an answer or patch. In contrast, practical coding assistants often work through… See the full description on the dataset page: https://huggingface.co/datasets/squeezebits/synthesized-coding-assistant-dataset.tabulartext-generationn<1K0 likes62 downloads4mo agoHugging Face27neurocheckout-ai /synthetic-abandoned-cart-email-examples Synthetic Abandoned Cart Email Examples An entirely synthetic, bilingual collection of abandoned-cart email drafts with transparent checklist annotations. It is intended for education, prototyping, and evaluation, and contains no real recipients, customer messages, orders, merchant data, or campaign results. Dataset Description The dataset mirrors the five visible checks in NeuroCheckout's public Abandoned Cart Email Checker: message clarity; primary call to… See the full description on the dataset page: https://huggingface.co/datasets/neurocheckout-ai/synthetic-abandoned-cart-email-examples.tabulartext-classificationn<1K0 likes62 downloads26d agoHugging Face28boda /RevUtil_synthetic RevUtil: Measuring the Utility of Peer Reviews for Authors 📄 Paper 💻 GitHub Repository 📚 Overview Providing constructive feedback to authors is a key goal of peer review. To support research on evaluating and generating useful peer review comments, we introduce RevUtil, a dataset for measuring the utility of peer review feedback. RevUtil focuses on four main aspects of review comments: Actionability – Can the author act on the comment? Grounding & Specificity –… See the full description on the dataset page: https://huggingface.co/datasets/boda/RevUtil_synthetic.tabulartext-classification10K<n<100K0 likes59 downloads10mo agoHugging Face29haiyewon /Strudel-Synth Strudel-Synth Strudel-Synth is a synthetic corpus of 21,174 (MIDI, Strudel) pairs for training and evaluating MIDI-to-Strudel decompilation, introduced in Decomposer: Learning to Decompile Symbolic Music to Programs. 📄 Paper: arXiv:2607.01849 🌐 Project page: yewon-kim.com/decomposer 🎹 Live demo: haiyewon/decomposer-demo 🤗 Model: haiyewon/Decomposer-Qwen3-8B 💻 Code: github.com/elianakim/Decomposer Each pair consists of a Strudel program distilled from Claude-Opus-4.6… See the full description on the dataset page: https://huggingface.co/datasets/haiyewon/Strudel-Synth.tabulartext-generation10K<n<100K1 likes59 downloads2mo agoHugging Face30CurryOvO /Heterogenous_Synthesis_Benchmark Heterogenous_Synthesis_Benchmark This repository presents a diverse tabular data generation benchmark. We invite you to refer to our paper on arxiv to explore the mechanism behind our data diversity, which we called Distribution-Guided-Rule (DGR). Within this benchmark, you can experience how diverse preference data coverage combined with customized generation enhances post-training performance. Additionally, Heterogenous_Synthesis_Benchmark includes a comprehensive toolkit for… See the full description on the dataset page: https://huggingface.co/datasets/CurryOvO/Heterogenous_Synthesis_Benchmark.texttext-classification10K<n<100K1 likes56 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.