CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01PleIAs /SYNTH SYNTH Blog announcement SYNTH is the first open generalist synthetic dataset for training small reasoning model end-to-end, jointly released by Pleias and the AI Alliance. SYNTH includes 79,648,272 individual text samples, comprising over 41 billion words (about 75 billion tokens with Pleias tokenizer). It is based on the amplification of 58,698 articles from Wikipedia and made possible thanks to the Structured Wikipedia dataset from Wikimedia Enterprise. SYNTH differs… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/SYNTH.texttext-generation10M<n<100M277 likes13k downloads5mo agoHugging Face02SynthLabsAI /Big-Math-RL-Verifiedgated Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models Big-Math is the largest open-source dataset of high-quality mathematical problems, curated specifically for reinforcement learning (RL) training in language models. With over 250,000 rigorously filtered and verified problems, Big-Math bridges the gap between quality and quantity, establishing a robust foundation for advancing reasoning in LLMs. Request Early Access to Private… See the full description on the dataset page: https://huggingface.co/datasets/SynthLabsAI/Big-Math-RL-Verified.textquestion-answering100K<n<1M243 likes5.3k downloads2y agoHugging Face03SYNTH-Initiative /SYNTH SYNTH SYNTH is the first open generalist synthetic dataset for training small reasoning model end-to-end, jointly released by Pleias and the AI Alliance. SYNTH includes 79,648,272 individual text samples, comprising over 41 billion words (about 75 billion tokens with Pleias tokenizer). It is based on the amplification of 58,698 articles from Wikipedia and made possible thanks to the Structured Wikipedia dataset from Wikimedia Enterprise. SYNTH differs from existing open synthetic… See the full description on the dataset page: https://huggingface.co/datasets/SYNTH-Initiative/SYNTH.texttext-generation10M<n<100M0 likes4.7k downloads11mo agoHugging Face04yuyijiong /context_qa_sum_qwen3_synthetic Context-based QA and Summarization Synthetic Dataset Overview This dataset contains synthetic context-based question-answering (QA) and summarization data. The data was synthesized using: Source context: openbmb/Ultra-FineWeb Synthesis model: Qwen3-30B-A3B-Instruct-2507 Each context is obtained by taking the initial segment of raw pretraining text from Ultra-FineWeb, truncated to at most the corresponding number of tokens, while ensuring the truncation does not occur in… See the full description on the dataset page: https://huggingface.co/datasets/yuyijiong/context_qa_sum_qwen3_synthetic.texttext-generation10M<n<100M5 likes3.7k downloads6mo agoHugging Face05McGill-NLP /A3-Synth A3-Synth 💾 Code 📄 Paper 🌐 Website 🤗 Dataset 🤖 Models 📦 PyPI Structured Distillation of Web Agent Capabilities Enables Generalization Xing Han Lù, Siva Reddy A3-Synth is a synthetic training dataset for web agents, generated using the Agent-as-Annotators (A3) framework. It contains ~16k SFT training examples produced by Gemini 3 Pro acting as the Annotator across 3,000 tasks on 6 WebArena environments. Dataset Structure A3-Synth/ training/… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/A3-Synth.text-generation10K<n<100K1 likes3.5k downloads6mo agoHugging Face06gretelai /synthetic_text_to_sql Image generated by DALL-E. See prompt for more details synthetic_text_to_sql gretelai/synthetic_text_to_sql is a rich dataset of high quality synthetic Text-to-SQL samples, designed and generated using Gretel Navigator, and released under Apache 2.0. Please see our release blogpost for more details. The dataset includes: 105,851 records partitioned into 100,000 train and 5,851 test records ~23M total tokens, including ~12M SQL tokens Coverage across 100 distinct… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_text_to_sql.textquestion-answering100K<n<1M703 likes2.9k downloads9mo agoHugging Face07argilla /Synth-APIGen-v0.1 Dataset card for Synth-APIGen-v0.1 This dataset has been created with distilabel. Pipeline script: pipeline_apigen_train.py. Dataset creation It has been created with distilabel==1.4.0 version. This dataset is an implementation of APIGen: Automated Pipeline for Generating Verifiable and Diverse Function-Calling Datasets in distilabel, generated from synthetic functions. The process can be summarized as follows: Generate (or in this case modify) python… See the full description on the dataset page: https://huggingface.co/datasets/argilla/Synth-APIGen-v0.1.texttext-generation10K<n<100K65 likes2.8k downloads2y agoHugging Face08Azzindani /ID_Legal_QA_SynThink 🧠 Indonesian Legal QA Synthetic Think Dataset (ID_Legal_QA_SynThink) This repository features an advanced Synthetic Question-and-Answer dataset for the Indonesian legal domain, distinguished by the inclusion of an explicit Thinking Phase (Chain-of-Thought). 🏛️ 💡 The Concept: Transparent Legal Reasoning Standard QA datasets often provide just the "final answer." This dataset goes deeper by capturing the internal reasoning process of the model before it arrives at a… See the full description on the dataset page: https://huggingface.co/datasets/Azzindani/ID_Legal_QA_SynThink.tabulartext-generation1K<n<10K1 likes2.7k downloads7mo agoHugging Face09UCB-team /unclickbait-synthetic-27b-trajectories Unclickbait Synthetic 27B Trajectories Synthetic trajectory dataset generated by using rich structured JSON prompts and validated by two-stage judging pipeline. Contents : Full generated trajectories (current snapshot: 48,623 records out of 152,369 pristine event candidates). : 30 benchmark test samples audited end-to-end through the 122B two-stage judge (Stage 1 integrity gate + Stage 2 4D scoring). texttext-generationn<1K0 likes2.7k downloads12d agoHugging Face10qvac /TranslatePsy-AfriSLM-Synthetic-Mix TranslatePsy-AfriSLM Synthetic Mix TranslatePsy-AfriSLM Synthetic Mix is a quality-filtered synthetic parallel corpus for machine translation between English and 19 Sub-Saharan African languages. It contains 215,653,192 bidirectional training examples and was selected as the primary African translation component used to post-train the TranslatePsy-AfriSLM model family. The dataset accompanies the EMNLP 2026 paper TranslatePsy-AfriSLM: High-Quality Data Scaling For Low-Resource… See the full description on the dataset page: https://huggingface.co/datasets/qvac/TranslatePsy-AfriSLM-Synthetic-Mix.texttranslation100M<n<1B0 likes1.4k downloads1mo agoHugging Face11synthetix-institute /latex-data-pub Hyperion: Scientific LaTeX Corpus (Public) This dataset constitutes the Public Scientific Corpus for the Hyperion Project at the Synthetix Institute. It contains high-fidelity LaTeX source text extracted from diverse scientific repositories, optimized for topological knowledge discovery and relational mapping. Dataset Details Total Documents: ~1,318,468 Average Document Length: Variable (approx. 32KB - 256KB) Primary Domain: Mathematics, Physics, and Chemistry. Goal:… See the full description on the dataset page: https://huggingface.co/datasets/synthetix-institute/latex-data-pub.texttext-generation1M<n<10M1 likes1.3k downloads6mo agoHugging Face12Synthetic-Entropy-Labs /entropy-shard-002 Synthetic Entropy Shards (v2) Dataset Description This dataset consists of high-entropy binary shards generated for stress-testing data loading pipelines and training robustness against random noise injection in large-scale tensor operations. Usage These files are intended to be consumed as raw byte streams. Due to the stochastic nature of the generation process, the data mimics encrypted traffic patterns or high-density compression artifacts. Warning: Do not… See the full description on the dataset page: https://huggingface.co/datasets/Synthetic-Entropy-Labs/entropy-shard-002.feature-extraction0 likes1.3k downloads7mo agoHugging Face13to-be /OpenHand-Synth Dataset Card for OpenHand-Synth 📜 Paper: OpenHand-Synth: A Large-Scale Synthetic Handwriting Dataset for Multimodal Language Models Sample Images Image Ground Truth Source Language CER JW 02-10-1436 faker-date por 0.10 0.96 Stephan Thomsen-Johansen faker-name dan 0.0 1.0 Le chat mange. tatoeba fra 0.0 1.0 Classical musicsoothes me.She took the risk, knowing that shemight lose a lot of money.I could not catcha single word of their talk.In the old days… See the full description on the dataset page: https://huggingface.co/datasets/to-be/OpenHand-Synth.imagefeature-extraction10K<n<100K3 likes1.2k downloads7mo agoHugging Face14Arko007 /zenyx-v3-synthetic-sft-v4 Zenyx-V3-Synthetic-SFT-V4 (Master Expansion) This is the finalized V4 Master Collection for the Zenyx project, expanding upon the previous V3. This version focuses on high-reasoning, code, and math capabilities through massive distillation and thinking-chain integration. Dataset Summary Total Samples: 101,523 Branding: Rebranded to Zenyx / Zenyx Lab. Filtering: Strict English, Math, and Code filter applied. Chinese and non-standard characters removed. Format:… See the full description on the dataset page: https://huggingface.co/datasets/Arko007/zenyx-v3-synthetic-sft-v4.text-generation100K<n<1M0 likes1.1k downloads3mo agoHugging Face15Synthetic-Entropy-Labs /entropy-shard-001 Synthetic Entropy Shards (v1) Dataset Description This dataset consists of high-entropy binary shards generated for stress-testing data loading pipelines and training robustness against random noise injection in large-scale tensor operations. Usage These files are intended to be consumed as raw byte streams. Due to the stochastic nature of the generation process, the data mimics encrypted traffic patterns or high-density compression artifacts. Warning: Do not… See the full description on the dataset page: https://huggingface.co/datasets/Synthetic-Entropy-Labs/entropy-shard-001.feature-extraction0 likes966 downloads8mo agoHugging Face16starmpcc /Asclepius-Synthetic-Clinical-Notes Asclepius: Synthetic Clincal Notes & Instruction Dataset Dataset Summary This dataset is official dataset for Asclepius (arxiv) This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs. We first synthesized synthetic notes from PMC-Patients case reports with GPT-3.5 Then, we generate instruction-answer pairs for 157k synthetic discharge summaries Supported Tasks This dataset covers below 8 tasks Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/starmpcc/Asclepius-Synthetic-Clinical-Notes.textquestion-answering100K<n<1M117 likes961 downloads2y agoHugging Face17tokyotech-llm /lmsys-chat-1m-synth LMSYS-Chat-1M-Synth: Japanese/English Synthetic Conversation Dataset Derived from LMSYS-Chat-1M This repository contains a series of Japanese and English conversation datasets derived from LMSYS-Chat-1M. Llama-3.1-LMSYS-Chat-1M-Synth Utilized in the post-training of Llama-3.1-Swallow-8B-Instruct-v0.1 and Llama-3.1-Swallow-70B-Instruct-v0.1 Gemma-2-LMSYS-Chat-1M-Synth Utilized in the post-training of Llama-3.1-Swallow-8B-Instruct-v0.3 and Llama-3.1-Swallow-70B-Instruct-v0.3… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/lmsys-chat-1m-synth.text-generation100K<n<1M23 likes868 downloads7mo agoHugging Face18MachineLearningLM /machinelearninglm-scm-synthetic-tabularml MachineLearningLM Pretraining Corpus This repository contains the pretraining corpus for MachineLearningLM, a framework designed to equip large language models (LLMs) with robust in-context machine learning (ML) capabilities. The dataset consists of ML tasks synthesized from millions of structural causal models (SCMs), spanning various shot counts up to 1,024. It is designed to enable LLMs to learn from many in-context examples on standard ML tasks purely via in-context learning… See the full description on the dataset page: https://huggingface.co/datasets/MachineLearningLM/machinelearninglm-scm-synthetic-tabularml.texttext-generation1M<n<10M4 likes814 downloads10mo agoHugging Face19TMoC /SYNTH-Swallow-Math-Code-Mix Mixed dataset: SYNTH + SwallowMath-v2 + SwallowCode-v2 This high-signal, all synthetic dataset is a complete shuffled mix of the following four sources: SYNTH ~63.5% SwallowCode-v2 ~15.5% SwallowMath-v2-textbook ~10.5% SwallowMath-v2-qa ~10.0% The motivation to provide this on HF was the need for a convenient, pre-shuffled merge of the highest quality synthetic / augmented datasets for small language model pre-training experiments as of… See the full description on the dataset page: https://huggingface.co/datasets/TMoC/SYNTH-Swallow-Math-Code-Mix.texttext-generation100M<n<1B1 likes761 downloads8mo agoHugging Face20synthetix-institute /latex-data Hyperion: Scientific LaTeX Corpus This dataset constitutes the Internal Scientific Corpus for the Hyperion Project of Synthetix Institute. It contains LaTeX source documents used for high-fidelity relational extraction and internal benchmarking of the Epistemic Manifold. Dataset Details Total Documents: ~1,192,727 (Internal Base) Status: Private Access: Restricted to Synthetix Institute authorized personnel. Primary Use: Training the private sector of the Epistemic… See the full description on the dataset page: https://huggingface.co/datasets/synthetix-institute/latex-data.texttext-generation1M<n<10M0 likes756 downloads6mo agoHugging Face21rajistics /openhands-synthetic-conversations OpenHands Synthetic Conversations Overview 28 synthetic OpenHands V1 agent conversations generated by running diverse coding-task prompts against the OpenHands SDK with 4 models rotated round-robin. Each conversation captures a complete agentic session: system prompt, user message, tool calls, terminal observations, and the agent's final reply — exactly as produced by the app.all-hands.dev "Download Conversation" export. Intended use: raw material for indexing /… See the full description on the dataset page: https://huggingface.co/datasets/rajistics/openhands-synthetic-conversations.text-generationn<1K0 likes712 downloads4mo agoHugging Face22richardyoung /synthea-575k-patients Synthea Synthetic Patient Records (575K Patients) A comprehensive synthetic healthcare dataset containing 575,415 patients with complete medical histories, generated using Synthea — the gold standard for synthetic EHR data. No real patient data. Fully synthetic, HIPAA-safe, and ready for ML research and education. Why This Dataset? 575K patients with realistic demographics, conditions, medications, and encounters Privacy-safe: No real PHI — use freely in research… See the full description on the dataset page: https://huggingface.co/datasets/richardyoung/synthea-575k-patients.tabulartext-generation1B<n<10B12 likes678 downloads6mo agoHugging Face23argilla-warehouse /synth-apigen-qwen Dataset Card for argilla-warehouse/synth-apigen-qwen This dataset has been created with distilabel. The pipeline script was uploaded to easily reproduce the dataset: synth_apigen.py. Dataset creation This dataset is a replica in distilabel of the framework defined in: APIGen: Automated Pipeline for Generating Verifiable and Diverse Function-Calling Datasets. Using the seed dataset of synthetic python functions in argilla-warehouse/python-seed-tools, the… See the full description on the dataset page: https://huggingface.co/datasets/argilla-warehouse/synth-apigen-qwen.texttext-generation10K<n<100K7 likes663 downloads2y agoHugging Face24devanshamin /synthetic-pii-function-calling Dataset Summary A function calling dataset created by filtering the urchade/synthetic-pii-ner-mistral-v1 dataset. texttext-generation1K<n<10K0 likes531 downloads2y agoHugging Face25KirillR /Infinity-Instruct-RU-Synthetic Infinity-Instruct-RU-Synthetic A large-scale Russian-language instructional dataset based on Infinity-Instruct by BAAI. This is not a translation of English answers — it is an independent Russian-language dataset, where only the instructions are sourced from the original set, and all answers are newly generated in Russian from scratch. To translate the instructions, YandexGPT-5-Lite-8B-instruct was used with a specially fine-tuned LoRA adapter designed for this dataset. The original… See the full description on the dataset page: https://huggingface.co/datasets/KirillR/Infinity-Instruct-RU-Synthetic.texttext-generation1M<n<10M6 likes522 downloads1y agoHugging Face26open-athena /recursive-task-synthesis-glm-5.3-rollouts GLM 5.3 agentic rollouts on Recursive-Task-Synthesis This dataset catalogs the full collection made from the pinned Recursive-Task-Synthesis dataset revision be44f96808d5a9b599d5cb024341ff00091adeb7. The repository includes approximately 260.5 GiB of trajectory payload tar shards. Contents at a glance Item Count Source tasks considered 37,284 Source candidates inspected 19,368 Converted tasks after source filters 18,600 Tasks passing gold… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/recursive-task-synthesis-glm-5.3-rollouts.tabulartext-generation100K<n<1M0 likes509 downloads6d agoHugging Face27vibhuiitj /UltraData-Math-L3-Textbook-Exercise-Synthetic-split UltraData-Math L3 Textbook Exercise Synthetic Split Source dataset: openbmb/UltraData-Math Source config: UltraData-Math-L3-Textbook-Exercise-Synthetic Each row contains: uid question answer The original content field was split using the literal markers The exercise: and The solution:. texttext-generation10M<n<100M1 likes502 downloads6mo agoHugging Face28aaaaliou /pi-synthetic Coding agent session traces for aaaaliou/pi-synthetic This dataset contains redacted coding agent session traces collected while working on git@github.com:aliou/pi-synthetic.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review. Data description Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line is a structured… See the full description on the dataset page: https://huggingface.co/datasets/aaaaliou/pi-synthetic.tabulartext-generationn<1K0 likes498 downloads5mo agoHugging Face29pdelobelle /staatsblad-synth-nl Synthetic Dutch from the Belgisch Staatsblad Diverse, fluent Dutch pretraining text synthesized from guust-franssens/belgisch-staatsblad (CC0, Belgian official-gazette filings). Adds Belgium/Flanders coverage to Dutch LM pretraining mixes, where clean Belgian-Dutch prose is otherwise scarce. The source text is noisy OCR from scanned PDFs, but its metadata (company, juridical form, act type, city, date) is clean. A local LLM (google/gemma-2-9b-it) "launders" the OCR + metadata… See the full description on the dataset page: https://huggingface.co/datasets/pdelobelle/staatsblad-synth-nl.texttext-generation100K<n<1M0 likes477 downloads2mo agoHugging Face30philschmid /gretel-synthetic-text-to-sql Fork of gretelai/synthetic_text_to_sql The gretelai/synthetic_text_to_sql dataset is a large, Apache 2.0 licensed, synthetic Text-to-SQL dataset consisting of 105,851 high-quality records across 100 diverse domains, designed for training language models. It includes comprehensive SQL tasks with varying complexities, database contexts, natural language explanations, and contextual tags, outperforming existing datasets in SQL correctness and standards compliance. textquestion-answering100K<n<1M9 likes470 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.