CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SynthLabsAI /Big-Math-RL-Verifiedgated Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models Big-Math is the largest open-source dataset of high-quality mathematical problems, curated specifically for reinforcement learning (RL) training in language models. With over 250,000 rigorously filtered and verified problems, Big-Math bridges the gap between quality and quantity, establishing a robust foundation for advancing reasoning in LLMs. Request Early Access to Private… See the full description on the dataset page: https://huggingface.co/datasets/SynthLabsAI/Big-Math-RL-Verified.textquestion-answering100K<n<1M243 likes5.3k downloads2y agoHugging Face02yuyijiong /context_qa_sum_qwen3_synthetic Context-based QA and Summarization Synthetic Dataset Overview This dataset contains synthetic context-based question-answering (QA) and summarization data. The data was synthesized using: Source context: openbmb/Ultra-FineWeb Synthesis model: Qwen3-30B-A3B-Instruct-2507 Each context is obtained by taking the initial segment of raw pretraining text from Ultra-FineWeb, truncated to at most the corresponding number of tokens, while ensuring the truncation does not occur in… See the full description on the dataset page: https://huggingface.co/datasets/yuyijiong/context_qa_sum_qwen3_synthetic.texttext-generation10M<n<100M5 likes3.7k downloads6mo agoHugging Face03agicommies /synthiaThe Synthia Dataset is a continuously growing aggregate of validated synthetic explanations of subjects picked from Claude Opus latent space based on varying esotericity in a large general list of technical fields. The explanations are varying in their target audience, level of detail and abstraction while incentivized to target Claude3-grade quality. The Synthia subnet leverages Commune's incentives to create a permissionless mining market around distilling knowledge out of SOTA closed-source… See the full description on the dataset page: https://huggingface.co/datasets/agicommies/synthia.texttable-question-answering1M<n<10M14 likes3.3k downloads2y agoHugging Face04gretelai /synthetic_text_to_sql Image generated by DALL-E. See prompt for more details synthetic_text_to_sql gretelai/synthetic_text_to_sql is a rich dataset of high quality synthetic Text-to-SQL samples, designed and generated using Gretel Navigator, and released under Apache 2.0. Please see our release blogpost for more details. The dataset includes: 105,851 records partitioned into 100,000 train and 5,851 test records ~23M total tokens, including ~12M SQL tokens Coverage across 100 distinct… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_text_to_sql.textquestion-answering100K<n<1M703 likes2.9k downloads9mo agoHugging Face05Azzindani /ID_Legal_QA_SynThink 🧠 Indonesian Legal QA Synthetic Think Dataset (ID_Legal_QA_SynThink) This repository features an advanced Synthetic Question-and-Answer dataset for the Indonesian legal domain, distinguished by the inclusion of an explicit Thinking Phase (Chain-of-Thought). 🏛️ 💡 The Concept: Transparent Legal Reasoning Standard QA datasets often provide just the "final answer." This dataset goes deeper by capturing the internal reasoning process of the model before it arrives at a… See the full description on the dataset page: https://huggingface.co/datasets/Azzindani/ID_Legal_QA_SynThink.tabulartext-generation1K<n<10K1 likes2.7k downloads7mo agoHugging Face06lmqg /qa_squadshifts_synthetic Dataset Card for "lmqg/qa_squadshifts_synthetic" Dataset Summary This is a synthetic QA dataset generated with fine-tuned QG models over lmqg/qa_squadshifts, made for question-answering based evaluation (QAE) for question generation model proposed by Zhang and Bansal, 2019. The test split is the original validation set of lmqg/qa_squadshifts, where the model should be evaluate on. Supported Tasks and Leaderboards question-answering Languages… See the full description on the dataset page: https://huggingface.co/datasets/lmqg/qa_squadshifts_synthetic.textquestion-answering1M<n<10M1 likes2.4k downloads4y agoHugging Face07synthetix-institute /latex-data-pub Hyperion: Scientific LaTeX Corpus (Public) This dataset constitutes the Public Scientific Corpus for the Hyperion Project at the Synthetix Institute. It contains high-fidelity LaTeX source text extracted from diverse scientific repositories, optimized for topological knowledge discovery and relational mapping. Dataset Details Total Documents: ~1,318,468 Average Document Length: Variable (approx. 32KB - 256KB) Primary Domain: Mathematics, Physics, and Chemistry. Goal:… See the full description on the dataset page: https://huggingface.co/datasets/synthetix-institute/latex-data-pub.texttext-generation1M<n<10M1 likes1.3k downloads6mo agoHugging Face08to-be /OpenHand-Synth Dataset Card for OpenHand-Synth 📜 Paper: OpenHand-Synth: A Large-Scale Synthetic Handwriting Dataset for Multimodal Language Models Sample Images Image Ground Truth Source Language CER JW 02-10-1436 faker-date por 0.10 0.96 Stephan Thomsen-Johansen faker-name dan 0.0 1.0 Le chat mange. tatoeba fra 0.0 1.0 Classical musicsoothes me.She took the risk, knowing that shemight lose a lot of money.I could not catcha single word of their talk.In the old days… See the full description on the dataset page: https://huggingface.co/datasets/to-be/OpenHand-Synth.imagefeature-extraction10K<n<100K3 likes1.2k downloads7mo agoHugging Face09timschwa /lc_quad_synth LC-QuAD 2.0-synth Dataset Summary This dataset is an updated version of the LC-QuAD 2.0 dataset which includes LLM-based natural language translations of the corresponding wikidata queries. It also includes verifier scores for the LLM translations and the original translations indicating the probability that the translation is correct (for details see our linked GitHub Repository). It contains 19000 examples of queries and translations. It can be used for training and… See the full description on the dataset page: https://huggingface.co/datasets/timschwa/lc_quad_synth.tabularquestion-answering10K<n<100K2 likes989 downloads2y agoHugging Face10starmpcc /Asclepius-Synthetic-Clinical-Notes Asclepius: Synthetic Clincal Notes & Instruction Dataset Dataset Summary This dataset is official dataset for Asclepius (arxiv) This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs. We first synthesized synthetic notes from PMC-Patients case reports with GPT-3.5 Then, we generate instruction-answer pairs for 157k synthetic discharge summaries Supported Tasks This dataset covers below 8 tasks Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/starmpcc/Asclepius-Synthetic-Clinical-Notes.textquestion-answering100K<n<1M117 likes961 downloads2y agoHugging Face11synthetix-institute /latex-data Hyperion: Scientific LaTeX Corpus This dataset constitutes the Internal Scientific Corpus for the Hyperion Project of Synthetix Institute. It contains LaTeX source documents used for high-fidelity relational extraction and internal benchmarking of the Epistemic Manifold. Dataset Details Total Documents: ~1,192,727 (Internal Base) Status: Private Access: Restricted to Synthetix Institute authorized personnel. Primary Use: Training the private sector of the Epistemic… See the full description on the dataset page: https://huggingface.co/datasets/synthetix-institute/latex-data.texttext-generation1M<n<10M0 likes756 downloads6mo agoHugging Face12instruction-pretrain /ft-instruction-synthesizer-collection Instruction Pre-Training: Language Models are Supervised Multitask Learners (EMNLP 2024) This repo contains the fine-tuning data collection for the context-based instruction synthesizer used in our paper Instruction Pre-Training: Language Models are Supervised Multitask Learners. We explore supervised multitask pre-training by proposing Instruction Pre-Training, a framework that scalably augments massive raw corpora with instruction-response pairs to pre-train language models. The… See the full description on the dataset page: https://huggingface.co/datasets/instruction-pretrain/ft-instruction-synthesizer-collection.texttext-classification100K<n<1M62 likes549 downloads7mo agoHugging Face13Kylan12 /Synthetic-AI-ML-Dataset Synthetic-AI-ML-Dataset Synthetic Q&A dataset on AI and Machine Learning Dataset Details Metric Value Topic AI and Machine Learning Total Q&A Pairs 14021 Valid Pairs 14021 Provider/Model ollama/gpt-oss:120b Generation Cost Metric Value Prompt Tokens 14,941,957 Completion Tokens 17,159,263 Total Tokens 32,101,220 GPU Energy 12.9628 kWh Sources This dataset was generated from 474 scholarly papers: #… See the full description on the dataset page: https://huggingface.co/datasets/Kylan12/Synthetic-AI-ML-Dataset.textquestion-answering10K<n<100K2 likes473 downloads6mo agoHugging Face14philschmid /gretel-synthetic-text-to-sql Fork of gretelai/synthetic_text_to_sql The gretelai/synthetic_text_to_sql dataset is a large, Apache 2.0 licensed, synthetic Text-to-SQL dataset consisting of 105,851 high-quality records across 100 diverse domains, designed for training language models. It includes comprehensive SQL tasks with varying complexities, database contexts, natural language explanations, and contextual tags, outperforming existing datasets in SQL correctness and standards compliance. textquestion-answering100K<n<1M9 likes470 downloads2y agoHugging Face15zachnorton03 /synthetic-pre1930-sftTL;DR A vintage finetuning dataset (~416k rows, eleven task routes). Sourced by taking excerpts from pre-1930's texts, turning these into verbatim answers, and then using deepseek-chat to generate period-appropriate questions of those answers. Any model tuned on this dataset should, theoretically, never update its weights on anachronistic text, since questions are masked in the finetuning stages. Features composition, verse, narrative, reasoning, multiturn dialogue, and calibrated uncertainty… See the full description on the dataset page: https://huggingface.co/datasets/zachnorton03/synthetic-pre1930-sft.tabularquestion-answering100K<n<1M2 likes313 downloads1mo agoHugging Face16gretelai /gsm8k-synthetic-diverse-8b gretelai/gsm8k-synthetic-diverse-8b This dataset is a synthetically generated version inspired by the GSM8K https://huggingface.co/datasets/openai/gsm8k dataset, created entirely using Gretel Navigator with meta-llama/Meta-Llama-3.1-8B as the agent LLM. It contains ~1500 Grade School-level math word problems with step-by-step solutions, focusing on age group, difficulty, and domain diversity. Key Features: Synthetically Generated: Math problems created using Gretel… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/gsm8k-synthetic-diverse-8b.textquestion-answering1K<n<10K0 likes293 downloads2y agoHugging Face17nvidia /Retrieval-Synthetic-NVDocs-v1 Dataset Description: Retrieval-Synthetic-NVDocs-v1 is a synthetic retrieval dataset with question–answer supervision designed to train and evaluate embedding and RAG systems. The dataset was generated on top of NVIDIA's publicly available content using NeMo Data Designer, NVIDIA's open-source framework for generating high-quality synthetic data from scratch or based on seed data. The dataset contains document chunks paired with semantically rich question-answer pairs across multiple… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Retrieval-Synthetic-NVDocs-v1.textquestion-answering10K<n<100K24 likes278 downloads6mo agoHugging Face18lianghsun /tw-legal-synthetic-qa Dataset Card for tw-legal-synthetic-qa Dataset Summary 本合成對話資料集(下稱本資料集)由 THUDM/chatglm3-6b-32k 和 lianghsun/tw-processed-judgments,由實驗後的 prompt 去生成繁體中文法律對話合成集。 Supported Tasks and Leaderboards 本資料集可以運用在 SFT,讓模型學會如何回答法律問題。 Languages 繁體中文。 Dataset Structure Data Instances 一個資料樣本如下,首先由 user 發問了一個具有(或可能有)法律情境的問題,然後 assistant 回答法律相關知識。 { "messages":[ { "role":"user"… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-legal-synthetic-qa.textquestion-answering1K<n<10K9 likes245 downloads2y agoHugging Face19Kasher13 /prospire-synth-global-personas 🌍 Prospire Synth Global Personas The World's Largest Unified Synthetic Persona Database 512M+ records · 82 columns · 77+ countries · 39 languages · DuckDB-native 🎯 What Is This? Prospire Synth Global Personas is a unified, query-ready database of synthetic human personas built for AI agent simulations, market research, and cultural analysis. It merges 18 open-source datasets into a single coherent Parquet warehouse — partitioned, compressed, and… See the full description on the dataset page: https://huggingface.co/datasets/Kasher13/prospire-synth-global-personas.texttext-generation100M<n<1B1 likes208 downloads6mo agoHugging Face20ameau01 /synthetic-it-support-tickets Synthetic IT Support Tickets — PII-Enriched + Redaction Ground Truth 745 synthetic IT service-management incident records for LLM wiki and retrieval-augmented-generation experiments. Each record is a help-desk/IT-ops incident with submitted ticket text, timestamped troubleshooting correspondence, structured diagnostics, root cause, and resolution steps. The free text is enriched with realistic technical detail and injected synthetic PII. The corpus ships two authored… See the full description on the dataset page: https://huggingface.co/datasets/ameau01/synthetic-it-support-tickets.texttext-generationn<1K0 likes203 downloads3mo agoHugging Face21CohereLabs /fusion-synth-data-geofactx Offline Synthetic Data (GeoFactX) for: Making, not taking, the Best-of-N Content This data contains completions for the GeoFactX training split prompts from 5 different teacher models and 2 aggregations: Teachers: We sample one completion from each of the following models at temperature T=0.3. For kimik2, qwen3, and deepseek-v3 we use TogetherAI, for gemma3-27b and command-a we use locally hosted images. gemma3-27b: GEMMA3-27B-IT kimik2: KIMI-K2-INSTRUCT qwen3:… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/fusion-synth-data-geofactx.texttext-generation1K<n<10K1 likes171 downloads1y agoHugging Face22CohereLabs /fusion-synth-data-s1kx Offline Synthetic Data (s1K-X) for: Making, not taking, the Best-of-N Content This data contains completions for the s1K-X training split prompts from 5 different teacher models and 2 aggregations: Teachers: We sample one completion from each of the following models at temperature T=0.3. For kimik2, qwen3, and deepseek-v3 we use TogetherAI, for gemma3-27b and command-a we use locally hosted images. gemma3-27b: GEMMA3-27B-IT kimik2: KIMI-K2-INSTRUCT qwen3: QWEN3-235B… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/fusion-synth-data-s1kx.texttext-generation10K<n<100K1 likes170 downloads1y agoHugging Face23ryokamoi /VisOnlyQA_Eval_Synthetic VisOnlyQA 🌐 Project Website | 📄 Paper | 🤗 Dataset | 🔥 VLMEvalKit This repository contains the code and data for the paper "VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information" (COLM 2025). VisOnlyQA is designed to evaluate the visual perception capability of large vision language models (LVLMs) on geometric information of scientific figures. The evaluation set includes 1,200 mlutiple choice questions in 12 visual perception… See the full description on the dataset page: https://huggingface.co/datasets/ryokamoi/VisOnlyQA_Eval_Synthetic.imagemultiple-choicen<1K2 likes165 downloads1y agoHugging Face24amd /Instella-GSM8K-synthetic Instella-GSM8K-synthetic The Instella-GSM8K-synthetic dataset was used in the second stage pre-training of Instella-3B model, which was trained on top of the Instella-3B-Stage1 model. This synthetic dataset was generated using the training set of GSM8k dataset, where we first used Qwen2.5-72B-Instruct to Abstract numerical values as function parameters and generate a Python program to solve the math question. Identify and replace numerical values in the existing question with… See the full description on the dataset page: https://huggingface.co/datasets/amd/Instella-GSM8K-synthetic.textquestion-answering1M<n<10M7 likes154 downloads11mo agoHugging Face25ZennyKenny /synthetic_vc_financial_decisions_reasoning_dataset Best Curator Use Case in the Reasoning Datasets Competition: https://www.linkedin.com/feed/update/urn:li:activity:7330998995990781952/ Synthetic VC Financial Decisions Reasoning Dataset Dataset Summary The Synthetic VC Financial Decisions Reasoning Dataset is a large-scale collection designed to train, evaluate, and fine-tune language models on subjective, abstract financial reasoning tasks. It simulates venture capital (VC) workflows by capturing multiple… See the full description on the dataset page: https://huggingface.co/datasets/ZennyKenny/synthetic_vc_financial_decisions_reasoning_dataset.textreinforcement-learningn<1K15 likes144 downloads1y agoHugging Face26wayneworkman2012 /ww2-synthetic-corpusgated WWII Synthetic LLM Training Corpus A large synthetic dataset of conversational training examples about World War II, generated by a custom synthetic-data pipeline. The knowledge originates in curated Wikipedia articles; large language models were used only as transformation tools to reformat that source knowledge into diverse training patterns — they are not the source of the facts. Total examples: 16,397,795 Subsets (configs): 92 Format: ChatML-style messages (role/content)… See the full description on the dataset page: https://huggingface.co/datasets/wayneworkman2012/ww2-synthetic-corpus.texttext-generation10M<n<100M2 likes144 downloads16d agoHugging Face27dendriteholdings /Dendrite-Synth-Multi-Domain Dendrite Synth Multi-Domain A verified, difficulty-filtered, style-amplified synthetic corpus of question / reasoning / answer triples spanning the 14 MMLU-Pro categories - mathematics, computer science, natural sciences, chemistry, physics, engineering, health, law, business, economics, psychology, philosophy, history and others (expanded to 122 fine-grained categories and 695 subcategories). Problems are written by a pool of generator models, solved with explicit reasoning by… See the full description on the dataset page: https://huggingface.co/datasets/dendriteholdings/Dendrite-Synth-Multi-Domain.texttext-generation100K<n<1M0 likes132 downloads25d agoHugging Face28CJJones /LLM_Electrical_Engineering_Educational_Synthetic_DialogDataset Card for LLM_Electrical_Engineering_Educational_Synthetic_Dialog The full CJ Jones' synthetic dataset catalog is available at: https://datadeveloper1.gumroad.com Want more? 🚀 Get the AI Startup Bundle from Gumroad. Dataset Description The LLM_Electrical_Engineering_Educational_Synthetic_Dialog dataset contains AI-generated conversational interactions designed for training large language models in electrical engineering education. This synthetic dialogue corpus simulates tutor-student… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/LLM_Electrical_Engineering_Educational_Synthetic_Dialog.textquestion-answering10K<n<100K3 likes129 downloads7mo agoHugging Face29gretelai /synthetic-gsm8k-evolutionary-405b gretelai/synthetic-gsm8k-evolutionary-405b This dataset is a synthetically generated version inspired by the GSM8K dataset, created entirely using Gretel Navigator with meta-llama/Meta-Llama-3.1-405B as the agent LLM. It contains Grade School-level reasoning tasks with step-by-step solutions, focusing on multi-step reasoning problems. Key Features: Synthetically Generated: Built using Gretel Navigator, leveraging evolutionary approach for diversity to create both the… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic-gsm8k-evolutionary-405b.textquestion-answering1K<n<10K5 likes124 downloads2y agoHugging Face30gratex /GNOTHEIA-synthetic-insurance-dataset GNOTHEIA Synthetic Insurance Dataset Published by: Gratex International a.s.Project: InnovAIte — InnovAIte Slovakia License: Apache 2.0Version: 1.0.0Contact: info@gratex.com A synthetic insurance claims dataset designed for AI systems that evaluate insurance claims using OMG SBVR business rules, structured claim polycontexts and synthetic claim-related documents. The dataset main goal is to support: LLM fine-tuning pipeline SBVR reasoning benchmarks insurance claim AI… See the full description on the dataset page: https://huggingface.co/datasets/gratex/GNOTHEIA-synthetic-insurance-dataset.tabulartext-classification1K<n<10K0 likes122 downloads5d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.