CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nvidia /Nemotron-Terminal-Synthetic-Tasks Terminal-Corpus: Task Structure Specification This repository contains the skill-based synthetic tasks within the Terminal-Corpus. These tasks are designed to evaluate and train autonomous agents in realistic Linux terminal environments. 🏗️ Task Anatomy Each task is contained within a dedicated directory and follows a strict four-component architecture: 1. Instruction (instruction.md) Purpose: Provides the natural language description of the objective.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Terminal-Synthetic-Tasks.question-answering100K<n<1M31 likes4.8k downloads7mo agoHugging Face02yuyijiong /context_qa_sum_qwen3_synthetic Context-based QA and Summarization Synthetic Dataset Overview This dataset contains synthetic context-based question-answering (QA) and summarization data. The data was synthesized using: Source context: openbmb/Ultra-FineWeb Synthesis model: Qwen3-30B-A3B-Instruct-2507 Each context is obtained by taking the initial segment of raw pretraining text from Ultra-FineWeb, truncated to at most the corresponding number of tokens, while ensuring the truncation does not occur in… See the full description on the dataset page: https://huggingface.co/datasets/yuyijiong/context_qa_sum_qwen3_synthetic.texttext-generation10M<n<100M5 likes3.7k downloads6mo agoHugging Face03gretelai /synthetic_text_to_sql Image generated by DALL-E. See prompt for more details synthetic_text_to_sql gretelai/synthetic_text_to_sql is a rich dataset of high quality synthetic Text-to-SQL samples, designed and generated using Gretel Navigator, and released under Apache 2.0. Please see our release blogpost for more details. The dataset includes: 105,851 records partitioned into 100,000 train and 5,851 test records ~23M total tokens, including ~12M SQL tokens Coverage across 100 distinct… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_text_to_sql.textquestion-answering100K<n<1M703 likes2.9k downloads9mo agoHugging Face04lmqg /qa_squadshifts_synthetic Dataset Card for "lmqg/qa_squadshifts_synthetic" Dataset Summary This is a synthetic QA dataset generated with fine-tuned QG models over lmqg/qa_squadshifts, made for question-answering based evaluation (QAE) for question generation model proposed by Zhang and Bansal, 2019. The test split is the original validation set of lmqg/qa_squadshifts, where the model should be evaluate on. Supported Tasks and Leaderboards question-answering Languages… See the full description on the dataset page: https://huggingface.co/datasets/lmqg/qa_squadshifts_synthetic.textquestion-answering1M<n<10M1 likes2.7k downloads4y agoHugging Face05starmpcc /Asclepius-Synthetic-Clinical-Notes Asclepius: Synthetic Clincal Notes & Instruction Dataset Dataset Summary This dataset is official dataset for Asclepius (arxiv) This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs. We first synthesized synthetic notes from PMC-Patients case reports with GPT-3.5 Then, we generate instruction-answer pairs for 157k synthetic discharge summaries Supported Tasks This dataset covers below 8 tasks Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/starmpcc/Asclepius-Synthetic-Clinical-Notes.textquestion-answering100K<n<1M117 likes967 downloads2y agoHugging Face06Kylan12 /Synthetic-AI-ML-Dataset Synthetic-AI-ML-Dataset Synthetic Q&A dataset on AI and Machine Learning Dataset Details Metric Value Topic AI and Machine Learning Total Q&A Pairs 14021 Valid Pairs 14021 Provider/Model ollama/gpt-oss:120b Generation Cost Metric Value Prompt Tokens 14,941,957 Completion Tokens 17,159,263 Total Tokens 32,101,220 GPU Energy 12.9628 kWh Sources This dataset was generated from 474 scholarly papers: #… See the full description on the dataset page: https://huggingface.co/datasets/Kylan12/Synthetic-AI-ML-Dataset.textquestion-answering10K<n<100K2 likes472 downloads6mo agoHugging Face07philschmid /gretel-synthetic-text-to-sql Fork of gretelai/synthetic_text_to_sql The gretelai/synthetic_text_to_sql dataset is a large, Apache 2.0 licensed, synthetic Text-to-SQL dataset consisting of 105,851 high-quality records across 100 diverse domains, designed for training language models. It includes comprehensive SQL tasks with varying complexities, database contexts, natural language explanations, and contextual tags, outperforming existing datasets in SQL correctness and standards compliance. textquestion-answering100K<n<1M9 likes454 downloads2y agoHugging Face08zachnorton03 /synthetic-pre1930-sftTL;DR A vintage finetuning dataset (~416k rows, eleven task routes). Sourced by taking excerpts from pre-1930's texts, turning these into verbatim answers, and then using deepseek-chat to generate period-appropriate questions of those answers. Any model tuned on this dataset should, theoretically, never update its weights on anachronistic text, since questions are masked in the finetuning stages. Features composition, verse, narrative, reasoning, multiturn dialogue, and calibrated uncertainty… See the full description on the dataset page: https://huggingface.co/datasets/zachnorton03/synthetic-pre1930-sft.tabularquestion-answering100K<n<1M2 likes333 downloads1mo agoHugging Face09gretelai /gsm8k-synthetic-diverse-8b gretelai/gsm8k-synthetic-diverse-8b This dataset is a synthetically generated version inspired by the GSM8K https://huggingface.co/datasets/openai/gsm8k dataset, created entirely using Gretel Navigator with meta-llama/Meta-Llama-3.1-8B as the agent LLM. It contains ~1500 Grade School-level math word problems with step-by-step solutions, focusing on age group, difficulty, and domain diversity. Key Features: Synthetically Generated: Math problems created using Gretel… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/gsm8k-synthetic-diverse-8b.textquestion-answering1K<n<10K0 likes300 downloads2y agoHugging Face10nvidia /Retrieval-Synthetic-NVDocs-v1 Dataset Description: Retrieval-Synthetic-NVDocs-v1 is a synthetic retrieval dataset with question–answer supervision designed to train and evaluate embedding and RAG systems. The dataset was generated on top of NVIDIA's publicly available content using NeMo Data Designer, NVIDIA's open-source framework for generating high-quality synthetic data from scratch or based on seed data. The dataset contains document chunks paired with semantically rich question-answer pairs across multiple… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Retrieval-Synthetic-NVDocs-v1.textquestion-answering10K<n<100K24 likes273 downloads6mo agoHugging Face11lianghsun /tw-legal-synthetic-qa Dataset Card for tw-legal-synthetic-qa Dataset Summary 本合成對話資料集(下稱本資料集)由 THUDM/chatglm3-6b-32k 和 lianghsun/tw-processed-judgments,由實驗後的 prompt 去生成繁體中文法律對話合成集。 Supported Tasks and Leaderboards 本資料集可以運用在 SFT,讓模型學會如何回答法律問題。 Languages 繁體中文。 Dataset Structure Data Instances 一個資料樣本如下,首先由 user 發問了一個具有(或可能有)法律情境的問題,然後 assistant 回答法律相關知識。 { "messages":[ { "role":"user"… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-legal-synthetic-qa.textquestion-answering1K<n<10K9 likes259 downloads2y agoHugging Face12therayelab /synthetic-stroke-patients-v3 Synthetic Stroke Patient Episodes v3 Release status: PILOT_ONLY. Clinical use is not approved. This repository contains 1,000 wholly synthetic adult stroke episodes with 130 structured dimensions, compositional clinical notes, and 5,000 time-indexed educational task examples. It contains no real patient records. Splits Train: 700 patients Validation: 100 patients Test: 100 patients Challenge: 100 patients Each episode is also represented at five information… See the full description on the dataset page: https://huggingface.co/datasets/therayelab/synthetic-stroke-patients-v3.text-classification1K<n<10K2 likes225 downloads9d agoHugging Face13ameau01 /synthetic-it-support-tickets Synthetic IT Support Tickets — PII-Enriched + Redaction Ground Truth 745 synthetic IT service-management incident records for LLM wiki and retrieval-augmented-generation experiments. Each record is a help-desk/IT-ops incident with submitted ticket text, timestamped troubleshooting correspondence, structured diagnostics, root cause, and resolution steps. The free text is enriched with realistic technical detail and injected synthetic PII. The corpus ships two authored… See the full description on the dataset page: https://huggingface.co/datasets/ameau01/synthetic-it-support-tickets.texttext-generationn<1K0 likes197 downloads3mo agoHugging Face14amd /Instella-GSM8K-synthetic Instella-GSM8K-synthetic The Instella-GSM8K-synthetic dataset was used in the second stage pre-training of Instella-3B model, which was trained on top of the Instella-3B-Stage1 model. This synthetic dataset was generated using the training set of GSM8k dataset, where we first used Qwen2.5-72B-Instruct to Abstract numerical values as function parameters and generate a Python program to solve the math question. Identify and replace numerical values in the existing question with… See the full description on the dataset page: https://huggingface.co/datasets/amd/Instella-GSM8K-synthetic.textquestion-answering1M<n<10M7 likes167 downloads10mo agoHugging Face15ryokamoi /VisOnlyQA_Eval_Synthetic VisOnlyQA 🌐 Project Website | 📄 Paper | 🤗 Dataset | 🔥 VLMEvalKit This repository contains the code and data for the paper "VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information" (COLM 2025). VisOnlyQA is designed to evaluate the visual perception capability of large vision language models (LVLMs) on geometric information of scientific figures. The evaluation set includes 1,200 mlutiple choice questions in 12 visual perception… See the full description on the dataset page: https://huggingface.co/datasets/ryokamoi/VisOnlyQA_Eval_Synthetic.imagemultiple-choicen<1K2 likes158 downloads1y agoHugging Face16ZennyKenny /synthetic_vc_financial_decisions_reasoning_dataset Best Curator Use Case in the Reasoning Datasets Competition: https://www.linkedin.com/feed/update/urn:li:activity:7330998995990781952/ Synthetic VC Financial Decisions Reasoning Dataset Dataset Summary The Synthetic VC Financial Decisions Reasoning Dataset is a large-scale collection designed to train, evaluate, and fine-tune language models on subjective, abstract financial reasoning tasks. It simulates venture capital (VC) workflows by capturing multiple… See the full description on the dataset page: https://huggingface.co/datasets/ZennyKenny/synthetic_vc_financial_decisions_reasoning_dataset.textreinforcement-learningn<1K15 likes151 downloads1y agoHugging Face17wayneworkman2012 /ww2-synthetic-corpusgated WWII Synthetic LLM Training Corpus A large synthetic dataset of conversational training examples about World War II, generated by a custom synthetic-data pipeline. The knowledge originates in curated Wikipedia articles; large language models were used only as transformation tools to reformat that source knowledge into diverse training patterns — they are not the source of the facts. Total examples: 16,397,795 Subsets (configs): 92 Format: ChatML-style messages (role/content)… See the full description on the dataset page: https://huggingface.co/datasets/wayneworkman2012/ww2-synthetic-corpus.texttext-generation10M<n<100M2 likes147 downloads14d agoHugging Face18vinhnx90 /synthetic-swift-data-single-turn Dataset Card for synthetic-swift-data-single-turn This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/vinhnx90/synthetic-swift-data-single-turn/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/vinhnx90/synthetic-swift-data-single-turn.texttext-generationn<1K0 likes144 downloads2y agoHugging Face19gratex /GNOTHEIA-synthetic-insurance-dataset GNOTHEIA Synthetic Insurance Dataset Published by: Gratex International a.s.Project: InnovAIte — InnovAIte Slovakia License: Apache 2.0Version: 1.0.0Contact: info@gratex.com A synthetic insurance claims dataset designed for AI systems that evaluate insurance claims using OMG SBVR business rules, structured claim polycontexts and synthetic claim-related documents. The dataset main goal is to support: LLM fine-tuning pipeline SBVR reasoning benchmarks insurance claim AI… See the full description on the dataset page: https://huggingface.co/datasets/gratex/GNOTHEIA-synthetic-insurance-dataset.tabulartext-classification1K<n<10K0 likes136 downloads3d agoHugging Face20yamaTK /japanese-math-synthetic-108k clean_v1_noassert_allcats LLM-JP チューニングコンテスト 2026 向けに作成した日本語数学問題の学習データセットです。 中学1年〜高校数学IIICまでの各カテゴリを網羅しています。問題作成・解法生成・検証のすべてを GPT-OSS (120B) で行い、重複除去を経てクリーニング済みです。 データ概要 2種類の分類軸でデータを収録しています。 学年別(中学〜高校カリキュラム準拠) ファイル カテゴリ 行数 主な単元 chu1.jsonl 中学1年 9,610 正負の数, 文字式, 一次方程式, 比例反比例 chu2.jsonl 中学2年 9,761 文字式, 一次関数, 連立方程式, 確率 chu3.jsonl 中学3年 9,454 二次方程式, 二次関数, 平方根, 展開と因数分解 IA.jsonl 数学IA 9,881 整数の性質, 場合の数と確率, 2次関数, 数と式 IIB.jsonl 数学IIB 11,089 数列, いろいろな式… See the full description on the dataset page: https://huggingface.co/datasets/yamaTK/japanese-math-synthetic-108k.question-answering100K<n<1M2 likes135 downloads7mo agoHugging Face21gretelai /synthetic_multilingual_llm_prompts Image generated by DALL-E. See prompt for more details 📝🌐 Synthetic Multilingual LLM Prompts Welcome to the "Synthetic Multilingual LLM Prompts" dataset! This comprehensive collection features 1,250 synthetic LLM prompts generated using Gretel Navigator, available in seven different languages. To ensure accuracy and diversity in prompts, and translation quality and consistency across the different languages, we employed Gretel Navigator both as a generation tool and as an… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_multilingual_llm_prompts.tabulartext-generation1K<n<10K11 likes128 downloads2y agoHugging Face22gretelai /synthetic-gsm8k-evolutionary-405b gretelai/synthetic-gsm8k-evolutionary-405b This dataset is a synthetically generated version inspired by the GSM8K dataset, created entirely using Gretel Navigator with meta-llama/Meta-Llama-3.1-405B as the agent LLM. It contains Grade School-level reasoning tasks with step-by-step solutions, focusing on multi-step reasoning problems. Key Features: Synthetically Generated: Built using Gretel Navigator, leveraging evolutionary approach for diversity to create both the… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic-gsm8k-evolutionary-405b.textquestion-answering1K<n<10K5 likes113 downloads2y agoHugging Face23CJJones /LLM_Electrical_Engineering_Educational_Synthetic_DialogDataset Card for LLM_Electrical_Engineering_Educational_Synthetic_Dialog The full CJ Jones' synthetic dataset catalog is available at: https://datadeveloper1.gumroad.com Want more? 🚀 Get the AI Startup Bundle from Gumroad. Dataset Description The LLM_Electrical_Engineering_Educational_Synthetic_Dialog dataset contains AI-generated conversational interactions designed for training large language models in electrical engineering education. This synthetic dialogue corpus simulates tutor-student… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/LLM_Electrical_Engineering_Educational_Synthetic_Dialog.textquestion-answering10K<n<100K3 likes113 downloads7mo agoHugging Face24leo-bjpark /nmr-dataset-synthetic-retraction NMR Dataset: Synthetic Retraction This dataset contains 360 complete belief-revision episodes: 90 each for monotonic, nmr_new_evidence, nmr_retraction, and nmr_mixed. Each JSONL row is one episode, with a fixed dependency graph, an initial belief base, three revisions, and complete gold T/F/U states at all four checkpoints. It is intentionally not expanded into one row per proposition. An evaluator can construct a query for every proposition at any checkpoint from the episode… See the full description on the dataset page: https://huggingface.co/datasets/leo-bjpark/nmr-dataset-synthetic-retraction.question-answering0 likes113 downloads10d agoHugging Face25Kylan12 /synthetic-superconductor-materials-dataset synthetic-superconductor-materials-dataset Synthetic Q&A dataset on Superconductor Materials, generated with SDGS (Synthetic Dataset Generation Suite). Dataset Details Metric Value Topic Superconductor Materials Total Q&A Pairs 2649 Valid Pairs 2649 Provider/Model ollama/gpt-oss:120b Sources This dataset was generated from 170 scholarly papers: # Title Authors Year Source QA Pairs 1 Observation of a large-gap… See the full description on the dataset page: https://huggingface.co/datasets/Kylan12/synthetic-superconductor-materials-dataset.textquestion-answering1K<n<10K1 likes111 downloads7mo agoHugging Face26Taylor658 /synthetic-legal ⚖️ Synthetic Legal (Query, Response) Dataset 📚 140,000 synthetic (legal query, legal response) pairs across 13 legal domains, built to resemble the structure of real-world fact patterns and citation-backed answers. ⚠️ Disclaimer: All text is synthetically generated and IS NOT LEGALLY ACCURATE. Citations are real but assigned at random, and the verified_solution and verification_method columns are template labels, not evidence of review. This dataset is not legal advice.… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/synthetic-legal.texttext-generation100K<n<1M10 likes108 downloads12d agoHugging Face27Gabriel8 /tiny-llm-synthetic-qa Tiny-LLM: Synthetic Question-Answering Dataset Dataset Description This dataset was created for the fine-tuning stage of the Tiny-LLM Project, a project focused on training and evaluating compact language models from scratch. It contains 706,727 high-quality, synthetic multi-turn Question-Answering (Q&A) conversations in English, generated using the Gemini API. The dataset was designed to teach small models instruction-following capabilities across a diverse range of… See the full description on the dataset page: https://huggingface.co/datasets/Gabriel8/tiny-llm-synthetic-qa.textquestion-answering100K<n<1M2 likes101 downloads11mo agoHugging Face28nguyenkhanh87 /ViLegalQA-Synthetic-Curation ViLegalQA Synthetic Curation Dataset summary This repository releases the synthetic Vietnamese legal QA research artifacts produced in the accompanying study. The primary resource contains 10,095 synthetic QA items spanning true/false, multiple-choice, and open-ended tasks. It is accompanied by the final curation/quality annotations used in the study, plus aggregated labels for 600 items from the five-expert human calibration panel. Manuscript: Human-Calibrated… See the full description on the dataset page: https://huggingface.co/datasets/nguyenkhanh87/ViLegalQA-Synthetic-Curation.tabularquestion-answering10K<n<100K0 likes100 downloads20d agoHugging Face29liodon-ai /math-dow-mod-synthetic-v2 Cyclic Calendar Story Reasoning v2 v2 of liodon-ai/math-dow-mod-synthetic-v1, rebuilt around one piece of feedback from running v1 at scale: story and reasoning diversity has to scale to millions of rows, or it reads as one template repeated. v1's story setups were picked off a fixed list of 8 full sentences per task/direction, and every reasoning trace followed one fixed 4-slot skeleton (opener / op / mod-reduce / conclude) with only the wording swapped — fine at 10-50k rows… See the full description on the dataset page: https://huggingface.co/datasets/liodon-ai/math-dow-mod-synthetic-v2.texttext-generation1M<n<10M0 likes97 downloads27d agoHugging Face30liodon-ai /math-dow-mod-synthetic-v1 Math + Cyclic-Time Synthetic Dataset Synthetic dataset for training a small (~10M-100M param), task-specialized LLM on arithmetic (addition, multiplication), cyclic time arithmetic (days-of-week, months, 12-hour and 24-hour clock), and mod-k remainder probes — generalization-focused rather than memorization, following on from the T2/T5/T10 modular-circuit discussion. Days-of-week, months, and hours are all instances of the same underlying cyclic/modular-addition structure… See the full description on the dataset page: https://huggingface.co/datasets/liodon-ai/math-dow-mod-synthetic-v1.texttext-generation10K<n<100K0 likes90 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.