CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sailor2 /sea-synthetictext10M<n<100M0 likes3.2k downloads2y agoHugging Face02UCB-team /unclickbait-synthetic-27b-trajectories Unclickbait Synthetic 27B Trajectories Synthetic trajectory dataset generated by using rich structured JSON prompts and validated by two-stage judging pipeline. Contents : Full generated trajectories (current snapshot: 48,623 records out of 152,369 pristine event candidates). : 30 benchmark test samples audited end-to-end through the 122B two-stage judge (Stage 1 integrity gate + Stage 2 4D scoring). texttext-generationn<1K0 likes2.4k downloads10d agoHugging Face03MachineLearningLM /machinelearninglm-scm-synthetic-tabularml MachineLearningLM Pretraining Corpus This repository contains the pretraining corpus for MachineLearningLM, a framework designed to equip large language models (LLMs) with robust in-context machine learning (ML) capabilities. The dataset consists of ML tasks synthesized from millions of structural causal models (SCMs), spanning various shot counts up to 1,024. It is designed to enable LLMs to learn from many in-context examples on standard ML tasks purely via in-context learning… See the full description on the dataset page: https://huggingface.co/datasets/MachineLearningLM/machinelearninglm-scm-synthetic-tabularml.texttext-generation1M<n<10M4 likes857 downloads10mo agoHugging Face04devanshamin /synthetic-pii-function-calling Dataset Summary A function calling dataset created by filtering the urchade/synthetic-pii-ner-mistral-v1 dataset. texttext-generation1K<n<10K0 likes546 downloads2y agoHugging Face05aaaaliou /pi-synthetic Coding agent session traces for aaaaliou/pi-synthetic This dataset contains redacted coding agent session traces collected while working on git@github.com:aliou/pi-synthetic.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review. Data description Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line is a structured… See the full description on the dataset page: https://huggingface.co/datasets/aaaaliou/pi-synthetic.tabulartext-generationn<1K0 likes498 downloads5mo agoHugging Face06Kylan12 /Synthetic-AI-ML-Dataset Synthetic-AI-ML-Dataset Synthetic Q&A dataset on AI and Machine Learning Dataset Details Metric Value Topic AI and Machine Learning Total Q&A Pairs 14021 Valid Pairs 14021 Provider/Model ollama/gpt-oss:120b Generation Cost Metric Value Prompt Tokens 14,941,957 Completion Tokens 17,159,263 Total Tokens 32,101,220 GPU Energy 12.9628 kWh Sources This dataset was generated from 474 scholarly papers: #… See the full description on the dataset page: https://huggingface.co/datasets/Kylan12/Synthetic-AI-ML-Dataset.textquestion-answering10K<n<100K2 likes472 downloads6mo agoHugging Face07davidfoss /Synthetic-Causal-Reasoning-50k 🏭 Sovereign Synthetic Reasoning Dataset (400k) "High-Quality Chain-of-Thought Data at Scale." 📊 Overview This dataset contains 400,000 synthetic reasoning samples spanning 16 enterprise domains (Finance, Pharma, Legal, Cybersecurity, Supply Chain, etc.). It was generated using the Sovereign Generator, which produced 1.6 million samples and applied a strict quality filter (Top 25%) to retain only the most logically consistent and complex chains. Average Quality… See the full description on the dataset page: https://huggingface.co/datasets/davidfoss/Synthetic-Causal-Reasoning-50k.text100K<n<1M1 likes406 downloads9mo agoHugging Face08din0s /synthetic-beir-datatext1M<n<10M0 likes396 downloads3y agoHugging Face09BILGEM-AI /BILGE-Synthetic-Web BILGE-Synthetic-Web Dataset BILGE-Synthetic-Web was created following the methodology presented in the Cosmopedia blog/article. All content was generated using a 27B-parameter model. Further details on the methodology are available at: 🔗 https://huggingface.co/blog/cosmopedia text1M<n<10M9 likes338 downloads10mo agoHugging Face10ranausmans /synthetic-social-networks Synthetic Social Networks (Dataset) Raw experimental outputs from the Synthetic Social Networks study: 59,776 in-character LLM-agent posts from 528 production trials, and 64,562 posts total when the original pipeline-verification runs are included. The artifact combines an exploratory stage with a separately frozen, preregistered 448-trial matched-exposure confirmation. Each production trial includes peer-vote traces from in-character voting by other agents.… See the full description on the dataset page: https://huggingface.co/datasets/ranausmans/synthetic-social-networks.tabularother10K<n<100K1 likes337 downloads1mo agoHugging Face11zachnorton03 /synthetic-pre1930-sftTL;DR A vintage finetuning dataset (~416k rows, eleven task routes). Sourced by taking excerpts from pre-1930's texts, turning these into verbatim answers, and then using deepseek-chat to generate period-appropriate questions of those answers. Any model tuned on this dataset should, theoretically, never update its weights on anachronistic text, since questions are masked in the finetuning stages. Features composition, verse, narrative, reasoning, multiturn dialogue, and calibrated uncertainty… See the full description on the dataset page: https://huggingface.co/datasets/zachnorton03/synthetic-pre1930-sft.tabularquestion-answering100K<n<1M2 likes333 downloads1mo agoHugging Face12fineinstructions-pretraining /nemotron_synthetic_1T_exp If you use this project in your research please cite: @article{patel2025fineinstructions, title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale}, author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris}, year = {2026}, month = jan, day = {28}, } text100M<n<1B0 likes332 downloads8mo agoHugging Face13anywaylabs /synthetic-mvtec-ad-defect-detection Synthetic MVTec AD – Defect Detection Dataset by AnywayLabs.ai Need a custom synthetic dataset for your own defect detection use case? This dataset is an open-source sample of our synthetic data generation work at AnywayLabs. If you're working on: industrial defect detection visual inspection supervised anomaly detection hard-to-collect defect classes synthetic data for computer vision training You can request a custom synthetic dataset here, or email:… See the full description on the dataset page: https://huggingface.co/datasets/anywaylabs/synthetic-mvtec-ad-defect-detection.imageobject-detectionn<1K1 likes317 downloads4mo agoHugging Face14SZLHOLDINGS /oac-clinical-transport-observability-synthetic OAC Clinical Transport Observability — Synthetic This dataset contains 1,200 fixed-seed, entirely synthetic operational transport-health examples for the companion OAC System Health v1 model. It contains no records collected from a patient, laboratory, analyzer, instrument, LIS, EHR, network, or health-care site. Companion model: OAC System Health v1. Canonical source: szl-forge clinical gateway. Data boundary The closed schema contains only eight bounded… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/oac-clinical-transport-observability-synthetic.texttabular-classification1K<n<10K0 likes301 downloads3h agoHugging Face15ritaranx /clinical-synthetic-text-kg Data Description We release the synthetic data generated using the method described in the paper Knowledge-Infused Prompting: Assessing and Advancing Clinical Text Data Generation with Large Language Models (ACL 2024 Findings). The external knowledge we use is based on external knowledge graphs. Generated Datasets The original train/validation/test data, and the generated synthetic training data are listed as follows. For each dataset, we generate 5000 synthetic… See the full description on the dataset page: https://huggingface.co/datasets/ritaranx/clinical-synthetic-text-kg.texttext-classification1K<n<10K0 likes263 downloads2y agoHugging Face16hotdogsalesman /unit-price-evidence-synthetic Unit Price Evidence: Synthetic This dataset contains rendered synthetic shopping pages and evidence-pointer targets for product-card discovery and unit-price field extraction. It was built to warm-start small encoder-decoder models without redistributing retailer HTML, screenshots, product data, account data, or browsing history. Release Version: 0.1.0 Source code: erichasinternet/apples-to-apples Source manifest SHA-256:… See the full description on the dataset page: https://huggingface.co/datasets/hotdogsalesman/unit-price-evidence-synthetic.imageimage-text-to-text10K<n<100K0 likes261 downloads2mo agoHugging Face17koushikcs09 /mitre-attack-synthetic-scenarios MITRE ATT&CK Synthetic Scenario Logs v3.0 Expanded Dataset: 30 scenarios × 8 events = 240 synthetic events Axis Coverage Environment endpoint, cloud, SaaS, identity, CI/CD, OT/IoT Actor Type external_apt, ransomware, insider, compromised_vendor, careless_admin, automated_threat Intent exfiltration, impact, fraud, persistence, reconnaissance, cryptomining, espionage Detection Source EDR, IAM, SIEM, DLP, DNS, proxy, cloud_audit, email_gateway, CASB, NDR, PAM, firewall… See the full description on the dataset page: https://huggingface.co/datasets/koushikcs09/mitre-attack-synthetic-scenarios.textn<1K0 likes260 downloads4mo agoHugging Face18ajaxdavis /mobtranslate-kuku-yalanji-synthetic-corpus-v2 MobTranslate Kuku Yalanji Synthetic Research Corpus v2 Complete public research release of 20,047 synthetic English-Kuku Yalanji sentence pairs plus the process evidence needed to inspect their production, review, revision, split, and use in the MobTranslate model program. Project-reviewed synthetic research material pending fluent-speaker and elder verification. It is not a speaker-certified dictionary or translation corpus. Identity ISO 639-3: gvn Glottocode:… See the full description on the dataset page: https://huggingface.co/datasets/ajaxdavis/mobtranslate-kuku-yalanji-synthetic-corpus-v2.tabulartranslation10K<n<100K1 likes220 downloads2mo agoHugging Face19stighellemans /meddeid-english-synthetic-benchmark MedDeID English synthetic clinical benchmark This is a fixed, human-validated benchmark containing 300 synthetic English clinical documents: 150 en-GB and 150 en-US. It contains 1,717 primary PII spans and 7,358 confirmed core-PII subannotation segments. It contains no real patient notes or personal information. Use the entire test split only for final evaluation: from datasets import load_dataset benchmark = load_dataset( "stighellemans/meddeid-english-synthetic-benchmark"… See the full description on the dataset page: https://huggingface.co/datasets/stighellemans/meddeid-english-synthetic-benchmark.documenttoken-classificationn<1K0 likes214 downloads9d agoHugging Face20ContextReq /Synthetic-Dataset-Childrens-Stories**Status: released 13-09-2026, repacked 14-09-2026.** The 14-09-2026 repack replaced 58 items after the acceptance gates were strengthened (prompt-instruction leaks, markdown bullet lists and blockquotes); the other 29,942 are unchanged. Development stopped, pipeline released 17/09/26. SAMPLE RELEASE: 30,000 synthetic children's short stories for early-reader language modelling. Metrics Value genres 26 stories per genre 1.153-1.154K stories total characters 38… See the full description on the dataset page: https://huggingface.co/datasets/ContextReq/Synthetic-Dataset-Childrens-Stories.texttext-generation10K<n<100K1 likes210 downloads8d agoHugging Face21stighellemans /meddeid-dutch-synthetic-benchmark MedDeID Dutch synthetic benchmark This repository contains the fixed 300-document synthetic Dutch evaluation benchmark. It contains no real patient notes and must not be mixed into a training or validation partition when reporting MedDeID benchmark results. This is the openly shareable synthetic benchmark described in the manuscript. It is not the separate 300-note hospital benchmark, which contains personal information and is not publicly distributed. Subannotations… See the full description on the dataset page: https://huggingface.co/datasets/stighellemans/meddeid-dutch-synthetic-benchmark.documenttoken-classificationn<1K0 likes209 downloads9d agoHugging Face22BAAI /OpenSeek-Synthetic-Reasoning-Data-Examples OpenSeek-Reasoning-Data OpenSeek [Github|Blog] Recent reseach has demonstrated that the reasoning ability of LLMs originates from the pre-training stage, activated by RL training. Massive raw corpus containing complex human reasoning process, but lack of generalized and effective synthesis method to extract these reasoning process. News 🔥🔥🔥[2025/02/25] We publish some math, code, and general knowledge domain reasoning data synthesized from the current pipeline.… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/OpenSeek-Synthetic-Reasoning-Data-Examples.text1M<n<10M27 likes204 downloads2y agoHugging Face23keisuke-miyako /chatml-synthetic-2026-0420text1K<n<10K0 likes203 downloads5mo agoHugging Face24Hodfa71 /pstu-synthetic-secrets PSTU Synthetic Secrets Dataset Synthetic secrets benchmark for evaluating LLM memorization and unlearning, from the paper: Not All Secrets Are Equal: Type-Aware Unlearning for Language Model Secret Removal Hoda Fakhar — ECML PKDD 2026 Dataset Description 175 synthetic secrets across 25 types, each paired with 100 structurally similar decoys for computing the Carlini exposure metric. All data is synthetically generated. No real credentials, PII, or sensitive information… See the full description on the dataset page: https://huggingface.co/datasets/Hodfa71/pstu-synthetic-secrets.texttext-generationn<1K0 likes201 downloads6mo agoHugging Face25legal-hackathon-2024 /synthetictabular100K<n<1M0 likes197 downloads2y agoHugging Face26ritaranx /clinical-synthetic-text-llm Data Description We release the synthetic data generated using the method described in the paper Knowledge-Infused Prompting: Assessing and Advancing Clinical Text Data Generation with Large Language Models (ACL 2024 Findings). The external knowledge we use is based on LLM-generated topics and writing styles. Generated Datasets The original train/validation/test data, and the generated synthetic training data are listed as follows. For each dataset, we generate 5000… See the full description on the dataset page: https://huggingface.co/datasets/ritaranx/clinical-synthetic-text-llm.texttext-classification1K<n<10K3 likes195 downloads2y agoHugging Face27stighellemans /meddeid-english-synthetic-corpus MedDeID English synthetic clinical corpus This repository contains 6,700 synthetic English clinical documents with character-offset de-identification spans: 3,350 en-GB and 3,350 en-US documents. It contains no real patient notes or personal information. The complete corpus was used to train meddeid-english-synth. Split policy All records are exposed in one train split. There is no publisher-defined validation split. Users who tune a model must create and report… See the full description on the dataset page: https://huggingface.co/datasets/stighellemans/meddeid-english-synthetic-corpus.documenttoken-classification1K<n<10K0 likes186 downloads9d agoHugging Face28hoanganhphamqn /unclickbait-synthetic-27b-trajectoriestext100K<n<1M0 likes177 downloads15d agoHugging Face29aseifert /pie-synthetic PIE synthetic dataset Repo: https://github.com/awasthiabhijeet/PIE Paper: https://aclanthology.org/D19-1435.pdf text10M<n<100M1 likes172 downloads4y agoHugging Face30tohio /slm-synthetic-distillation-sft SLM Synthetic Distillation SFT Summary Synthetic response-distillation dataset containing prompt-response examples generated with a teacher model. Dataset Dataset type: response distillation Total rows: 14,717 Language: English Signal Distribution Signal Rows arithmetic 2,153 cloud 842 code 994 data_transform 1,078 database 1,190 debugging 1,331 educational_qa 2,124 factual_restraint 2,050 instruction 1… See the full description on the dataset page: https://huggingface.co/datasets/tohio/slm-synthetic-distillation-sft.text10K<n<100K0 likes161 downloads20d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.