CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01efficient-deep-research /synthesized_datasettext10K<n<100K0 likes34k downloads1y agoHugging Face02lhbit20010120 /Jiuzhang3.0_synthtext1M<n<10M0 likes4.1k downloads2y agoHugging Face03sailor2 /sea-synthetictext10M<n<100M0 likes3.3k downloads2y agoHugging Face04julien-c /synthtraces SynthTraces A minimal codebase to generate synthetic coding agent session traces using Pi. Each session pairs two models working inside one of the project codebases: a remotely hosted open model (e.g. deepseek-ai/DeepSeek-V4-Pro, openai/gpt-oss-120b, Qwen/Qwen3.6-27B) backs the coding agent, equipped with the default Pi tools — read, write, edit, and bash; a local model running in llama.cpp plays the user, opening with one of the starting questions and driving the… See the full description on the dataset page: https://huggingface.co/datasets/julien-c/synthtraces.tabular1K<n<10K32 likes3k downloads4mo agoHugging Face05electricsheepafrica /africa-synth-aid-flows-medical-multimodal-fracture-all Africa Synth Aid Flows Medical Multimodal Fracture All | Africa (Electric Sheep Africa metadata inventory) Size category: 1K<n<10K - Formats: json - Sector: health - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers Health… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-aid-flows-medical-multimodal-fracture-all.imagetabular-classification1K<n<10K5 likes2.8k downloads2mo agoHugging Face06UCB-team /unclickbait-synthetic-27b-trajectories Unclickbait Synthetic 27B Trajectories Synthetic trajectory dataset generated by using rich structured JSON prompts and validated by two-stage judging pipeline. Contents : Full generated trajectories (current snapshot: 48,623 records out of 152,369 pristine event candidates). : 30 benchmark test samples audited end-to-end through the 122B two-stage judge (Stage 1 integrity gate + Stage 2 4D scoring). texttext-generationn<1K0 likes2.7k downloads12d agoHugging Face07semran1 /synth-cc-unfilteredtext100M<n<1B1 likes978 downloads1y agoHugging Face08MachineLearningLM /machinelearninglm-scm-synthetic-tabularml MachineLearningLM Pretraining Corpus This repository contains the pretraining corpus for MachineLearningLM, a framework designed to equip large language models (LLMs) with robust in-context machine learning (ML) capabilities. The dataset consists of ML tasks synthesized from millions of structural causal models (SCMs), spanning various shot counts up to 1,024. It is designed to enable LLMs to learn from many in-context examples on standard ML tasks purely via in-context learning… See the full description on the dataset page: https://huggingface.co/datasets/MachineLearningLM/machinelearninglm-scm-synthetic-tabularml.texttext-generation1M<n<10M4 likes814 downloads10mo agoHugging Face09instruction-pretrain /ft-instruction-synthesizer-collection Instruction Pre-Training: Language Models are Supervised Multitask Learners (EMNLP 2024) This repo contains the fine-tuning data collection for the context-based instruction synthesizer used in our paper Instruction Pre-Training: Language Models are Supervised Multitask Learners. We explore supervised multitask pre-training by proposing Instruction Pre-Training, a framework that scalably augments massive raw corpora with instruction-response pairs to pre-train language models. The… See the full description on the dataset page: https://huggingface.co/datasets/instruction-pretrain/ft-instruction-synthesizer-collection.texttext-classification100K<n<1M62 likes549 downloads7mo agoHugging Face10devanshamin /synthetic-pii-function-calling Dataset Summary A function calling dataset created by filtering the urchade/synthetic-pii-ner-mistral-v1 dataset. texttext-generation1K<n<10K0 likes531 downloads2y agoHugging Face11aaaaliou /pi-synthetic Coding agent session traces for aaaaliou/pi-synthetic This dataset contains redacted coding agent session traces collected while working on git@github.com:aliou/pi-synthetic.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review. Data description Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line is a structured… See the full description on the dataset page: https://huggingface.co/datasets/aaaaliou/pi-synthetic.tabulartext-generationn<1K0 likes498 downloads5mo agoHugging Face12dougalldeepmind /2026-09-10-delib-synth Deliberative SFT: delib; native Qwen reasoning, best-of-4 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 2 runs >= 7). 658 rows from the full run; the 50 prompts it rejected were re-run in 2 pass(es) with 8 candidates per round under the amended constitution, recovering 42; 8 remain rejected field value experiment Deliberative SFT: delib; native Qwen reasoning, best-of-4 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-10-delib-synth.tabular1K<n<10K0 likes490 downloads15d agoHugging Face13Kylan12 /Synthetic-AI-ML-Dataset Synthetic-AI-ML-Dataset Synthetic Q&A dataset on AI and Machine Learning Dataset Details Metric Value Topic AI and Machine Learning Total Q&A Pairs 14021 Valid Pairs 14021 Provider/Model ollama/gpt-oss:120b Generation Cost Metric Value Prompt Tokens 14,941,957 Completion Tokens 17,159,263 Total Tokens 32,101,220 GPU Energy 12.9628 kWh Sources This dataset was generated from 474 scholarly papers: #… See the full description on the dataset page: https://huggingface.co/datasets/Kylan12/Synthetic-AI-ML-Dataset.textquestion-answering10K<n<100K2 likes473 downloads6mo agoHugging Face14SZLHOLDINGS /oac-clinical-transport-observability-synthetic OAC Clinical Transport Observability — Synthetic This dataset contains 1,200 fixed-seed, entirely synthetic operational transport-health examples for the companion OAC System Health v1 model. It contains no records collected from a patient, laboratory, analyzer, instrument, LIS, EHR, network, or health-care site. Companion model: OAC System Health v1. Canonical source: szl-forge clinical gateway. Data boundary The closed schema contains only eight bounded… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/oac-clinical-transport-observability-synthetic.texttabular-classification1K<n<10K0 likes393 downloads2d agoHugging Face15davidfoss /Synthetic-Causal-Reasoning-50k 🏭 Sovereign Synthetic Reasoning Dataset (400k) "High-Quality Chain-of-Thought Data at Scale." 📊 Overview This dataset contains 400,000 synthetic reasoning samples spanning 16 enterprise domains (Finance, Pharma, Legal, Cybersecurity, Supply Chain, etc.). It was generated using the Sovereign Generator, which produced 1.6 million samples and applied a strict quality filter (Top 25%) to retain only the most logically consistent and complex chains. Average Quality… See the full description on the dataset page: https://huggingface.co/datasets/davidfoss/Synthetic-Causal-Reasoning-50k.text100K<n<1M1 likes388 downloads9mo agoHugging Face16din0s /synthetic-beir-datatext1M<n<10M0 likes385 downloads3y agoHugging Face17SynthForensics /SynthForensicsgatedSynthForensics: Benchmarking and Evaluating People-Centric Synthetic Video Deepfakes Official Repository for the SynthForensics (SF) Benchmark Abstract Modern T2V/I2V generators synthesize people increasingly hard to distinguish from authentic footage, while current evaluation suites lag: legacy benchmarks target manipulation-based forgeries, and recent synthetic-video benchmarks prioritize scale over realistic human depiction. We introduce SynthForensics, a… See the full description on the dataset page: https://huggingface.co/datasets/SynthForensics/SynthForensics.textvideo-classification10K<n<100K0 likes384 downloads5mo agoHugging Face18dougalldeepmind /2026-07-29-synthdoc-approved-constitution-sft Dialogue dataset: Synthetic SFT corpus generated by synthdoc from the approved constitution, to regenerate the difficult-advice training data against a revised specification at a scale comparable to v1, so that the constitution is the intended difference between the two datasets. 1,443 documents / 1,531,369 Qwen3 tokens across five sub-corpora, matching v1's 1.52M. Required metadata field value experiment Synthetic SFT corpus generated by synthdoc from… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-29-synthdoc-approved-constitution-sft.text1K<n<10K0 likes366 downloads1mo agoHugging Face19BILGEM-AI /BILGE-Synthetic-Web BILGE-Synthetic-Web Dataset BILGE-Synthetic-Web was created following the methodology presented in the Cosmopedia blog/article. All content was generated using a 27B-parameter model. Further details on the methodology are available at: 🔗 https://huggingface.co/blog/cosmopedia text1M<n<10M9 likes364 downloads10mo agoHugging Face20fineinstructions-pretraining /nemotron_synthetic_1T_exp If you use this project in your research please cite: @article{patel2025fineinstructions, title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale}, author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris}, year = {2026}, month = jan, day = {28}, } text100M<n<1B0 likes337 downloads8mo agoHugging Face21ranausmans /synthetic-social-networks Synthetic Social Networks (Dataset) Raw experimental outputs from the Synthetic Social Networks study: 59,776 in-character LLM-agent posts from 528 production trials, and 64,562 posts total when the original pipeline-verification runs are included. The artifact combines an exploratory stage with a separately frozen, preregistered 448-trial matched-exposure confirmation. Each production trial includes peer-vote traces from in-character voting by other agents.… See the full description on the dataset page: https://huggingface.co/datasets/ranausmans/synthetic-social-networks.tabularother10K<n<100K1 likes335 downloads1mo agoHugging Face22n4ze3m /typed-decisions-synth Typed Decisions Synth This is the synthetic dataset I made for Hmm, a small open model that answers questions about your data with probabilities instead of text. It has 7,414 cases with 25,859 questions across 149 domains and workflows. Every question has an answer and a soft label (a probability for every option), so you can train a model to be unsure when it should be. Code and the model: github.com/n4ze3m/hmm Note: Everything here is written and labelled by an LLM. Nobody… See the full description on the dataset page: https://huggingface.co/datasets/n4ze3m/typed-decisions-synth.texttext-classification1K<n<10K2 likes325 downloads6d agoHugging Face23anywaylabs /synthetic-mvtec-ad-defect-detection Synthetic MVTec AD – Defect Detection Dataset by AnywayLabs.ai Need a custom synthetic dataset for your own defect detection use case? This dataset is an open-source sample of our synthetic data generation work at AnywayLabs. If you're working on: industrial defect detection visual inspection supervised anomaly detection hard-to-collect defect classes synthetic data for computer vision training You can request a custom synthetic dataset here, or email:… See the full description on the dataset page: https://huggingface.co/datasets/anywaylabs/synthetic-mvtec-ad-defect-detection.imageobject-detectionn<1K1 likes318 downloads4mo agoHugging Face24zachnorton03 /synthetic-pre1930-sftTL;DR A vintage finetuning dataset (~416k rows, eleven task routes). Sourced by taking excerpts from pre-1930's texts, turning these into verbatim answers, and then using deepseek-chat to generate period-appropriate questions of those answers. Any model tuned on this dataset should, theoretically, never update its weights on anachronistic text, since questions are masked in the finetuning stages. Features composition, verse, narrative, reasoning, multiturn dialogue, and calibrated uncertainty… See the full description on the dataset page: https://huggingface.co/datasets/zachnorton03/synthetic-pre1930-sft.tabularquestion-answering100K<n<1M2 likes313 downloads1mo agoHugging Face25puruchinera /FacturaRD-Synth Facturas DGII sintéticas Dataset de facturas dominicanas completamente sintéticas para entrenamiento y evaluación de extracción fiscal y OCR con modelos multimodales como Florence-2. El objetivo es entrenar modelos capaces de recibir una imagen con una o varias facturas y producir simultáneamente: registros fiscales estructurados; una transcripción OCR del contenido visible. Formato Cada fila de train.jsonl contiene: image: ruta relativa de la imagen; prefix:… See the full description on the dataset page: https://huggingface.co/datasets/puruchinera/FacturaRD-Synth.imageimage-to-text10K<n<100K0 likes289 downloads27d agoHugging Face26RobinSta /SynthPAI Dataset Card for SynthPAI SynthPAI was created to provide a dataset that can be used to investigate the personal attribute inference (PAI) capabilities of LLM on online texts. Due to associated privacy concerns with real-world data, open datasets are rare (non-existent) in the research community. SynthPAI is a synthetic dataset that aims to fill this gap. Dataset Details Dataset Description SynthPAI was created using 300 GPT-4 agents seeded with individual… See the full description on the dataset page: https://huggingface.co/datasets/RobinSta/SynthPAI.textzero-shot-classification1K<n<10K20 likes287 downloads2y agoHugging Face27ritaranx /clinical-synthetic-text-kg Data Description We release the synthetic data generated using the method described in the paper Knowledge-Infused Prompting: Assessing and Advancing Clinical Text Data Generation with Large Language Models (ACL 2024 Findings). The external knowledge we use is based on external knowledge graphs. Generated Datasets The original train/validation/test data, and the generated synthetic training data are listed as follows. For each dataset, we generate 5000 synthetic… See the full description on the dataset page: https://huggingface.co/datasets/ritaranx/clinical-synthetic-text-kg.texttext-classification1K<n<10K0 likes269 downloads2y agoHugging Face28hotdogsalesman /unit-price-evidence-synthetic Unit Price Evidence: Synthetic This dataset contains rendered synthetic shopping pages and evidence-pointer targets for product-card discovery and unit-price field extraction. It was built to warm-start small encoder-decoder models without redistributing retailer HTML, screenshots, product data, account data, or browsing history. Release Version: 0.1.0 Source code: erichasinternet/apples-to-apples Source manifest SHA-256:… See the full description on the dataset page: https://huggingface.co/datasets/hotdogsalesman/unit-price-evidence-synthetic.imageimage-text-to-text10K<n<100K0 likes263 downloads2mo agoHugging Face29koushikcs09 /mitre-attack-synthetic-scenarios MITRE ATT&CK Synthetic Scenario Logs v3.0 Expanded Dataset: 30 scenarios × 8 events = 240 synthetic events Axis Coverage Environment endpoint, cloud, SaaS, identity, CI/CD, OT/IoT Actor Type external_apt, ransomware, insider, compromised_vendor, careless_admin, automated_threat Intent exfiltration, impact, fraud, persistence, reconnaissance, cryptomining, espionage Detection Source EDR, IAM, SIEM, DLP, DNS, proxy, cloud_audit, email_gateway, CASB, NDR, PAM, firewall… See the full description on the dataset page: https://huggingface.co/datasets/koushikcs09/mitre-attack-synthetic-scenarios.textn<1K0 likes252 downloads4mo agoHugging Face30schneiderkamplab /sapient-synth-tasksource-reclor sapient-synth-tasksource-reclor Chat-template-ready synthetic anonymous replacement examples for one Sapient source excluded from the DFM5 data mix. Contents Format: gzip-compressed JSON Lines under data/train.jsonl.gz Schema: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]} Files: 1 Rows: 4633 Task: synthetic anonymous instruction replacement Generation Rows were generated with google/gemma-4-31B-it and… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/sapient-synth-tasksource-reclor.text1K<n<10K0 likes246 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.