CoolFace
26 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01starmpcc /Asclepius-Synthetic-Clinical-Notes Asclepius: Synthetic Clincal Notes & Instruction Dataset Dataset Summary This dataset is official dataset for Asclepius (arxiv) This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs. We first synthesized synthetic notes from PMC-Patients case reports with GPT-3.5 Then, we generate instruction-answer pairs for 157k synthetic discharge summaries Supported Tasks This dataset covers below 8 tasks Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/starmpcc/Asclepius-Synthetic-Clinical-Notes.textquestion-answering100K<n<1M117 likes961 downloads2y agoHugging Face02gretelai /synthetic_multilingual_llm_prompts Image generated by DALL-E. See prompt for more details 📝🌐 Synthetic Multilingual LLM Prompts Welcome to the "Synthetic Multilingual LLM Prompts" dataset! This comprehensive collection features 1,250 synthetic LLM prompts generated using Gretel Navigator, available in seven different languages. To ensure accuracy and diversity in prompts, and translation quality and consistency across the different languages, we employed Gretel Navigator both as a generation tool and as an… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_multilingual_llm_prompts.tabulartext-generation1K<n<10K11 likes115 downloads2y agoHugging Face03Taylor658 /synthetic-legal ⚖️ Synthetic Legal (Query, Response) Dataset 📚 140,000 synthetic (legal query, legal response) pairs across 13 legal domains, built to resemble the structure of real-world fact patterns and citation-backed answers. ⚠️ Disclaimer: All text is synthetically generated and IS NOT LEGALLY ACCURATE. Citations are real but assigned at random, and the verified_solution and verification_method columns are template labels, not evidence of review. This dataset is not legal advice.… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/synthetic-legal.texttext-generation100K<n<1M10 likes102 downloads14d agoHugging Face04dn-institute /cyberattack-blockchain-synth ELTEX-Blockchain: A Domain-Specific Dataset for Cybersecurity 🔐 12k Synthetic Social Media Messages for Early Cyberattack Detection on Blockchain Dataset Statistics Category Samples Description Cyberattack 6,941 Early warning signals and indicators of cyberattacks General 4,507 Regular blockchain discussions (non-security related) Dataset Structure Each entry in the dataset contains: message_id: Unique identifier for each message… See the full description on the dataset page: https://huggingface.co/datasets/dn-institute/cyberattack-blockchain-synth.texttext-classification10K<n<100K1 likes86 downloads2y agoHugging Face05Taylor658 /synthetic-fine-arts 🎨 Synthetic Fine Arts (Challenge, Solution) Dataset 🖼️ 225,000 synthetic (artistic challenge, proposed solution) pairs spanning Visual Arts, Performing Arts, Musical Arts, Literary Arts, Digital Arts, Art History, and Art Theory, with metadata fields that mirror the shape of a curation workflow. ⚠️ Disclaimer: All text is synthetically generated and should not be relied on for artistic, historical, or technical accuracy. The VerificationMethod, ReferenceMaterial, and… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/synthetic-fine-arts.texttext-generation100K<n<1M2 likes66 downloads14d agoHugging Face06Glebkaa /MedSyn-synthetic Synthetic dataset: generated.csv - synthetic datasets containing 41,185 clinical note samples spanning 219 ICD-10 codes. Data field Description idx Unique sample identifier. ICD-10 The targeted ICD-10 code used for prior data sampling. generation_model The model used for sample generation (GTP-3.5, GPT-4, LLaMA-7b, LLaMA-13b) prompt Prompt used for sample generation. prior Type of prior data used for sample generation. example Bool variable for the presence or… See the full description on the dataset page: https://huggingface.co/datasets/Glebkaa/MedSyn-synthetic.tabulartext-classification10K<n<100K2 likes48 downloads2y agoHugging Face07kiddothe2b /synthetic_polistance Fully Synthetic Prompts for LLM Political Stance Detection All resources developed in the article "Templated or fully Synthetic? Prompt construction as a confound in measuring LLM political stance beyond writing assistance" (Chalkidis, 2026). Paper Abstract Political stance detection in LLMs has long been dominated by closed-ended, multiple-choice political survey questions—originally designed for humans, and thus lacks the realism and nuance of human-AI… See the full description on the dataset page: https://huggingface.co/datasets/kiddothe2b/synthetic_polistance.tabulartext-generation1K<n<10K0 likes48 downloads2mo agoHugging Face08bluesky333 /synthetic_discharge_summ Asclepius: Synthetic Clincal Notes & Instruction Dataset Dataset Summary This dataset is a subset of the dataset for Asclepius model (arxiv). The original dataset is made up of synthetic notes generated from PMC-Patients case reports with GPT-3.5. We filtered the summarization task for discharge notes. The dataset contains 13,584 notes. Supported Tasks This dataset covers below summarization task Languages English Dataset… See the full description on the dataset page: https://huggingface.co/datasets/bluesky333/synthetic_discharge_summ.textquestion-answeringn<1K1 likes41 downloads2y agoHugging Face09shekar-ai /SynPerForm-synthetic-persian-formality-pairs SynPerForm SynPerForm is a paired Persian dataset for formality style transfer. Each informal text is paired with a freely written formal rewrite that preserves its meaning without requiring lexical or structural equivalence. The formal rewrites were generated with OpenAI GPT-5.6 Luna. Columns Informal: the original informal Persian text. Formal: a free formal rewrite that preserves the original meaning. Intended uses Formality style transfer:… See the full description on the dataset page: https://huggingface.co/datasets/shekar-ai/SynPerForm-synthetic-persian-formality-pairs.texttext-generation100K<n<1M1 likes40 downloads2mo agoHugging Face10adwaith06 /indic-synthetic-profiles 🇮🇳 Indian Synthetic Identity Dataset 10,000 realistic Indian synthetic identities across 8 languages — generated by indic-faker Dataset Description This dataset contains 10,000 rows of realistic, synthetic Indian identity data generated using the indic-faker Python library. Every record is algorithmically valid — Aadhaar numbers pass Verhoeff checksum verification, GSTINs have correct state codes, and names are culturally authentic across 8 Indian languages.… See the full description on the dataset page: https://huggingface.co/datasets/adwaith06/indic-synthetic-profiles.tabulartext-generation10K<n<100K0 likes33 downloads6mo agoHugging Face11maharshipandya /synthanime-openhermes2.5 What is this dataset? This is a hybrid Instruction-Response dataset (scraped + synthetic) for anime synopses. Given around 10,000 scraped anime synopses, the user instructions for this dataset were generated using Teknium's Openhermes 2.5 The assistant column consists of synopsis for different animes (which were previously scraped) The user column consists of the instructions that "might have" generated the synopsis (synthetically generated) The goal of this dataset is: to be used… See the full description on the dataset page: https://huggingface.co/datasets/maharshipandya/synthanime-openhermes2.5.texttext-generation10K<n<100K2 likes29 downloads3y agoHugging Face12prithivMLmods /Synthetic-Context-Conversations Synthetic-Context-Conversations Overview The Synthetic-Context-Conversations dataset is a collection of synthetic conversations designed to simulate empathetic and context-rich dialogues. It is particularly useful for tasks such as text generation, summarization, and question answering. The dataset is available in English and contains between 10,000 to 100,000 entries. Dataset Details Modalities: Text Languages: English Size: 10K-100K Formats: Parquet License:… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Synthetic-Context-Conversations.texttext-generation10K<n<100K3 likes25 downloads2y agoHugging Face13WT-solutions /BG-GSM-Synthetic Overview This dataset aims to create a bulgarian version of the GSM-Symbolic dataset Code to generate samples is available at our Github BG-GSM-Symbolic Synthetic Dataset Generation for Large Language Model evaluation. The dataset is generated in the style of GSM-8k, thus it contains relatively easy mathematical questions that a bright grade school student should be able to answer. This project is intended for automatic evaluation of LLMs, which have already been trained… See the full description on the dataset page: https://huggingface.co/datasets/WT-solutions/BG-GSM-Synthetic.texttext-generationn<1K2 likes22 downloads6mo agoHugging Face14elizaveta-dev /Referencing_Errors_Synthetic_EN Synthetic Dataset for Automatic Error Correction in Referencing This dataset includes 4,600 parallel sentences for Automatic Error Correction in referencing in German.It was synthetically created with gpt-4o-mini model according to the Institutional Guidelines of the Center for Translation Studies (CTS), University of Vienna. Dataset Description corrupted_sentence: the sentence containing the referencing error clean_sentence: the correct version of the corrupted sentence… See the full description on the dataset page: https://huggingface.co/datasets/elizaveta-dev/Referencing_Errors_Synthetic_EN.texttext-generation1K<n<10K0 likes20 downloads10mo agoHugging Face15Xavier1234 /Asclepius-Synthetic-Clinical-Notes Asclepius: Synthetic Clincal Notes & Instruction Dataset Dataset Summary This dataset is official dataset for Asclepius (arxiv) This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs. We first synthesized synthetic notes from PMC-Patients case reports with GPT-3.5 Then, we generate instruction-answer pairs for 157k synthetic discharge summaries Supported Tasks This dataset covers below 8 tasks Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/Xavier1234/Asclepius-Synthetic-Clinical-Notes.textquestion-answering100K<n<1M1 likes20 downloads5mo agoHugging Face16TigreGotico /arabic-mantoq-synthetic-g2ptexttext-generation1M<n<10M0 likes19 downloads9mo agoHugging Face17cmp81 /cyberattack-blockchain-synth ELTEX-Blockchain: A Domain-Specific Dataset for Cybersecurity 🔐 12k Synthetic Social Media Messages for Early Cyberattack Detection on Blockchain Dataset Statistics Category Samples Description Cyberattack 6,941 Early warning signals and indicators of cyberattacks General 4,507 Regular blockchain discussions (non-security related) Dataset Structure Each entry in the dataset contains: message_id: Unique identifier for each message… See the full description on the dataset page: https://huggingface.co/datasets/cmp81/cyberattack-blockchain-synth.texttext-classification10K<n<100K0 likes17 downloads1mo agoHugging Face18ping98k /synthetic-thai-follow-up-questiontexttext-generationn<1K0 likes16 downloads2y agoHugging Face19nadizik /3D_Synthetic_Petroleum_Derived_GEMS 3D Synthetic Petroleum-Derived GEMS (MVP Release) 📌 Dataset Overview This dataset contains 197 elite, highly complex 3D molecular structures derived from petroleum fractions. Designed specifically for petrochemicals, materials science, organic semiconductors, and specialized additives, these compounds represent a curated "Golden Fund" of stable, complex hydrocarbons. Unlike drug-like molecules, this dataset focuses on polycyclic architectures, rigid 3-ring… See the full description on the dataset page: https://huggingface.co/datasets/nadizik/3D_Synthetic_Petroleum_Derived_GEMS.texttabular-classificationn<1K0 likes16 downloads2mo agoHugging Face20LampsteR /Asclepius-Synthetic-Clinical-Notes Asclepius: Synthetic Clincal Notes & Instruction Dataset Dataset Summary This dataset is official dataset for Asclepius (arxiv) This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs. We first synthesized synthetic notes from PMC-Patients case reports with GPT-3.5 Then, we generate instruction-answer pairs for 157k synthetic discharge summaries Supported Tasks This dataset covers below 8 tasks Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/LampsteR/Asclepius-Synthetic-Clinical-Notes.textquestion-answering100K<n<1M0 likes14 downloads10mo agoHugging Face21KirillNik /tweepfake_synthetic TweepFake Synthetic Tweets Synthetic tweets generated by three open / API LLMs, conditioned on topics extracted from real human tweets. Built as a controlled-variable corpus for studying — and detecting — machine-generated social-media text. The design holds the topic constant across conditions: every topic is fed to multiple models under multiple prompting strategies, so differences in the output can be attributed to the model and prompt, not the subject matter.… See the full description on the dataset page: https://huggingface.co/datasets/KirillNik/tweepfake_synthetic.texttext-generation10K<n<100K1 likes13 downloads4mo agoHugging Face22shak9345 /hopepet-ai-synthetic-dataset HOPEPET AI Synthetic Pet-Care Dataset This dataset was created for the HOPEPET AI final project. HOPEPET AI is an AI-based pet-care prototype for dog and cat owners.The goal of the system is to provide responsible first-step guidance when a user notices a pet-related problem, symptom, or behavior change. The system does not provide veterinary diagnosis, does not replace a veterinarian, and does not recommend medication or dosages. Dataset Description The dataset… See the full description on the dataset page: https://huggingface.co/datasets/shak9345/hopepet-ai-synthetic-dataset.texttext-classification10K<n<100K0 likes13 downloads3mo agoHugging Face23TigreGotico /portuguese-sentences-synthetic-g2p Dataset Card for 'TigreGotico/portuguese_g2p' Dataset Description Dataset Summary TigreGotico/portuguese_g2p is a Grapheme-to-Phoneme (G2P) dataset for Portuguese, offering phonetic transcriptions for sentences across ten different regional variants. It is derived from the portuguese_phonetic_lexicon and is designed to aid in the development of robust Speech Recognition (ASR) and Text-to-Speech (TTS) models that account for dialectal variation in Portuguese.… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/portuguese-sentences-synthetic-g2p.texttext-classification10K<n<100K0 likes9 downloads1y agoHugging Face24elizaveta-dev /Referencing_Errors_Synthetic_DE Synthetic Dataset for Automatic Error Correction in Referencing This dataset includes 4,600 parallel sentences for Automatic Error Correction in referencing in German.It was synthetically created with gpt-4o-mini model according to the Institutional Guidelines of the Center for Translation Studies (CTS), University of Vienna. Dataset Description corrupted_sentence: the sentence containing the referencing error clean_sentence: the correct version of the corrupted sentence… See the full description on the dataset page: https://huggingface.co/datasets/elizaveta-dev/Referencing_Errors_Synthetic_DE.texttext-generation1K<n<10K0 likes9 downloads10mo agoHugging Face25dkhundley /synthetic-knowledge-itemstexttext-classificationn<1K0 likes7 downloads2y agoHugging Face26pixelbombe /synth-dent-testtexttext-generationn<1K0 likes5 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.