datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
synthiaThe Synthia Dataset is a continuously growing aggregate of validated synthetic explanations of subjects picked from Claude Opus latent space based on varying esotericity in a large general list of technical fields. The explanations are varying in their target audience, level of detail and abstraction while incentivized to target Claude3-grade quality. The Synthia subnet leverages Commune's incentives to create a permissionless mining market around distilling knowledge out of SOTA closed-source… See the full description on the dataset page: https://huggingface.co/datasets/agicommies/synthia.lc_quad_synth
LC-QuAD 2.0-synth
Dataset Summary
This dataset is an updated version of the LC-QuAD 2.0 dataset which includes LLM-based natural language translations of the corresponding wikidata queries. It also includes
verifier scores for the LLM translations and the original translations indicating the probability that the translation is correct (for details see our linked GitHub Repository).
It contains 19000 examples of queries and translations. It can be used for training and… See the full description on the dataset page: https://huggingface.co/datasets/timschwa/lc_quad_synth.Asclepius-Synthetic-Clinical-Notes
Asclepius: Synthetic Clincal Notes & Instruction Dataset
Dataset Summary
This dataset is official dataset for Asclepius (arxiv)
This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs.
We first synthesized synthetic notes from PMC-Patients case reports with GPT-3.5
Then, we generate instruction-answer pairs for 157k synthetic discharge summaries
Supported Tasks
This dataset covers below 8 tasks
Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/starmpcc/Asclepius-Synthetic-Clinical-Notes.synthetic_multilingual_llm_prompts
Image generated by DALL-E. See prompt for more details
📝🌐 Synthetic Multilingual LLM Prompts
Welcome to the "Synthetic Multilingual LLM Prompts" dataset! This comprehensive collection features 1,250 synthetic LLM prompts generated using Gretel Navigator, available in seven different languages. To ensure accuracy and diversity in prompts, and translation quality and consistency across the different languages, we employed Gretel Navigator both as a generation tool and as an… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_multilingual_llm_prompts.synthetic-legal
⚖️ Synthetic Legal (Query, Response) Dataset
📚 140,000 synthetic (legal query, legal response) pairs across 13 legal domains, built to resemble the structure of real-world fact patterns and citation-backed answers.
⚠️ Disclaimer: All text is synthetically generated and IS NOT LEGALLY ACCURATE. Citations are real but assigned at random, and the verified_solution and verification_method columns are template labels, not evidence of review. This dataset is not legal advice.… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/synthetic-legal.synthetic-fine-arts
🎨 Synthetic Fine Arts (Challenge, Solution) Dataset
🖼️ 225,000 synthetic (artistic challenge, proposed solution) pairs spanning Visual Arts, Performing Arts, Musical Arts, Literary Arts, Digital Arts, Art History, and Art Theory, with metadata fields that mirror the shape of a curation workflow.
⚠️ Disclaimer: All text is synthetically generated and should not be relied on for artistic, historical, or technical accuracy. The VerificationMethod, ReferenceMaterial, and… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/synthetic-fine-arts.synthetic-enterprise-operations-pack
Solstice Synthetic Enterprise Operations Pack (Sample)
A curated synthetic internal company dataset spanning engineering, task systems, collaboration, CRM, support, incidents, documents, and account-health workflows. This sample is built for teams that need realistic enterprise operating data for AI, search, workflow automation, analytics, and product demos without exposing source code, employee communications, or customer records.
Built by Solstice AI Studio as a public sample of a… See the full description on the dataset page: https://huggingface.co/datasets/solsticestudioai/synthetic-enterprise-operations-pack.synthetic-confidential-information-injected-business-excerpts
Synthetic Confidential Information Injected Business Excerpts
This dataset aims to provide business report excerpts which contain relevant confidential/sensitive information.
This includes mentions of :
1. Internal Marketing Strategies.
2. Proprietary Product Composition.
3. License Internals.
4. Internal Sales Projections.
5. Confidential Patent Details.
6. others.
The dataset contains around 1k business excerpt - Reasons pairs. The Reason field contains the… See the full description on the dataset page: https://huggingface.co/datasets/Rohit-D/synthetic-confidential-information-injected-business-excerpts.synthetic_discharge_summ
Asclepius: Synthetic Clincal Notes & Instruction Dataset
Dataset Summary
This dataset is a subset of the dataset for Asclepius model (arxiv).
The original dataset is made up of synthetic notes generated from PMC-Patients case reports with GPT-3.5.
We filtered the summarization task for discharge notes. The dataset contains 13,584 notes.
Supported Tasks
This dataset covers below summarization task
Languages
English
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/bluesky333/synthetic_discharge_summ.albanian-synthetic
Albanian Synthetic Q&A Dataset
NOTE: DATASET MAY NOT BE WITH ACCURATE INFORMATION, AS IT IS AI-GENERATED.
A high-quality synthetic dataset of Albanian question-answer pairs covering diverse topics, generated using Mistral-Medium via the Le Platforme API.
⚠️ Note: Code for synthetic data generation can be found in the project repository.
Dataset Details
Language: Albanian (Shqip)
Format: CSV (UTF-8 encoded)
Columns:
Prompt: Question in Albanian
Response: Detailed… See the full description on the dataset page: https://huggingface.co/datasets/LTS-VVE/albanian-synthetic.synthetic_text_to_sql_th
Synthetic Text-to-SQL Thai Dataset
Thai translation of the gretelai/synthetic_text_to_sql dataset.
Dataset Description
This dataset contains Thai translations of synthetic text-to-SQL examples covering various domains and SQL patterns.
Source
Original Dataset: gretelai/synthetic_text_to_sql
Created by: Gretel.ai
Statistics
Split
Rows
Train
100,000
Test
5,851
Total
105,851
Columns
Column
Description… See the full description on the dataset page: https://huggingface.co/datasets/Porameht/synthetic_text_to_sql_th.synthetic-enterprise-ops-pack-sample
Solstice Synthetic Enterprise Operations Pack (Sample)
A multi-system graph dataset for agent evaluation and RAG benchmarking. This dataset simulates the interconnected operations of a modern technology company, linking sales activities, engineering workflows, IT support, and internal communications.
Built by Solstice AI Studio as a free sample of a larger commercial pack. 100% synthetic — no real company or employee data.
What's in the box
This dataset consists of 32… See the full description on the dataset page: https://huggingface.co/datasets/solsticestudioai/synthetic-enterprise-ops-pack-sample.Synthetic-Context-Conversations
Synthetic-Context-Conversations
Overview
The Synthetic-Context-Conversations dataset is a collection of synthetic conversations designed to simulate empathetic and context-rich dialogues. It is particularly useful for tasks such as text generation, summarization, and question answering. The dataset is available in English and contains between 10,000 to 100,000 entries.
Dataset Details
Modalities: Text
Languages: English
Size: 10K-100K
Formats: Parquet
License:… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Synthetic-Context-Conversations.Asclepius-Synthetic-Clinical-Notes
Asclepius: Synthetic Clincal Notes & Instruction Dataset
Dataset Summary
This dataset is official dataset for Asclepius (arxiv)
This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs.
We first synthesized synthetic notes from PMC-Patients case reports with GPT-3.5
Then, we generate instruction-answer pairs for 157k synthetic discharge summaries
Supported Tasks
This dataset covers below 8 tasks
Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/Xavier1234/Asclepius-Synthetic-Clinical-Notes.Asclepius-Synthetic-Clinical-Notes
Asclepius: Synthetic Clincal Notes & Instruction Dataset
Dataset Summary
This dataset is official dataset for Asclepius (arxiv)
This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs.
We first synthesized synthetic notes from PMC-Patients case reports with GPT-3.5
Then, we generate instruction-answer pairs for 157k synthetic discharge summaries
Supported Tasks
This dataset covers below 8 tasks
Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/LampsteR/Asclepius-Synthetic-Clinical-Notes.synthiaThe Synthia Dataset is a continuously growing aggregate of validated synthetic explanations of subjects picked from Claude Opus latent space based on varying esotericity in a large general list of technical fields. The explanations are varying in their target audience, level of detail and abstraction while incentivized to target Claude3-grade quality. The Synthia subnet leverages Commune's incentives to create a permissionless mining market around distilling knowledge out of SOTA closed-source… See the full description on the dataset page: https://huggingface.co/datasets/mubasit/synthia.frc-rules-syntheticFRC manual specific synthetic dataset!
synthiaThe Synthia Dataset is a continuously growing aggregate of validated synthetic explanations of subjects picked from Claude Opus latent space based on varying esotericity in a large general list of technical fields. The explanations are varying in their target audience, level of detail and abstraction while incentivized to target Claude3-grade quality. The Synthia subnet leverages Commune's incentives to create a permissionless mining market around distilling knowledge out of SOTA closed-source… See the full description on the dataset page: https://huggingface.co/datasets/lanba1/synthia.
