datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Synthetic-AI-ML-Dataset
Synthetic-AI-ML-Dataset
Synthetic Q&A dataset on AI and Machine Learning
Dataset Details
Metric
Value
Topic
AI and Machine Learning
Total Q&A Pairs
14021
Valid Pairs
14021
Provider/Model
ollama/gpt-oss:120b
Generation Cost
Metric
Value
Prompt Tokens
14,941,957
Completion Tokens
17,159,263
Total Tokens
32,101,220
GPU Energy
12.9628 kWh
Sources
This dataset was generated from 474 scholarly papers:
#… See the full description on the dataset page: https://huggingface.co/datasets/Kylan12/Synthetic-AI-ML-Dataset.synthetic-beir-dataSynthetic-Dataset-Childrens-Stories**Status: released 13-09-2026, repacked 14-09-2026.** The 14-09-2026 repack replaced 58
items after the acceptance gates were strengthened (prompt-instruction leaks, markdown
bullet lists and blockquotes); the other 29,942 are unchanged. Development stopped, pipeline released 17/09/26.
SAMPLE RELEASE: 30,000 synthetic children's short stories for early-reader language modelling.
Metrics
Value
genres
26
stories per genre
1.153-1.154K stories
total characters
38… See the full description on the dataset page: https://huggingface.co/datasets/ContextReq/Synthetic-Dataset-Childrens-Stories.OpenSeek-Synthetic-Reasoning-Data-Examples
OpenSeek-Reasoning-Data
OpenSeek [Github|Blog]
Recent reseach has demonstrated that the reasoning ability of LLMs originates from the pre-training stage, activated by RL training. Massive raw corpus containing complex human reasoning process, but lack of generalized and effective synthesis method to extract these reasoning process.
News
🔥🔥🔥[2025/02/25] We publish some math, code, and general knowledge domain reasoning data synthesized from the current pipeline.… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/OpenSeek-Synthetic-Reasoning-Data-Examples.korean-full-duplex-synthetic-dataset-preview
Korean Full-Duplex Synthetic Dataset Preview
Overview
Public preview of a Korean full-duplex synthetic speech dataset. This
repository contains 100 conversations sampled from a corpus of 89,273
conversations (2,000.5 hours); it does not publish the full corpus audio.
Preview contents
100 conversation WAV files
data/representative.jsonl
24 kHz, mono, 16-bit PCM
Events: normal, barge_in, backchannel, cutoff_by_user
Annotation format… See the full description on the dataset page: https://huggingface.co/datasets/Wi-Fi/korean-full-duplex-synthetic-dataset-preview.synthetic-superconductor-materials-dataset
synthetic-superconductor-materials-dataset
Synthetic Q&A dataset on Superconductor Materials, generated with SDGS (Synthetic Dataset Generation Suite).
Dataset Details
Metric
Value
Topic
Superconductor Materials
Total Q&A Pairs
2649
Valid Pairs
2649
Provider/Model
ollama/gpt-oss:120b
Sources
This dataset was generated from 170 scholarly papers:
#
Title
Authors
Year
Source
QA Pairs
1
Observation of a large-gap… See the full description on the dataset page: https://huggingface.co/datasets/Kylan12/synthetic-superconductor-materials-dataset.Spatial-Scene-Synthetic-Datasetsyntheticdatasynthetic-neurology-QA-datasetSynthetic-JP-EN-Translation-Dataset-Magpie-Nemotron-4-20k
Synthetic-JP-EN-Translation-Dataset-Magpie-Nemotron-4-20k
Magpieの手法をnvidia/Nemotron-4-340B-Instructに対して適用し作成した、20000件の日⇔英翻訳データセットです。
データセットの作成にはDeepInfraを利用しました。
また、このリポジトリでデータセット作成に用いたコードを公開しています。これをベースに、システムプロンプトとstopを一部変更することで生成しています。
特に事後的なフィルタ処理は加えていないため、クオリティの低いレコードが含まれている可能性があります。ご注意ください。
makisu-dataset-syntheticSynthetic-JP-Coding-Dataset-Magpie-Nemotron-4-10k
Synthetic-JP-Coding-Dataset-Magpie-Nemotron-4-10k
Magpieの手法をnvidia/Nemotron-4-340B-Instructに対して適用し作成した、約10000件の日本語のコーディング用対話データセットです。
データセットの作成にはDeepInfraを利用しました。
また、このリポジトリでデータセット作成に用いたコードを公開しています。これをベースに、システムプロンプトとstopを一部変更することで生成しています。
特に事後的なフィルタ処理は加えていないため、クオリティの低いレコードが含まれている可能性があります。ご注意ください。
Synthetic-JP-EN-Coding-Dataset-Magpie-69k
Synthetic-JP-EN-Coding-Dataset-Magpie-69k
Magpieの手法を様々なモデルに対して適用し作成した、約69000件の日本語・英語のコーディング対話データセットです。
作成に利用したモデルは以下の通りです。modelキーに該当レコードの作成に利用したモデル情報があります。
nvidia/Nemotron-4-340B-Instruct
microsoft/Phi-3-medium-4k-instruct
mistralai/Mixtral-8x22B-Instruct-v0.1
cyberagent/calm3-22b-chat
データセットの作成にはDeepInfraを利用しました。
また、このリポジトリでデータセット作成に用いたコードを公開しています。これをベースに、プロンプトテンプレートやシステムプロンプト等を一部変更することで生成しています。特に事後的なフィルタ処理は加えていないため、クオリティの低いレコードが含まれている可能性があります。ご注意ください。
Bitcoin_synthetic_data
🧠 Bitcoin Synthetic Dataset Collection (AI-generated)
A collection of synthetic Bitcoin transaction datasets enriched with generative AI explanations.
🛠️ Topics
Whale Transactions
OP_RETURN rare patterns
Each transaction includes:
Fee, size, rarity score
Semantic AI-generated description
💡 Use Cases
Training predictive models of Bitcoin activity
Network and anomaly simulation
Financial behavior studies
Temporal analysis and outlier detection
📜License: CC… See the full description on the dataset page: https://huggingface.co/datasets/syn-data/Bitcoin_synthetic_data.building-engineering-synthetic-dataset-v5
Building Engineering Synthetic Dataset (V5)
Repository: Irfanuruchi/building-engineering-synthetic-dataset-v5
This repository contains a synthetic dataset for training engineering reasoning models focused on building engineering calculations and sanity checks.
The dataset was generated using physics-based engineering equations and structured prompts suitable for LLM fine-tuning.
It was used to train:
Irfanuruchi/qwen2.5-1.5b-buildeng-precheck-lora-v5
Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/Irfanuruchi/building-engineering-synthetic-dataset-v5.russian_synthetic_datasetNEUDev_AI_as_code_evaluator_SyntheticDataset
Description
This synthetic dataset was generated using GPT-3.5 Turbo and contains programming challenges in Python, Java, and C#.
Each entry in the dataset includes:
language: The programming language of the solution (Python, Java, or C#)
question: The coding problem or challenge description
solution: A model-generated solution to the problem
label: A quality label indicating if the solution is efficient, inefficient, or buggy
comment: Model-generated feedback explaining the… See the full description on the dataset page: https://huggingface.co/datasets/Hananie/NEUDev_AI_as_code_evaluator_SyntheticDataset.synthetic_stoney_data
Dataset Card for Synthetic Stoney Nakoda Q&A
Dataset Description
This dataset contains 150,000 synthetic question-answer pairs designed for training language models in Stoney Nakoda and English. It was generated as a foundational resource to aid in the development of NLP tools for the low-resource Stoney Nakoda language. The pairs cover translations, grammatical nuances, contextual usage, and cultural relevance derived from bilingual dictionary entries.
Homepage:… See the full description on the dataset page: https://huggingface.co/datasets/HarleyCooper/synthetic_stoney_data.Synthetic-Hinglish-Finetuning-Dataset
Hinglish Conversations Dataset
Overview
This dataset contains synthetically generated conversational dialogues in Hinglish (a blend of Hindi and English). The conversations revolve around typical college life, cultural festivities, daily routines, and general discussions, designed to be relatable and engaging.
Dataset Details
Language: Hinglish (Hindi + English)
Domain: College life, daily interactions, cultural events, and general discussions
Size: 3576… See the full description on the dataset page: https://huggingface.co/datasets/prakharb01/Synthetic-Hinglish-Finetuning-Dataset.aragenre-synthetic-training-data
AraGenre Synthetic Training Data
Synthetic Arabic text corpus generated by NAMAA Community in support of a
submission to the AraGenre 2026 shared task
(hierarchical Arabic genre classification, ArabicNLP 2026 / EMNLP 2026). The
data was produced as a candidate training-augmentation resource for the
task's low-resource setting.
Dataset Summary
2,600 DeepSeek-V3-generated Arabic texts, each labelled with a
broad_genre and specific_genre pair from the AraGenre… See the full description on the dataset page: https://huggingface.co/datasets/HassanB4/aragenre-synthetic-training-data.synthetic-medical-mistakes-dataset
Synthetic Medical Mistakes Dataset (SFT Training Data)
A dataset of 350 synthetic clinical reports with gold-standard error annotations, generated by state-of-the-art LLMs for supervised fine-tuning of clinical error detection models. Created as part of the Clinipal project.
Dataset Description
Overview
This dataset was designed to train AI models to detect critical patient safety errors in clinical documentation. Each entry contains a synthetic emergency… See the full description on the dataset page: https://huggingface.co/datasets/Vrda/synthetic-medical-mistakes-dataset.Synthetic-Weakaura-Datasetsynthetic_qa_data
synthetic_qa_data
This dataset contains synthetic question-answer pairs generated and filtered using the following models:
Generation Models
Qwen/Qwen3-1.7B
Qwen/Qwen3-4B
Qwen/Qwen3-8B
Filtering Model
Qwen/Qwen3.5-35B-A3B — a 35B Mixture-of-Experts model with 3B active parameters
Dataset Structure
data/
├── unfiltered_qa/ # Raw generated QA pairs per model
├── both_filtered_qa/ # QA pairs passing both filters
├──… See the full description on the dataset page: https://huggingface.co/datasets/sohamb37lexsi/synthetic_qa_data.gemma_diverse_synthetic_datamyX-Burmese-Morpho-Synthetic
myX-Burmese-Morpho-Synthetic
myX-Burmese-Morpho-Synthetic is a high-volume, synthetically augmented dataset consisting of over 37.8 million rows of Burmese word formations. Developed by Khant Sint Heinn (Kalix Louis) under the DatarrX organization, this resource is designed to advance the structural understanding of the Burmese language in the field of Natural Language Processing (NLP).
📌 Purpose
The primary goal of this dataset is to improve Burmese NLP by providing… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/myX-Burmese-Morpho-Synthetic.synthetic-data-papers
Synthetic Data Papers — FineSet
A research-paper dataset on Synthetic Data Papers, assembled, deduplicated, and quality-scored by
FineSet from arXiv and Semantic Scholar.
📸 This is a dated snapshot — generated 2026-06-12.
It is not auto-updated. Research on Synthetic Data Papers moves fast — new papers land on arXiv every
week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓
Why this dataset
Quality-scored: quality_score float… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/synthetic-data-papers.Japanese-Knowledge-Base-Synthetic-Data
Japanese Knowledge Base Synthetic Dataset
Conversations: 240,585
File size: ~143 MB
Format: JSON array of chat conversations
Content: synthetic Japanese multi-turn conversations for language tasks
Overview
This dataset contains high-quality synthetic Japanese question-answering pairs, generated from various rich linguistic and encyclopedic sources. It is primarily designed to improve language models' understanding of Japanese grammar, vocabulary, idioms… See the full description on the dataset page: https://huggingface.co/datasets/bunbohue/Japanese-Knowledge-Base-Synthetic-Data.Norwegian-Synthetic-HR-data-v-1
Synthetic norwegian public sector HR dataset
Dataset description
This dataset contains 4,000 rows of synthetic instructional data focused on Human Resources (HR) topics within the Norwegian public sector.
The license for the dataset follows the license of the LLMs used to generate the data. Users are advised to review the specific terms associated with the source models before use.
The datasets includes Chain of Thought (CoT) reasoning traces and is generated using a… See the full description on the dataset page: https://huggingface.co/datasets/Hebbelille/Norwegian-Synthetic-HR-data-v-1.synthetic_data_warmstart_3.25kElyza-qwen_thinking_synthetic_data-v001こちらのデータは、magpie手法を使って、生成した合成データセットです。
使用モデルはElyza社の「elyza/ELYZA-Thinking-1.0-Qwen-32B」です。
※合計84,197件のデータセット
※OpenAI Chatテンプレート形式で作成
magpieコード
!pip install bitsandbytes>=0.45.3
!pip install tokenizers==0.21.0
!pip install --upgrade pyzmq
!pip install vllm==0.8.4
import json
import logging
import torch
import random
from tqdm.auto import tqdm
from vllm import LLM, SamplingParams
# ロギングの設定
logging.basicConfig(level=logging.INFO)
# 設定定数
CONFIG = {
"MODEL_NAME":… See the full description on the dataset page: https://huggingface.co/datasets/kazuyamaa/Elyza-qwen_thinking_synthetic_data-v001.
