CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Kylan12 /Synthetic-AI-ML-Dataset Synthetic-AI-ML-Dataset Synthetic Q&A dataset on AI and Machine Learning Dataset Details Metric Value Topic AI and Machine Learning Total Q&A Pairs 14021 Valid Pairs 14021 Provider/Model ollama/gpt-oss:120b Generation Cost Metric Value Prompt Tokens 14,941,957 Completion Tokens 17,159,263 Total Tokens 32,101,220 GPU Energy 12.9628 kWh Sources This dataset was generated from 474 scholarly papers: #… See the full description on the dataset page: https://huggingface.co/datasets/Kylan12/Synthetic-AI-ML-Dataset.textquestion-answering10K<n<100K2 likes473 downloads6mo agoHugging Face02din0s /synthetic-beir-datatext1M<n<10M0 likes354 downloads3y agoHugging Face03ContextReq /Synthetic-Dataset-Childrens-Stories**Status: released 13-09-2026, repacked 14-09-2026.** The 14-09-2026 repack replaced 58 items after the acceptance gates were strengthened (prompt-instruction leaks, markdown bullet lists and blockquotes); the other 29,942 are unchanged. Development stopped, pipeline released 17/09/26. SAMPLE RELEASE: 30,000 synthetic children's short stories for early-reader language modelling. Metrics Value genres 26 stories per genre 1.153-1.154K stories total characters 38… See the full description on the dataset page: https://huggingface.co/datasets/ContextReq/Synthetic-Dataset-Childrens-Stories.texttext-generation10K<n<100K1 likes210 downloads7d agoHugging Face04BAAI /OpenSeek-Synthetic-Reasoning-Data-Examples OpenSeek-Reasoning-Data OpenSeek [Github|Blog] Recent reseach has demonstrated that the reasoning ability of LLMs originates from the pre-training stage, activated by RL training. Massive raw corpus containing complex human reasoning process, but lack of generalized and effective synthesis method to extract these reasoning process. News 🔥🔥🔥[2025/02/25] We publish some math, code, and general knowledge domain reasoning data synthesized from the current pipeline.… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/OpenSeek-Synthetic-Reasoning-Data-Examples.text1M<n<10M27 likes208 downloads2y agoHugging Face05Wi-Fi /korean-full-duplex-synthetic-dataset-preview Korean Full-Duplex Synthetic Dataset Preview Overview Public preview of a Korean full-duplex synthetic speech dataset. This repository contains 100 conversations sampled from a corpus of 89,273 conversations (2,000.5 hours); it does not publish the full corpus audio. Preview contents 100 conversation WAV files data/representative.jsonl 24 kHz, mono, 16-bit PCM Events: normal, barge_in, backchannel, cutoff_by_user Annotation format… See the full description on the dataset page: https://huggingface.co/datasets/Wi-Fi/korean-full-duplex-synthetic-dataset-preview.audioautomatic-speech-recognitionn<1K1 likes138 downloads1mo agoHugging Face06Kylan12 /synthetic-superconductor-materials-dataset synthetic-superconductor-materials-dataset Synthetic Q&A dataset on Superconductor Materials, generated with SDGS (Synthetic Dataset Generation Suite). Dataset Details Metric Value Topic Superconductor Materials Total Q&A Pairs 2649 Valid Pairs 2649 Provider/Model ollama/gpt-oss:120b Sources This dataset was generated from 170 scholarly papers: # Title Authors Year Source QA Pairs 1 Observation of a large-gap… See the full description on the dataset page: https://huggingface.co/datasets/Kylan12/synthetic-superconductor-materials-dataset.textquestion-answering1K<n<10K1 likes113 downloads7mo agoHugging Face07DeepAIResearch /Spatial-Scene-Synthetic-Datasettext10K<n<100K0 likes104 downloads2y agoHugging Face08Skorcht /syntheticdatatextn<1K0 likes76 downloads2y agoHugging Face09KryptoniteCrown /synthetic-neurology-QA-datasettext1K<n<10K3 likes62 downloads8mo agoHugging Face10Aratako /Synthetic-JP-EN-Translation-Dataset-Magpie-Nemotron-4-20k Synthetic-JP-EN-Translation-Dataset-Magpie-Nemotron-4-20k Magpieの手法をnvidia/Nemotron-4-340B-Instructに対して適用し作成した、20000件の日⇔英翻訳データセットです。 データセットの作成にはDeepInfraを利用しました。 また、このリポジトリでデータセット作成に用いたコードを公開しています。これをベースに、システムプロンプトとstopを一部変更することで生成しています。 特に事後的なフィルタ処理は加えていないため、クオリティの低いレコードが含まれている可能性があります。ご注意ください。 texttext-generation10K<n<100K6 likes55 downloads2y agoHugging Face11RinKana /makisu-dataset-synthetictext1K<n<10K1 likes54 downloads9mo agoHugging Face12Aratako /Synthetic-JP-Coding-Dataset-Magpie-Nemotron-4-10k Synthetic-JP-Coding-Dataset-Magpie-Nemotron-4-10k Magpieの手法をnvidia/Nemotron-4-340B-Instructに対して適用し作成した、約10000件の日本語のコーディング用対話データセットです。 データセットの作成にはDeepInfraを利用しました。 また、このリポジトリでデータセット作成に用いたコードを公開しています。これをベースに、システムプロンプトとstopを一部変更することで生成しています。 特に事後的なフィルタ処理は加えていないため、クオリティの低いレコードが含まれている可能性があります。ご注意ください。 texttext-generation10K<n<100K1 likes47 downloads2y agoHugging Face13Aratako /Synthetic-JP-EN-Coding-Dataset-Magpie-69k Synthetic-JP-EN-Coding-Dataset-Magpie-69k Magpieの手法を様々なモデルに対して適用し作成した、約69000件の日本語・英語のコーディング対話データセットです。 作成に利用したモデルは以下の通りです。modelキーに該当レコードの作成に利用したモデル情報があります。 nvidia/Nemotron-4-340B-Instruct microsoft/Phi-3-medium-4k-instruct mistralai/Mixtral-8x22B-Instruct-v0.1 cyberagent/calm3-22b-chat データセットの作成にはDeepInfraを利用しました。 また、このリポジトリでデータセット作成に用いたコードを公開しています。これをベースに、プロンプトテンプレートやシステムプロンプト等を一部変更することで生成しています。特に事後的なフィルタ処理は加えていないため、クオリティの低いレコードが含まれている可能性があります。ご注意ください。 texttext-generation10K<n<100K9 likes44 downloads2y agoHugging Face14syn-data /Bitcoin_synthetic_data 🧠 Bitcoin Synthetic Dataset Collection (AI-generated) A collection of synthetic Bitcoin transaction datasets enriched with generative AI explanations. 🛠️ Topics Whale Transactions OP_RETURN rare patterns Each transaction includes: Fee, size, rarity score Semantic AI-generated description 💡 Use Cases Training predictive models of Bitcoin activity Network and anomaly simulation Financial behavior studies Temporal analysis and outlier detection 📜License: CC… See the full description on the dataset page: https://huggingface.co/datasets/syn-data/Bitcoin_synthetic_data.tabular1K<n<10K1 likes44 downloads11mo agoHugging Face15Irfanuruchi /building-engineering-synthetic-dataset-v5 Building Engineering Synthetic Dataset (V5) Repository: Irfanuruchi/building-engineering-synthetic-dataset-v5 This repository contains a synthetic dataset for training engineering reasoning models focused on building engineering calculations and sanity checks. The dataset was generated using physics-based engineering equations and structured prompts suitable for LLM fine-tuning. It was used to train: Irfanuruchi/qwen2.5-1.5b-buildeng-precheck-lora-v5 Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/Irfanuruchi/building-engineering-synthetic-dataset-v5.texttext-generation10K<n<100K1 likes39 downloads7mo agoHugging Face16SashaTusur /russian_synthetic_datasettext10K<n<100K1 likes39 downloads3mo agoHugging Face17Hananie /NEUDev_AI_as_code_evaluator_SyntheticDataset Description This synthetic dataset was generated using GPT-3.5 Turbo and contains programming challenges in Python, Java, and C#. Each entry in the dataset includes: language: The programming language of the solution (Python, Java, or C#) question: The coding problem or challenge description solution: A model-generated solution to the problem label: A quality label indicating if the solution is efficient, inefficient, or buggy comment: Model-generated feedback explaining the… See the full description on the dataset page: https://huggingface.co/datasets/Hananie/NEUDev_AI_as_code_evaluator_SyntheticDataset.text10K<n<100K0 likes36 downloads1y agoHugging Face18HarleyCooper /synthetic_stoney_data Dataset Card for Synthetic Stoney Nakoda Q&A Dataset Description This dataset contains 150,000 synthetic question-answer pairs designed for training language models in Stoney Nakoda and English. It was generated as a foundational resource to aid in the development of NLP tools for the low-resource Stoney Nakoda language. The pairs cover translations, grammatical nuances, contextual usage, and cultural relevance derived from bilingual dictionary entries. Homepage:… See the full description on the dataset page: https://huggingface.co/datasets/HarleyCooper/synthetic_stoney_data.text10K<n<100K0 likes36 downloads1y agoHugging Face19prakharb01 /Synthetic-Hinglish-Finetuning-Dataset Hinglish Conversations Dataset Overview This dataset contains synthetically generated conversational dialogues in Hinglish (a blend of Hindi and English). The conversations revolve around typical college life, cultural festivities, daily routines, and general discussions, designed to be relatable and engaging. Dataset Details Language: Hinglish (Hindi + English) Domain: College life, daily interactions, cultural events, and general discussions Size: 3576… See the full description on the dataset page: https://huggingface.co/datasets/prakharb01/Synthetic-Hinglish-Finetuning-Dataset.texttext-generation1K<n<10K0 likes34 downloads1y agoHugging Face20HassanB4 /aragenre-synthetic-training-data AraGenre Synthetic Training Data Synthetic Arabic text corpus generated by NAMAA Community in support of a submission to the AraGenre 2026 shared task (hierarchical Arabic genre classification, ArabicNLP 2026 / EMNLP 2026). The data was produced as a candidate training-augmentation resource for the task's low-resource setting. Dataset Summary 2,600 DeepSeek-V3-generated Arabic texts, each labelled with a broad_genre and specific_genre pair from the AraGenre… See the full description on the dataset page: https://huggingface.co/datasets/HassanB4/aragenre-synthetic-training-data.texttext-classification1K<n<10K0 likes34 downloads1mo agoHugging Face21Vrda /synthetic-medical-mistakes-dataset Synthetic Medical Mistakes Dataset (SFT Training Data) A dataset of 350 synthetic clinical reports with gold-standard error annotations, generated by state-of-the-art LLMs for supervised fine-tuning of clinical error detection models. Created as part of the Clinipal project. Dataset Description Overview This dataset was designed to train AI models to detect critical patient safety errors in clinical documentation. Each entry contains a synthetic emergency… See the full description on the dataset page: https://huggingface.co/datasets/Vrda/synthetic-medical-mistakes-dataset.texttext-classificationn<1K1 likes33 downloads7mo agoHugging Face22Crystalcareai /Synthetic-Weakaura-Datasettext1K<n<10K0 likes32 downloads3y agoHugging Face23sohamb37lexsi /synthetic_qa_data synthetic_qa_data This dataset contains synthetic question-answer pairs generated and filtered using the following models: Generation Models Qwen/Qwen3-1.7B Qwen/Qwen3-4B Qwen/Qwen3-8B Filtering Model Qwen/Qwen3.5-35B-A3B — a 35B Mixture-of-Experts model with 3B active parameters Dataset Structure data/ ├── unfiltered_qa/ # Raw generated QA pairs per model ├── both_filtered_qa/ # QA pairs passing both filters ├──… See the full description on the dataset page: https://huggingface.co/datasets/sohamb37lexsi/synthetic_qa_data.text1K<n<10K0 likes31 downloads4mo agoHugging Face24kazuyamaa /gemma_diverse_synthetic_datatext10K<n<100K0 likes27 downloads1y agoHugging Face25DatarrX /myX-Burmese-Morpho-Synthetic myX-Burmese-Morpho-Synthetic myX-Burmese-Morpho-Synthetic is a high-volume, synthetically augmented dataset consisting of over 37.8 million rows of Burmese word formations. Developed by Khant Sint Heinn (Kalix Louis) under the DatarrX organization, this resource is designed to advance the structural understanding of the Burmese language in the field of Natural Language Processing (NLP). 📌 Purpose The primary goal of this dataset is to improve Burmese NLP by providing… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/myX-Burmese-Morpho-Synthetic.texttext-classification10M<n<100M4 likes26 downloads5mo agoHugging Face26fineset-io /synthetic-data-papers Synthetic Data Papers — FineSet A research-paper dataset on Synthetic Data Papers, assembled, deduplicated, and quality-scored by FineSet from arXiv and Semantic Scholar. 📸 This is a dated snapshot — generated 2026-06-12. It is not auto-updated. Research on Synthetic Data Papers moves fast — new papers land on arXiv every week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓ Why this dataset Quality-scored: quality_score float… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/synthetic-data-papers.tabulartext-classificationn<1K0 likes26 downloads3mo agoHugging Face27bunbohue /Japanese-Knowledge-Base-Synthetic-Data Japanese Knowledge Base Synthetic Dataset Conversations: 240,585 File size: ~143 MB Format: JSON array of chat conversations Content: synthetic Japanese multi-turn conversations for language tasks Overview This dataset contains high-quality synthetic Japanese question-answering pairs, generated from various rich linguistic and encyclopedic sources. It is primarily designed to improve language models' understanding of Japanese grammar, vocabulary, idioms… See the full description on the dataset page: https://huggingface.co/datasets/bunbohue/Japanese-Knowledge-Base-Synthetic-Data.text100K<n<1M0 likes26 downloads2mo agoHugging Face28Hebbelille /Norwegian-Synthetic-HR-data-v-1 Synthetic norwegian public sector HR dataset Dataset description This dataset contains 4,000 rows of synthetic instructional data focused on Human Resources (HR) topics within the Norwegian public sector. The license for the dataset follows the license of the LLMs used to generate the data. Users are advised to review the specific terms associated with the source models before use. The datasets includes Chain of Thought (CoT) reasoning traces and is generated using a… See the full description on the dataset page: https://huggingface.co/datasets/Hebbelille/Norwegian-Synthetic-HR-data-v-1.texttext-generation1K<n<10K0 likes25 downloads10mo agoHugging Face29fchesnay /synthetic_data_warmstart_3.25ktext1K<n<10K0 likes24 downloads1y agoHugging Face30kazuyamaa /Elyza-qwen_thinking_synthetic_data-v001こちらのデータは、magpie手法を使って、生成した合成データセットです。 使用モデルはElyza社の「elyza/ELYZA-Thinking-1.0-Qwen-32B」です。 ※合計84,197件のデータセット ※OpenAI Chatテンプレート形式で作成 magpieコード !pip install bitsandbytes>=0.45.3 !pip install tokenizers==0.21.0 !pip install --upgrade pyzmq !pip install vllm==0.8.4 import json import logging import torch import random from tqdm.auto import tqdm from vllm import LLM, SamplingParams # ロギングの設定 logging.basicConfig(level=logging.INFO) # 設定定数 CONFIG = { "MODEL_NAME":… See the full description on the dataset page: https://huggingface.co/datasets/kazuyamaa/Elyza-qwen_thinking_synthetic_data-v001.text10K<n<100K0 likes24 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.