CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Lego-X /SWE-Lego-Synthetic-Data Dataset Summary Paper | Github | HF Collection SWE-Lego-Synthetic-Data contains 11.5k synthetic github issues (Python language) and their multi-turn agent trajectories. The column named messages is collected using Qwen/Qwen3-Coder-480B-A35B-Instruct with OpenHands (v0.53.0) agent scaffolding, which can be directly used for SFT training. This dataset is part of the work presented in SWE-Lego, a supervised fine-tuning (SFT) recipe designed to achieve state-of-the-art performance… See the full description on the dataset page: https://huggingface.co/datasets/Lego-X/SWE-Lego-Synthetic-Data.texttext-generation10K<n<100K5 likes405 downloads9mo agoHugging Face02Stereotypes-in-LLMs /hiring-bias-mitigation-synthetic-data Hiring-bias mitigation — synthetic training data Semi-synthetic data for training LLMs to make hiring decisions that do not depend on a protected attribute (military status, gender, religion), in English and Ukrainian. Real inputs, synthetic labels. CVs and job descriptions are real, anonymised postings from the Djinni Recruitment Dataset (MIT). Decisions and rationales were written by the teacher model Qwen/Qwen3.5-122B-A10B-GPTQ-Int4. Code and results:… See the full description on the dataset page: https://huggingface.co/datasets/Stereotypes-in-LLMs/hiring-bias-mitigation-synthetic-data.tabulartext-generation100K<n<1M0 likes277 downloads20h agoHugging Face03ContextReq /Synthetic-Dataset-Childrens-Stories**Status: released 13-09-2026, repacked 14-09-2026.** The 14-09-2026 repack replaced 58 items after the acceptance gates were strengthened (prompt-instruction leaks, markdown bullet lists and blockquotes); the other 29,942 are unchanged. Development stopped, pipeline released 17/09/26. SAMPLE RELEASE: 30,000 synthetic children's short stories for early-reader language modelling. Metrics Value genres 26 stories per genre 1.153-1.154K stories total characters 38… See the full description on the dataset page: https://huggingface.co/datasets/ContextReq/Synthetic-Dataset-Childrens-Stories.texttext-generation10K<n<100K1 likes210 downloads8d agoHugging Face04Aratako /Synthetic-JP-EN-Coding-Dataset-801k Synthetic-JP-EN-Coding-Dataset-801k Magpieによって作成したコードSFTデータセットであるAratako/Synthetic-JP-EN-Coding-Dataset-Magpie-69kを元に、Evol-Instructのような手法を用いて複数のinstructionとresonseを生成し拡張して作成した、日英混合801262件のコードSFT用合成データセットです。 日本語: 173849件 英語: 627413件 元のinstructionの作成に利用したモデルは以下の通りです。modelキーに該当レコードの作成に利用したモデル情報があります。 nvidia/Nemotron-4-340B-Instruct microsoft/Phi-3-medium-4k-instruct mistralai/Mixtral-8x22B-Instruct-v0.1… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Synthetic-JP-EN-Coding-Dataset-801k.tabulartext-generation100K<n<1M17 likes169 downloads2y agoHugging Face05ZennyKenny /synthetic_vc_financial_decisions_reasoning_dataset Best Curator Use Case in the Reasoning Datasets Competition: https://www.linkedin.com/feed/update/urn:li:activity:7330998995990781952/ Synthetic VC Financial Decisions Reasoning Dataset Dataset Summary The Synthetic VC Financial Decisions Reasoning Dataset is a large-scale collection designed to train, evaluate, and fine-tune language models on subjective, abstract financial reasoning tasks. It simulates venture capital (VC) workflows by capturing multiple… See the full description on the dataset page: https://huggingface.co/datasets/ZennyKenny/synthetic_vc_financial_decisions_reasoning_dataset.textreinforcement-learningn<1K15 likes151 downloads1y agoHugging Face06vinhnx90 /synthetic-swift-data-single-turn Dataset Card for synthetic-swift-data-single-turn This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/vinhnx90/synthetic-swift-data-single-turn/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/vinhnx90/synthetic-swift-data-single-turn.texttext-generationn<1K0 likes144 downloads2y agoHugging Face07gratex /GNOTHEIA-synthetic-insurance-dataset GNOTHEIA Synthetic Insurance Dataset Published by: Gratex International a.s.Project: InnovAIte — InnovAIte Slovakia License: Apache 2.0Version: 1.0.0Contact: info@gratex.com A synthetic insurance claims dataset designed for AI systems that evaluate insurance claims using OMG SBVR business rules, structured claim polycontexts and synthetic claim-related documents. The dataset main goal is to support: LLM fine-tuning pipeline SBVR reasoning benchmarks insurance claim AI… See the full description on the dataset page: https://huggingface.co/datasets/gratex/GNOTHEIA-synthetic-insurance-dataset.tabulartext-classification1K<n<10K0 likes136 downloads3d agoHugging Face08Aratako /Synthetic-JP-Preference-Dataset-Qwen2.5_72B-191k Synthetic-JP-Preference-Dataset-Qwen2.5_72B-191k 概要 5種類のオープンモデルとQwen/Qwen2.5-72B-Instruct-GPTQ-Int8を使って作成した、190854件の日本語合成Preferenceデータセットです。 以下、データセットの詳細です。 instructionには、Aratako/Magpie-Tanuki-8B-annotated-96kのinput_qualityがexcellentのものを利用 回答生成には、以下の5つのApache 2.0ライセンスのモデルを利用 weblab-GENIAC/Tanuki-8B-dpo-v1.0 team-hatakeyama-phase2/Tanuki-8x8B-dpo-v1.0-GPTQ-8bit cyberagent/calm3-22b-chat llm-jp/llm-jp-3-13b-instruct Qwen/Qwen2.5-32B-Instruct-GPTQ-Int8… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Synthetic-JP-Preference-Dataset-Qwen2.5_72B-191k.texttext-generation100K<n<1M6 likes109 downloads2y agoHugging Face09robworks-software /database-query-logs-synthetic Database Query Logs (synthetic) 3,995 database query-log entries spanning 10 engines - MySQL, PostgreSQL, MongoDB, SQL Server, Oracle, MariaDB, SQLite, Cassandra, Redis, and Elasticsearch - with query text, type, complexity, execution timing, and row-count metadata. These queries are synthetic The queries were programmatically generated, not captured from production systems. They were produced by templating a set of query shapes across industry-flavored schema… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/database-query-logs-synthetic.tabulartext-classification1K<n<10K0 likes70 downloads2mo agoHugging Face10eagle0504 /synthetic-text2sql-dataset Dataset Card for "synthetic-text2sql-dataset" Dataset Summary The synthetic-text2sql-dataset is a large-scale, structured dataset containing 100,000 training and 5,851 test examples designed to support research and development in SQL semantic parsing, text-to-SQL generation, and chain-of-thought (CoT) reasoning. It was derived from an original DataFrame and converted into Hugging Face's datasets.Dataset format. Three new fields were added: question: alias for the… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/synthetic-text2sql-dataset.textquestion-answering100K<n<1M1 likes57 downloads1y agoHugging Face11Aratako /Synthetic-JP-EN-Translation-Dataset-Magpie-Nemotron-4-20k Synthetic-JP-EN-Translation-Dataset-Magpie-Nemotron-4-20k Magpieの手法をnvidia/Nemotron-4-340B-Instructに対して適用し作成した、20000件の日⇔英翻訳データセットです。 データセットの作成にはDeepInfraを利用しました。 また、このリポジトリでデータセット作成に用いたコードを公開しています。これをベースに、システムプロンプトとstopを一部変更することで生成しています。 特に事後的なフィルタ処理は加えていないため、クオリティの低いレコードが含まれている可能性があります。ご注意ください。 texttext-generation10K<n<100K6 likes54 downloads2y agoHugging Face12Aratako /Synthetic-JP-Coding-Dataset-Magpie-Nemotron-4-10k Synthetic-JP-Coding-Dataset-Magpie-Nemotron-4-10k Magpieの手法をnvidia/Nemotron-4-340B-Instructに対して適用し作成した、約10000件の日本語のコーディング用対話データセットです。 データセットの作成にはDeepInfraを利用しました。 また、このリポジトリでデータセット作成に用いたコードを公開しています。これをベースに、システムプロンプトとstopを一部変更することで生成しています。 特に事後的なフィルタ処理は加えていないため、クオリティの低いレコードが含まれている可能性があります。ご注意ください。 texttext-generation10K<n<100K1 likes48 downloads2y agoHugging Face13Aratako /Synthetic-JP-EN-Coding-Dataset-Magpie-69k Synthetic-JP-EN-Coding-Dataset-Magpie-69k Magpieの手法を様々なモデルに対して適用し作成した、約69000件の日本語・英語のコーディング対話データセットです。 作成に利用したモデルは以下の通りです。modelキーに該当レコードの作成に利用したモデル情報があります。 nvidia/Nemotron-4-340B-Instruct microsoft/Phi-3-medium-4k-instruct mistralai/Mixtral-8x22B-Instruct-v0.1 cyberagent/calm3-22b-chat データセットの作成にはDeepInfraを利用しました。 また、このリポジトリでデータセット作成に用いたコードを公開しています。これをベースに、プロンプトテンプレートやシステムプロンプト等を一部変更することで生成しています。特に事後的なフィルタ処理は加えていないため、クオリティの低いレコードが含まれている可能性があります。ご注意ください。 texttext-generation10K<n<100K9 likes46 downloads2y agoHugging Face14Irfanuruchi /building-engineering-synthetic-dataset-v5 Building Engineering Synthetic Dataset (V5) Repository: Irfanuruchi/building-engineering-synthetic-dataset-v5 This repository contains a synthetic dataset for training engineering reasoning models focused on building engineering calculations and sanity checks. The dataset was generated using physics-based engineering equations and structured prompts suitable for LLM fine-tuning. It was used to train: Irfanuruchi/qwen2.5-1.5b-buildeng-precheck-lora-v5 Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/Irfanuruchi/building-engineering-synthetic-dataset-v5.texttext-generation10K<n<100K1 likes39 downloads7mo agoHugging Face15prakharb01 /Synthetic-Hinglish-Finetuning-Dataset Hinglish Conversations Dataset Overview This dataset contains synthetically generated conversational dialogues in Hinglish (a blend of Hindi and English). The conversations revolve around typical college life, cultural festivities, daily routines, and general discussions, designed to be relatable and engaging. Dataset Details Language: Hinglish (Hindi + English) Domain: College life, daily interactions, cultural events, and general discussions Size: 3576… See the full description on the dataset page: https://huggingface.co/datasets/prakharb01/Synthetic-Hinglish-Finetuning-Dataset.texttext-generation1K<n<10K0 likes35 downloads1y agoHugging Face16jonasluehrs-jaai /synthetic_dataset_low-mid Synthetic Dataset: Low Context, Medium Generation Dataset Description This is a synthetic benchmark dataset designed to test LLM inference performance in low-context, mid-generation scenarios. The dataset consists of 2,000 samples with randomly generated tokens that simulate workloads where models receive short prompts but generate longer responses. Use Cases This dataset is ideal for benchmarking: Creative writing and content generation Code generation from… See the full description on the dataset page: https://huggingface.co/datasets/jonasluehrs-jaai/synthetic_dataset_low-mid.tabulartext-generation1K<n<10K0 likes34 downloads11mo agoHugging Face17Vrda /synthetic-medical-mistakes-dataset Synthetic Medical Mistakes Dataset (SFT Training Data) A dataset of 350 synthetic clinical reports with gold-standard error annotations, generated by state-of-the-art LLMs for supervised fine-tuning of clinical error detection models. Created as part of the Clinipal project. Dataset Description Overview This dataset was designed to train AI models to detect critical patient safety errors in clinical documentation. Each entry contains a synthetic emergency… See the full description on the dataset page: https://huggingface.co/datasets/Vrda/synthetic-medical-mistakes-dataset.texttext-classificationn<1K1 likes33 downloads7mo agoHugging Face18fineset-io /synthetic-data-papers Synthetic Data Papers — FineSet A research-paper dataset on Synthetic Data Papers, assembled, deduplicated, and quality-scored by FineSet from arXiv and Semantic Scholar. 📸 This is a dated snapshot — generated 2026-06-12. It is not auto-updated. Research on Synthetic Data Papers moves fast — new papers land on arXiv every week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓ Why this dataset Quality-scored: quality_score float… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/synthetic-data-papers.tabulartext-classificationn<1K0 likes28 downloads3mo agoHugging Face19jonasluehrs-jaai /synthetic_dataset_mid-mid Synthetic Dataset: Medium Context, Medium Generation Dataset Description This is a synthetic benchmark dataset designed to test LLM inference performance in balanced workload scenarios. The dataset consists of 2,000 samples with randomly generated tokens that simulate realistic, mixed-workload patterns where both input and output lengths are moderate. Use Cases This dataset is ideal for benchmarking: Chat applications with conversational exchanges API services… See the full description on the dataset page: https://huggingface.co/datasets/jonasluehrs-jaai/synthetic_dataset_mid-mid.tabulartext-generation1K<n<10K0 likes27 downloads11mo agoHugging Face20Hebbelille /Norwegian-Synthetic-HR-data-v-1 Synthetic norwegian public sector HR dataset Dataset description This dataset contains 4,000 rows of synthetic instructional data focused on Human Resources (HR) topics within the Norwegian public sector. The license for the dataset follows the license of the LLMs used to generate the data. Users are advised to review the specific terms associated with the source models before use. The datasets includes Chain of Thought (CoT) reasoning traces and is generated using a… See the full description on the dataset page: https://huggingface.co/datasets/Hebbelille/Norwegian-Synthetic-HR-data-v-1.texttext-generation1K<n<10K0 likes24 downloads10mo agoHugging Face21TumeloKonaite /synthetic-patient-dr-data Synthetic Patient DR Data Synthetic doctor-patient consultation dataset with structured clinical outputs and optional full-consultation audio. Dataset Summary This dataset was generated for research and prototyping in: clinical dialogue generation structured clinical extraction text-to-audio workflows conversational healthcare modeling All consultations are synthetic and should not be treated as real clinical encounters. Export Metadata Mode: audio Repo… See the full description on the dataset page: https://huggingface.co/datasets/TumeloKonaite/synthetic-patient-dr-data.audiotext-generationn<1K0 likes23 downloads6mo agoHugging Face22david-ar /synthetic-irc-data Synthetic IRC Conversation Dataset Dataset Description This dataset contains 1,500 synthetic IRC-style conversations featuring multiple participants, including an AI character named Em. The conversations were generated to replicate authentic IRC chat dynamics with natural flow, interruptions, and varied engagement levels. Dataset Summary Total conversations: 1,500 Total size: ~10MB Format: JSONL with IRC-style formatting Language: English License: Apache 2.0… See the full description on the dataset page: https://huggingface.co/datasets/david-ar/synthetic-irc-data.texttext-generation1K<n<10K2 likes21 downloads1y agoHugging Face23himanshunakrani9 /instruction-dataset-qwen-synthetic Synthetic Instruction Dataset A high-quality synthetic instruction-response dataset generated using Qwen 2.5 72B Instruct. Dataset Details Total Examples: 500 Train/Test Split: 475/25 Generator Model: Qwen 2.5 72B Instruct Generated: 2026-05-30 Categories The dataset covers diverse categories: Coding & Programming - Algorithm implementation, code explanation, debugging Mathematics & Reasoning - Problem solving with step-by-step explanations… See the full description on the dataset page: https://huggingface.co/datasets/himanshunakrani9/instruction-dataset-qwen-synthetic.texttext-generationn<1K0 likes20 downloads4mo agoHugging Face24sallani /matchcv-synthetic-dataset MatchCV Synthetic Training Dataset 500 synthetic, anonymous CV/Job matching pairs — generated algorithmically, zero real personal data. Used to fine-tune sallani/MatchCV-Qwen2.5-0.5B. What's inside Each record is an instruction-following example: { "instruction": "Analyse le profil candidat...", "input": "=== CV ANONYMISÉ ===\n...\n=== OFFRE ===\n...", "output": "Score de matching : 73.5/100\nCompétences correspondantes : ...\nGaps techniques :… See the full description on the dataset page: https://huggingface.co/datasets/sallani/matchcv-synthetic-dataset.texttext-generationn<1K0 likes19 downloads4mo agoHugging Face25DatarrX /myX-Burmese-Synthetic-Pseudo-Syllables 📝 Burmese Synthetic Pseudo-Syllables Dataset This dataset contains 5,814,699 computer-generated (synthetic) Myanmar pseudo-syllables structured systematically based on specific complex linguistic and orthographic patterns. Developed as part of the foundational research for low-resource language processing, this dataset serves as a rigorous baseline and stress-testing environment for Burmese Natural Language Processing (NLP), tokenization, font rendering, and spell-checking… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/myX-Burmese-Synthetic-Pseudo-Syllables.texttext-generation1M<n<10M5 likes17 downloads4mo agoHugging Face26m-newhauser /synthetic-rag-dataset Dataset Card for synthetic-rag-dataset This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/m-newhauser/synthetic-rag-dataset/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/m-newhauser/synthetic-rag-dataset.texttext-generationn<1K1 likes15 downloads2y agoHugging Face27kareem2808 /ADHD-Synthetic-Dataset ADHD Assistant Synthetic Chat Dataset (JSONL - ChatML) 📌 Latar Belakang & Urgensi Data Dataset ini dibangun karena absennya dataset open-source yang didedikasikan untuk melatih model bahasa (LLM/SLM) dalam memahami dan membantu kondisi neurodivergen, khususnya ADHD. Menjadi berbeda di tengah lingkungan yang kaku adalah proses yang berat bagi mereka yang belum mampu menerima dirinya sendiri. Dataset ini dirancang dengan penuh kehati-hatian untuk mengisi kekosongan… See the full description on the dataset page: https://huggingface.co/datasets/kareem2808/ADHD-Synthetic-Dataset.texttext-generation1K<n<10K0 likes15 downloads3mo agoHugging Face28DataSynGen /Synthetic_CoT_dataset_RUСинтетический русский датасет для Chain-of-Thought (CoT) представляет собой набор текстов, созданных для тренировки моделей в пошаговом рассуждении. Каждый элемент включает входной запрос, последовательность промежуточных шагов рассуждения и окончательный ответ. Цель датасета – улучшить способность моделей формировать логические объяснения и решения сложных задач. Примеры задач охватывают арифметические вычисления, вопросы по общей эрудиции, логические и аналитические задачи. Данные… See the full description on the dataset page: https://huggingface.co/datasets/DataSynGen/Synthetic_CoT_dataset_RU.texttext-generationn<1K2 likes14 downloads2y agoHugging Face29Abhiram1009 /synthetic-data-factory-5000-20260317 Synthetic Data Factory 5k This dataset contains 5,000 synthetic math and logic examples generated without LLM-based generation. File generated_5000.jsonl: JSONL rows with problem text, explanation text, final answer, tags, metadata, and quality report. Families arithmetic expression evaluation linear equation solving comparison logic / transitive reasoning Generation approach Examples are created through a world-model-first pipeline: latent… See the full description on the dataset page: https://huggingface.co/datasets/Abhiram1009/synthetic-data-factory-5000-20260317.texttext-generation1K<n<10K0 likes13 downloads6mo agoHugging Face30Algocean /gk-synthetic-data-2026-ko gk-synthetic-data-2026-ko.jsonl gk-synthetic-data-2026-ko.jsonl은 외부 teacher 모델의 고급 지식과 장문 설명 능력을 한국어 SFT용으로 증류한 데이터셋이다. 현재 행 수: 77,408 한 줄 용도: 2026년 기준 고급 지식과 장문 설명 능력을 보강하기 위한 한국어 합성 SFT 데이터셋 주 용도: kanana 기반 지식/추론/장문 응답 모델 SFT 형식: JSONL, 한 줄에 하나의 학습 예시 스키마: {"messages":[{"role":"user","content":"..."},{"role":"assistant","content":"..."}]} 일부 행은 assistant 답변의 첫 부분에 공개 판단 근거 형식의 <think>...</think> 블록을 포함한다. 현재 포함 행 수는 111개이며, 닫는 태그 뒤에는 장문 최종 답변이 이어진다. 데이터 의미… See the full description on the dataset page: https://huggingface.co/datasets/Algocean/gk-synthetic-data-2026-ko.texttext-generation10K<n<100K0 likes13 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.