CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01PleIAs /SYNTH SYNTH Blog announcement SYNTH is the first open generalist synthetic dataset for training small reasoning model end-to-end, jointly released by Pleias and the AI Alliance. SYNTH includes 79,648,272 individual text samples, comprising over 41 billion words (about 75 billion tokens with Pleias tokenizer). It is based on the amplification of 58,698 articles from Wikipedia and made possible thanks to the Structured Wikipedia dataset from Wikimedia Enterprise. SYNTH differs… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/SYNTH.texttext-generation10M<n<100M277 likes13k downloads5mo agoHugging Face02InternRobotics /SynthVerseViewer is explicitly configured to read the parquet split only. Citation If you find this dataset useful, please cite: @article{zhao2026SythnVerse, title={SynthVerse: A Large-Scale Diverse Synthetic Dataset for Point Tracking}, author={Weiguang Zhao and Haoran Xu and Xingyu Miao and Qin Zhao and Rui Zhang and Kaizhu Huang and Ning Gao and Peizhou Cao and Mingze Sun and Mulin Yu and Tao Lu and Linning Xu and Junting Dong and Jiangmiao Pang}, journal={arXiv preprint… See the full description on the dataset page: https://huggingface.co/datasets/InternRobotics/SynthVerse.imagerobotics1K<n<10K15 likes9.5k downloads4mo agoHugging Face03sophia1ch /zendo-synthetic-data Zendo Synthetic Visual Reasoning Dataset Synthetic Zendo-style scenes with associated rules and per-scene tensor representations. Each scene either follows ("positive", label=1) or violates ("negative", label=0) a rule that is given in natural language and as a Prolog query. Splits split scenes train 56475 test 3344 rules total 3439 Layout images/<split>/<batch>/<rule_id>/<scene_id>.png — rendered scene… See the full description on the dataset page: https://huggingface.co/datasets/sophia1ch/zendo-synthetic-data.imageimage-classification10K<n<100K1 likes7.7k downloads4mo agoHugging Face04annahbanannah /synthetic-math-toolcall-deception Synthetic Math Tool-Call Deception 200 paired multi-turn math-assistant trajectories (400 rows) for evaluating deception detectors on mid-trajectory tool-call misreporting. Each trajectory: a system prompt instructs the model to compute via an execute_python tool under a stated tool-call limit, and requires every call to carry a running call_index argument (1 for the first call, 2 for the second, …). The platform enforcing the limit is said to only see the reported call_index… See the full description on the dataset page: https://huggingface.co/datasets/annahbanannah/synthetic-math-toolcall-deception.tabulartext-classificationn<1K0 likes6.5k downloads2mo agoHugging Face05SynthLabsAI /Big-Math-RL-Verifiedgated Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models Big-Math is the largest open-source dataset of high-quality mathematical problems, curated specifically for reinforcement learning (RL) training in language models. With over 250,000 rigorously filtered and verified problems, Big-Math bridges the gap between quality and quantity, establishing a robust foundation for advancing reasoning in LLMs. Request Early Access to Private… See the full description on the dataset page: https://huggingface.co/datasets/SynthLabsAI/Big-Math-RL-Verified.textquestion-answering100K<n<1M243 likes5.3k downloads2y agoHugging Face06SYNTH-Initiative /SYNTH SYNTH SYNTH is the first open generalist synthetic dataset for training small reasoning model end-to-end, jointly released by Pleias and the AI Alliance. SYNTH includes 79,648,272 individual text samples, comprising over 41 billion words (about 75 billion tokens with Pleias tokenizer). It is based on the amplification of 58,698 articles from Wikipedia and made possible thanks to the Structured Wikipedia dataset from Wikimedia Enterprise. SYNTH differs from existing open synthetic… See the full description on the dataset page: https://huggingface.co/datasets/SYNTH-Initiative/SYNTH.texttext-generation10M<n<100M0 likes4.7k downloads11mo agoHugging Face07naver-clova-ix /synthdog-en Donut 🍩 : OCR-Free Document Understanding Transformer (ECCV 2022) -- SynthDoG datasets For more information, please visit https://github.com/clovaai/donut The links to the SynthDoG-generated datasets are here: synthdog-en: English, 0.5M. synthdog-zh: Chinese, 0.5M. synthdog-ja: Japanese, 0.5M. synthdog-ko: Korean, 0.5M. To generate synthetic datasets with our SynthDoG, please see ./synthdog/README.md and our paper for details. How to Cite If you find this work useful… See the full description on the dataset page: https://huggingface.co/datasets/naver-clova-ix/synthdog-en.image10K<n<100K27 likes4.7k downloads3y agoHugging Face08swesynth /SWE-Synthtext1K<n<10K1 likes4.2k downloads2y agoHugging Face09oolongbench /oolong-synthOolong-synth is a dataset from the paper Oolong: Evaluating Long Context Reasoning and Aggregation Capabilities. See the paper for more details on the dataset construction. To run the standard evaluation setting you will need: input: context_window_text + "\n" + question (these are separated because the context window text can be cached for reuse across multiple input queries) output: answer UPDATE 6/20/2026: Corrected 14 instances, mostly for very-long-context temporal queries. Thanks to… See the full description on the dataset page: https://huggingface.co/datasets/oolongbench/oolong-synth.tabular1K<n<10K5 likes4.1k downloads3mo agoHugging Face10SynthLabsAI /PERSONAgated Dataset Card for PERSONAS (Prism Filter) PERSONAS (Prism filter) is one of the largest datasets of synthetic preferences, with over 200k preferences over thousands of questions and 1k personas. Details on the PERSONAS dataset can be found here paper link. Note that you MUST also fill out the form on our site to receive access to the full dataset. The form is available here. Dataset Details Dataset Description The personas dataset is a pluralistic… See the full description on the dataset page: https://huggingface.co/datasets/SynthLabsAI/PERSONA.text100K<n<1M23 likes4k downloads2y agoHugging Face11yuyijiong /context_qa_sum_qwen3_synthetic Context-based QA and Summarization Synthetic Dataset Overview This dataset contains synthetic context-based question-answering (QA) and summarization data. The data was synthesized using: Source context: openbmb/Ultra-FineWeb Synthesis model: Qwen3-30B-A3B-Instruct-2507 Each context is obtained by taking the initial segment of raw pretraining text from Ultra-FineWeb, truncated to at most the corresponding number of tokens, while ensuring the truncation does not occur in… See the full description on the dataset page: https://huggingface.co/datasets/yuyijiong/context_qa_sum_qwen3_synthetic.texttext-generation10M<n<100M5 likes3.7k downloads6mo agoHugging Face12meganwei /syntheory Dataset Card for SynTheory Dataset Summary SynTheory is a synthetic dataset of music theory concepts, specifically rhythmic (tempos and time signatures) and tonal (notes, intervals, scales, chords, and chord progressions). Each of these 7 concepts has its own config. tempos consist of 161 total integer tempos (bpm) ranging from 50 BPM to 210 BPM (inclusive), 5 percussive instrument types (click_config_name), and 5 random start time offsets (offset_time). time_signatures… See the full description on the dataset page: https://huggingface.co/datasets/meganwei/syntheory.audioaudio-classification100K<n<1M13 likes3.5k downloads2y agoHugging Face13openbmb /VisRAG-Ret-Train-Synthetic-data Dataset Description This dataset is the synthetic part of the training set of VisRAG it includes 239,358 Query-Document (Q-D) Pairs from a synthetic dataset made up of pages from web-crawled PDF documents and augmented with VLM-generated (GPT-4o) pseudo-queries. Our training data is organized with a batch size of 128, ensuring that all data within the same batch comes from the same dataset. Name Source Description # Pages Textbooks https://openstax.org/ College-level… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/VisRAG-Ret-Train-Synthetic-data.image100K<n<1M20 likes3.3k downloads2y agoHugging Face14chrisjcundy /mbpp-synthetic-v3text1K<n<10K0 likes3.3k downloads11mo agoHugging Face15yxchng /laion_synthetic_filtered_large_part3image10M<n<100M0 likes3.3k downloads3y agoHugging Face16SynthLabsAI /PERSONA_subsetgated Dataset Card for PERSONAS (Prism Filter) PERSONAS (Prism filter) is one of the largest datasets of synthetic preferences, with over 200k preferences over thousands of questions and 1k personas. Details on the PERSONAS dataset can be found here paper link Note that this subset is 5% of the training split of PERSONAS. The full dataset is here, strictly available for academic use. You MUST request access to the full persona dataset here. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/SynthLabsAI/PERSONA_subset.text1K<n<10K3 likes3.3k downloads2y agoHugging Face17yxchng /laion_synthetic_filtered_large_part1image10M<n<100M2 likes3.2k downloads3y agoHugging Face18naver-clova-ix /synthdog-ko Donut 🍩 : OCR-Free Document Understanding Transformer (ECCV 2022) -- SynthDoG datasets For more information, please visit https://github.com/clovaai/donut The links to the SynthDoG-generated datasets are here: synthdog-en: English, 0.5M. synthdog-zh: Chinese, 0.5M. synthdog-ja: Japanese, 0.5M. synthdog-ko: Korean, 0.5M. To generate synthetic datasets with our SynthDoG, please see ./synthdog/README.md and our paper for details. How to Cite If you find this work useful… See the full description on the dataset page: https://huggingface.co/datasets/naver-clova-ix/synthdog-ko.image10K<n<100K18 likes3.1k downloads3y agoHugging Face19yxchng /laion_synthetic_filtered_large_part2image10M<n<100M0 likes3k downloads3y agoHugging Face20gretelai /synthetic_text_to_sql Image generated by DALL-E. See prompt for more details synthetic_text_to_sql gretelai/synthetic_text_to_sql is a rich dataset of high quality synthetic Text-to-SQL samples, designed and generated using Gretel Navigator, and released under Apache 2.0. Please see our release blogpost for more details. The dataset includes: 105,851 records partitioned into 100,000 train and 5,851 test records ~23M total tokens, including ~12M SQL tokens Coverage across 100 distinct… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_text_to_sql.textquestion-answering100K<n<1M703 likes2.9k downloads9mo agoHugging Face21docling-project /SynthCodeNet SynthCodeNet SynthCodeNet is a multimodal dataset created for training the SmolDocling model. It consists of over 9.3 million synthetically generated image-text pairs, covering code snippets from 56 different programming languages. Text data was sourced from permissively licensed sources, while images were synthetically generated at 120 DPI using LaTeX and Pygments to ensure visual diversity. Dataset Statistics Total samples: 9,334,257 Training set: 8,400… See the full description on the dataset page: https://huggingface.co/datasets/docling-project/SynthCodeNet.imageimage-text-to-text1M<n<10M15 likes2.8k downloads1y agoHugging Face22argilla /Synth-APIGen-v0.1 Dataset card for Synth-APIGen-v0.1 This dataset has been created with distilabel. Pipeline script: pipeline_apigen_train.py. Dataset creation It has been created with distilabel==1.4.0 version. This dataset is an implementation of APIGen: Automated Pipeline for Generating Verifiable and Diverse Function-Calling Datasets in distilabel, generated from synthetic functions. The process can be summarized as follows: Generate (or in this case modify) python… See the full description on the dataset page: https://huggingface.co/datasets/argilla/Synth-APIGen-v0.1.texttext-generation10K<n<100K65 likes2.8k downloads2y agoHugging Face23Azzindani /ID_Legal_QA_SynThink 🧠 Indonesian Legal QA Synthetic Think Dataset (ID_Legal_QA_SynThink) This repository features an advanced Synthetic Question-and-Answer dataset for the Indonesian legal domain, distinguished by the inclusion of an explicit Thinking Phase (Chain-of-Thought). 🏛️ 💡 The Concept: Transparent Legal Reasoning Standard QA datasets often provide just the "final answer." This dataset goes deeper by capturing the internal reasoning process of the model before it arrives at a… See the full description on the dataset page: https://huggingface.co/datasets/Azzindani/ID_Legal_QA_SynThink.tabulartext-generation1K<n<10K1 likes2.7k downloads7mo agoHugging Face24Zhongzhi1228 /Recursive-Task-Synthesis Recursive Task Synthesis This dataset contains 37,484 validated command-line task instances produced through recursive task synthesis. Public identifiers are opaque and stable. metadata/tasks.parquet: one searchable row per task instance. metadata/shard_manifest.jsonl: TAR sizes and SHA256 checksums. data/tasks-*.tar: sanitized runnable task packages. The searchable task rows include: instruction: contents of instruction.md. task_toml: contents of task.toml. solution:… See the full description on the dataset page: https://huggingface.co/datasets/Zhongzhi1228/Recursive-Task-Synthesis.tabularreinforcement-learning10K<n<100K19 likes2.3k downloads2mo agoHugging Face25gretelai /synthetic_pii_finance_multilingual Image generated by DALL-E. See prompt for more details 💼 📊 Synthetic Financial Domain Documents with PII Labels gretelai/synthetic_pii_finance_multilingual is a dataset of full length synthetic financial documents containing Personally Identifiable Information (PII), generated using Gretel Navigator and released under Apache 2.0. This dataset is designed to assist with the following use cases: 🏷️ Training NER (Named Entity Recognition) models to detect and label PII in… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_pii_finance_multilingual.tabulartext-classification10K<n<100K81 likes2.2k downloads2y agoHugging Face26naver-clova-ix /synthdog-ja Donut 🍩 : OCR-Free Document Understanding Transformer (ECCV 2022) -- SynthDoG datasets For more information, please visit https://github.com/clovaai/donut The links to the SynthDoG-generated datasets are here: synthdog-en: English, 0.5M. synthdog-zh: Chinese, 0.5M. synthdog-ja: Japanese, 0.5M. synthdog-ko: Korean, 0.5M. To generate synthetic datasets with our SynthDoG, please see ./synthdog/README.md and our paper for details. How to Cite If you find this work useful… See the full description on the dataset page: https://huggingface.co/datasets/naver-clova-ix/synthdog-ja.image10K<n<100K5 likes2.2k downloads3y agoHugging Face27allenai /MolmoWeb-SyntheticGround MolmoWeb-SyntheticGround This dataset was introduced in the paper MolmoWeb: Open Visual Web Agent and Open Data for the Open Web. A dataset of webpage screenshots paired with synthetic grounding tasks. Each example asks a model to identify a target element on the page, with ground-truth bounding boxes and (for GPT examples) natural-language thoughts. Dataset Usage from datasets import load_dataset # load the gpt subset ds = load_dataset("allenai/MolmoWeb-SyntheticGround"… See the full description on the dataset page: https://huggingface.co/datasets/allenai/MolmoWeb-SyntheticGround.imageimage-text-to-text100K<n<1M8 likes2.2k downloads6mo agoHugging Face28PrimeIntellect /SYNTHETIC-1 SYNTHETIC-1: Two Million Crowdsourced Reasoning Traces from Deepseek-R1 SYNTHETIC-1 is a reasoning dataset obtained from Deepseek-R1, generated with crowdsourced compute and annotated with diverse verifiers such as LLM judges or symbolic mathematics verifiers. This is the raw version of the dataset, without any filtering for correctness - Filtered datasets specifically for fine-tuning as well as our 7B model can be found in our 🤗 SYNTHETIC-1 Collection. The dataset consists of the… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/SYNTHETIC-1.text1M<n<10M64 likes2.1k downloads2y agoHugging Face29naver-clova-ix /synthdog-zh Donut 🍩 : OCR-Free Document Understanding Transformer (ECCV 2022) -- SynthDoG datasets For more information, please visit https://github.com/clovaai/donut The links to the SynthDoG-generated datasets are here: synthdog-en: English, 0.5M. synthdog-zh: Chinese, 0.5M. synthdog-ja: Japanese, 0.5M. synthdog-ko: Korean, 0.5M. To generate synthetic datasets with our SynthDoG, please see ./synthdog/README.md and our paper for details. How to Cite If you find this work useful… See the full description on the dataset page: https://huggingface.co/datasets/naver-clova-ix/synthdog-zh.image10K<n<100K18 likes2.1k downloads3y agoHugging Face30UWGZQ /Synthetic_Visual_Genome2 Synthetic Visual Genome 2 (SVG2) A large-scale panoptic video scene graph dataset containing object labels, attributes, relationships, and instance-level segmentation masks. Paper: Synthetic Visual Genome 2: Extracting Large-scale Spatio-Temporal Scene Graphs from Videos Website: Synthetic Visual Genome 2 Versions cleaned The cleaned version has two sources: PVD (~593K videos) and SA-V (~43K videos). SAM-3 outputs We also provide the instance masks and… See the full description on the dataset page: https://huggingface.co/datasets/UWGZQ/Synthetic_Visual_Genome2.textvideo-classification100K<n<1M8 likes2k downloads11d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.