CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sophia1ch /zendo-synthetic-data Zendo Synthetic Visual Reasoning Dataset Synthetic Zendo-style scenes with associated rules and per-scene tensor representations. Each scene either follows ("positive", label=1) or violates ("negative", label=0) a rule that is given in natural language and as a Prolog query. Splits split scenes train 56475 test 3344 rules total 3439 Layout images/<split>/<batch>/<rule_id>/<scene_id>.png — rendered scene… See the full description on the dataset page: https://huggingface.co/datasets/sophia1ch/zendo-synthetic-data.imageimage-classification10K<n<100K1 likes7.7k downloads4mo agoHugging Face02annahbanannah /synthetic-math-toolcall-deception Synthetic Math Tool-Call Deception 200 paired multi-turn math-assistant trajectories (400 rows) for evaluating deception detectors on mid-trajectory tool-call misreporting. Each trajectory: a system prompt instructs the model to compute via an execute_python tool under a stated tool-call limit, and requires every call to carry a running call_index argument (1 for the first call, 2 for the second, …). The platform enforcing the limit is said to only see the reported call_index… See the full description on the dataset page: https://huggingface.co/datasets/annahbanannah/synthetic-math-toolcall-deception.tabulartext-classificationn<1K0 likes6.5k downloads2mo agoHugging Face03oolongbench /oolong-synthOolong-synth is a dataset from the paper Oolong: Evaluating Long Context Reasoning and Aggregation Capabilities. See the paper for more details on the dataset construction. To run the standard evaluation setting you will need: input: context_window_text + "\n" + question (these are separated because the context window text can be cached for reuse across multiple input queries) output: answer UPDATE 6/20/2026: Corrected 14 instances, mostly for very-long-context temporal queries. Thanks to… See the full description on the dataset page: https://huggingface.co/datasets/oolongbench/oolong-synth.tabular1K<n<10K5 likes4.1k downloads3mo agoHugging Face04yxchng /laion_synthetic_filtered_large_part3image10M<n<100M0 likes3.3k downloads3y agoHugging Face05yxchng /laion_synthetic_filtered_large_part1image10M<n<100M2 likes3.2k downloads3y agoHugging Face06yxchng /laion_synthetic_filtered_large_part2image10M<n<100M0 likes3k downloads3y agoHugging Face07Azzindani /ID_Legal_QA_SynThink 🧠 Indonesian Legal QA Synthetic Think Dataset (ID_Legal_QA_SynThink) This repository features an advanced Synthetic Question-and-Answer dataset for the Indonesian legal domain, distinguished by the inclusion of an explicit Thinking Phase (Chain-of-Thought). 🏛️ 💡 The Concept: Transparent Legal Reasoning Standard QA datasets often provide just the "final answer." This dataset goes deeper by capturing the internal reasoning process of the model before it arrives at a… See the full description on the dataset page: https://huggingface.co/datasets/Azzindani/ID_Legal_QA_SynThink.tabulartext-generation1K<n<10K1 likes2.7k downloads7mo agoHugging Face08Zhongzhi1228 /Recursive-Task-Synthesis Recursive Task Synthesis This dataset contains 37,484 validated command-line task instances produced through recursive task synthesis. Public identifiers are opaque and stable. metadata/tasks.parquet: one searchable row per task instance. metadata/shard_manifest.jsonl: TAR sizes and SHA256 checksums. data/tasks-*.tar: sanitized runnable task packages. The searchable task rows include: instruction: contents of instruction.md. task_toml: contents of task.toml. solution:… See the full description on the dataset page: https://huggingface.co/datasets/Zhongzhi1228/Recursive-Task-Synthesis.tabularreinforcement-learning10K<n<100K19 likes2.3k downloads2mo agoHugging Face09gretelai /synthetic_pii_finance_multilingual Image generated by DALL-E. See prompt for more details 💼 📊 Synthetic Financial Domain Documents with PII Labels gretelai/synthetic_pii_finance_multilingual is a dataset of full length synthetic financial documents containing Personally Identifiable Information (PII), generated using Gretel Navigator and released under Apache 2.0. This dataset is designed to assist with the following use cases: 🏷️ Training NER (Named Entity Recognition) models to detect and label PII in… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_pii_finance_multilingual.tabulartext-classification10K<n<100K81 likes2.2k downloads2y agoHugging Face10physicl /synthetic-bathroom-dataset-for-robotic-perception Synthetic Bathroom Dataset for Robotic Perception Generated by datapack-import.ts This dataset mirrors public data-pack render outputs from Physicl. Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is enabled by the… See the full description on the dataset page: https://huggingface.co/datasets/physicl/synthetic-bathroom-dataset-for-robotic-perception.imagen<1K0 likes1.9k downloads3mo agoHugging Face11vidore /syntheticDocQA_artificial_intelligence_test_beirBEIR version of vidore/syntheticDocQA_artificial_intelligence_test. imagedocument-question-answering1K<n<10K0 likes1.9k downloads1y agoHugging Face12vidore /syntheticDocQA_energy_test_beirBEIR version of vidore/syntheticDocQA_energy_test. imagedocument-question-answering1K<n<10K0 likes1.8k downloads1y agoHugging Face13vidore /syntheticDocQA_healthcare_industry_test_beirBEIR version of vidore/syntheticDocQA_healthcare_industry_test. imagedocument-question-answering1K<n<10K0 likes1.8k downloads1y agoHugging Face14vidore /syntheticDocQA_government_reports_test_beirBEIR version of vidore/syntheticDocQA_government_reports_test. imagedocument-question-answering1K<n<10K1 likes1.8k downloads1y agoHugging Face15physicl /synthetic-living-room-dataset-for-robotic-perception Synthetic Living Room Dataset for Robotic Perception Generated by datapack-import.ts This dataset mirrors public data-pack render outputs from Physicl. Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is enabled by… See the full description on the dataset page: https://huggingface.co/datasets/physicl/synthetic-living-room-dataset-for-robotic-perception.imagen<1K0 likes1.5k downloads3mo agoHugging Face16yxchng /laion_synthetic_filtered_large_part4image10M<n<100M0 likes1.5k downloads3y agoHugging Face17GOD111111111 /synthetic-timeseries-data cruscy data — evaluation sample Three full days of real crypto market microstructure (Binance spot), prepared for public evaluation: absolute prices, dates, and the instrument are withheld — the shape of the day (tick-by-tick relative price, normalized volumes, trade side, book imbalance) is fully preserved. The full feed — 27+ streams (raw L2 depth, 1-second trade tape, order-book metrics, derived features, regime labels) with SQL console, backtest runner and MCP access for AI… See the full description on the dataset page: https://huggingface.co/datasets/GOD111111111/synthetic-timeseries-data.tabular1B<n<10B0 likes1.3k downloads4h agoHugging Face18ServiceNow-AI /SynthDocBench Built With Llama! SynthDocBench SynthDocBench is a fully synthetic benchmark for evaluating vision-language models (VLMs) on complex, multi-page PDF documents. Documents are generated end-to-end by an LLM pipeline that produces realistic multi-page reports with embedded D3.js charts, rich layouts, and deterministically grounded ground-truth answers — enabling controlled, noise-free evaluation impossible with real-world corpora. Paper: SynthDocBench: A Controlled… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow-AI/SynthDocBench.documentvisual-question-answering1K<n<10K4 likes1.1k downloads4mo agoHugging Face19Zhongzhi1228 /Recursive-Task-Synthesis-Trajectories Recursive Task Synthesis Trajectories This dataset contains 327,189 completed agent trajectories collected on recursively synthesized command-line tasks. Public identifiers are opaque and stable. The trajectory JSON retains messages, actions, observations, and token counts. Token-level log-probability arrays and duplicated debug/session captures are excluded from the public packages. metadata/trajectories.parquet: searchable trajectory metadata. metadata/shard_manifest.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/Zhongzhi1228/Recursive-Task-Synthesis-Trajectories.tabularreinforcement-learning100K<n<1M3 likes1.1k downloads1mo agoHugging Face20mteb /syntheticDocQA_healthcare_industry_test_beirBEIR version of vidore/syntheticDocQA_healthcare_industry_test. imagedocument-question-answering1K<n<10K0 likes1k downloads8mo agoHugging Face21mteb /syntheticDocQA_artificial_intelligence_test_beirBEIR version of vidore/syntheticDocQA_artificial_intelligence_test. imagedocument-question-answering1K<n<10K0 likes1k downloads8mo agoHugging Face22mteb /syntheticDocQA_government_reports_test_beirBEIR version of vidore/syntheticDocQA_government_reports_test. imagedocument-question-answering1K<n<10K0 likes981 downloads8mo agoHugging Face23mteb /syntheticDocQA_energy_test_beirBEIR version of vidore/syntheticDocQA_energy_test. imagedocument-question-answering1K<n<10K0 likes974 downloads8mo agoHugging Face24electricsheepafrica /africa-synth-snp-array-aims-all Genome-wide SNP Array with Ancestry Informative Markers (SSA-focused) | Africa (Electric Sheep Africa metadata inventory) Size category: 1M<n<10M - Formats: parquet - Sector: culture_language - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-snp-array-aims-all.tabulartabular-classification1M<n<10M0 likes738 downloads2mo agoHugging Face25Synthyra /AtlasFold-Data AtlasFold-Data AtlasFold-Data is a Parquet conversion of the structures that AtlasFold released for training its monomer and complex models (preprint). The original release is nine .tar.zst archives of LMDB databases in a Google Drive folder. This repository holds the same entries as typed Parquet that datasets, pyarrow, DuckDB, and Polars can stream, plus index tables that reproduce AtlasFold's training sampler exactly. No value was changed. Each Parquet row was read back and… See the full description on the dataset page: https://huggingface.co/datasets/Synthyra/AtlasFold-Data.tabular10M<n<100M0 likes730 downloads9d agoHugging Face26richardyoung /synthea-575k-patients Synthea Synthetic Patient Records (575K Patients) A comprehensive synthetic healthcare dataset containing 575,415 patients with complete medical histories, generated using Synthea — the gold standard for synthetic EHR data. No real patient data. Fully synthetic, HIPAA-safe, and ready for ML research and education. Why This Dataset? 575K patients with realistic demographics, conditions, medications, and encounters Privacy-safe: No real PHI — use freely in research… See the full description on the dataset page: https://huggingface.co/datasets/richardyoung/synthea-575k-patients.tabulartext-generation1B<n<10B12 likes678 downloads6mo agoHugging Face27hotchpotch /wikipedia-multilingual-synthetic-ir-query wikipedia-multilingual-synthetic-ir-query This dataset contains multilingual Wikipedia-derived synthetic query-document pairs for information retrieval training. It was created with the query-crafter-multilingual model, which generates search-like queries from Wikipedia text. The current release contains two different retrieval settings: short_doc: pairs of (query, short document) long_doc: pairs of (query, long document) These two subsets are not generated in the same way… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/wikipedia-multilingual-synthetic-ir-query.tabulartext-retrieval10M<n<100M0 likes585 downloads4mo agoHugging Face28jinaai /airbnb-synthetic-retrieval_beirThis is a copy of https://huggingface.co/datasets/jinaai/airbnb-synthetic-retrieval reformatted into the BEIR format. For any further information like license, please refer to the original dataset. Disclaimer This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at)… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/airbnb-synthetic-retrieval_beir.image1K<n<10K0 likes524 downloads1y agoHugging Face29open-athena /recursive-task-synthesis-glm-5.3-rollouts GLM 5.3 agentic rollouts on Recursive-Task-Synthesis This dataset catalogs the full collection made from the pinned Recursive-Task-Synthesis dataset revision be44f96808d5a9b599d5cb024341ff00091adeb7. The repository includes approximately 260.5 GiB of trajectory payload tar shards. Contents at a glance Item Count Source tasks considered 37,284 Source candidates inspected 19,368 Converted tasks after source filters 18,600 Tasks passing gold… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/recursive-task-synthesis-glm-5.3-rollouts.tabulartext-generation100K<n<1M0 likes509 downloads7d agoHugging Face30lum-ai /metal-python-synthetic-explanations-gpt4-graphcodeberttabular1M<n<10M0 likes487 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.