datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
zendo-synthetic-data
Zendo Synthetic Visual Reasoning Dataset
Synthetic Zendo-style scenes with associated rules and per-scene tensor
representations. Each scene either follows ("positive", label=1) or violates
("negative", label=0) a rule that is given in natural language and as a Prolog
query.
Splits
split
scenes
train
56475
test
3344
rules total
3439
Layout
images/<split>/<batch>/<rule_id>/<scene_id>.png — rendered scene… See the full description on the dataset page: https://huggingface.co/datasets/sophia1ch/zendo-synthetic-data.synthetic-math-toolcall-deception
Synthetic Math Tool-Call Deception
200 paired multi-turn math-assistant trajectories (400 rows) for evaluating deception
detectors on mid-trajectory tool-call misreporting.
Each trajectory: a system prompt instructs the model to compute via an execute_python
tool under a stated tool-call limit, and requires every call to carry a running
call_index argument (1 for the first call, 2 for the second, …). The platform enforcing
the limit is said to only see the reported call_index… See the full description on the dataset page: https://huggingface.co/datasets/annahbanannah/synthetic-math-toolcall-deception.oolong-synthOolong-synth is a dataset from the paper Oolong: Evaluating Long Context Reasoning and Aggregation Capabilities. See the paper for more details on the dataset construction.
To run the standard evaluation setting you will need:
input: context_window_text + "\n" + question (these are separated because the context window text can be cached for reuse across multiple input queries)
output: answer
UPDATE 6/20/2026: Corrected 14 instances, mostly for very-long-context temporal queries. Thanks to… See the full description on the dataset page: https://huggingface.co/datasets/oolongbench/oolong-synth.laion_synthetic_filtered_large_part3laion_synthetic_filtered_large_part1laion_synthetic_filtered_large_part2ID_Legal_QA_SynThink
🧠 Indonesian Legal QA Synthetic Think Dataset (ID_Legal_QA_SynThink)
This repository features an advanced Synthetic Question-and-Answer dataset for the Indonesian legal domain, distinguished by the inclusion of an explicit Thinking Phase (Chain-of-Thought). 🏛️
💡 The Concept: Transparent Legal Reasoning
Standard QA datasets often provide just the "final answer." This dataset goes deeper by capturing the internal reasoning process of the model before it arrives at a… See the full description on the dataset page: https://huggingface.co/datasets/Azzindani/ID_Legal_QA_SynThink.Recursive-Task-Synthesis
Recursive Task Synthesis
This dataset contains 37,484 validated command-line task instances produced
through recursive task synthesis. Public identifiers are opaque and stable.
metadata/tasks.parquet: one searchable row per task instance.
metadata/shard_manifest.jsonl: TAR sizes and SHA256 checksums.
data/tasks-*.tar: sanitized runnable task packages.
The searchable task rows include:
instruction: contents of instruction.md.
task_toml: contents of task.toml.
solution:… See the full description on the dataset page: https://huggingface.co/datasets/Zhongzhi1228/Recursive-Task-Synthesis.synthetic_pii_finance_multilingual
Image generated by DALL-E. See prompt for more details
💼 📊 Synthetic Financial Domain Documents with PII Labels
gretelai/synthetic_pii_finance_multilingual is a dataset of full length synthetic financial documents containing Personally Identifiable Information (PII), generated using Gretel Navigator and released under Apache 2.0.
This dataset is designed to assist with the following use cases:
🏷️ Training NER (Named Entity Recognition) models to detect and label PII in… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_pii_finance_multilingual.synthetic-bathroom-dataset-for-robotic-perception
Synthetic Bathroom Dataset for Robotic Perception
Generated by datapack-import.ts
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is enabled by the… See the full description on the dataset page: https://huggingface.co/datasets/physicl/synthetic-bathroom-dataset-for-robotic-perception.syntheticDocQA_artificial_intelligence_test_beirBEIR version of vidore/syntheticDocQA_artificial_intelligence_test.
syntheticDocQA_energy_test_beirBEIR version of vidore/syntheticDocQA_energy_test.
syntheticDocQA_healthcare_industry_test_beirBEIR version of vidore/syntheticDocQA_healthcare_industry_test.
syntheticDocQA_government_reports_test_beirBEIR version of vidore/syntheticDocQA_government_reports_test.
synthetic-living-room-dataset-for-robotic-perception
Synthetic Living Room Dataset for Robotic Perception
Generated by datapack-import.ts
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is enabled by… See the full description on the dataset page: https://huggingface.co/datasets/physicl/synthetic-living-room-dataset-for-robotic-perception.laion_synthetic_filtered_large_part4synthetic-timeseries-data
cruscy data — evaluation sample
Three full days of real crypto market microstructure (Binance spot), prepared
for public evaluation: absolute prices, dates, and the instrument are withheld —
the shape of the day (tick-by-tick relative price, normalized volumes, trade
side, book imbalance) is fully preserved.
The full feed — 27+ streams (raw L2 depth, 1-second trade tape, order-book
metrics, derived features, regime labels) with SQL console, backtest runner and
MCP access for AI… See the full description on the dataset page: https://huggingface.co/datasets/GOD111111111/synthetic-timeseries-data.SynthDocBench
Built With Llama!
SynthDocBench
SynthDocBench is a fully synthetic benchmark for evaluating vision-language models (VLMs)
on complex, multi-page PDF documents.
Documents are generated end-to-end by an LLM pipeline that produces realistic multi-page reports
with embedded D3.js charts, rich layouts, and deterministically grounded ground-truth answers —
enabling controlled, noise-free evaluation impossible with real-world corpora.
Paper: SynthDocBench: A Controlled… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow-AI/SynthDocBench.Recursive-Task-Synthesis-Trajectories
Recursive Task Synthesis Trajectories
This dataset contains 327,189 completed agent trajectories collected on
recursively synthesized command-line tasks. Public identifiers are opaque and
stable.
The trajectory JSON retains messages, actions, observations, and token counts.
Token-level log-probability arrays and duplicated debug/session captures are
excluded from the public packages.
metadata/trajectories.parquet: searchable trajectory metadata.
metadata/shard_manifest.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/Zhongzhi1228/Recursive-Task-Synthesis-Trajectories.syntheticDocQA_healthcare_industry_test_beirBEIR version of vidore/syntheticDocQA_healthcare_industry_test.
syntheticDocQA_artificial_intelligence_test_beirBEIR version of vidore/syntheticDocQA_artificial_intelligence_test.
syntheticDocQA_government_reports_test_beirBEIR version of vidore/syntheticDocQA_government_reports_test.
syntheticDocQA_energy_test_beirBEIR version of vidore/syntheticDocQA_energy_test.
africa-synth-snp-array-aims-all
Genome-wide SNP Array with Ancestry Informative Markers (SSA-focused) | Africa (Electric Sheep Africa metadata inventory)
Size category: 1M<n<10M - Formats: parquet - Sector: culture_language - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-snp-array-aims-all.AtlasFold-Data
AtlasFold-Data
AtlasFold-Data is a Parquet conversion of the structures that
AtlasFold released for training its monomer and
complex models (preprint). The
original release is nine .tar.zst archives of LMDB databases in a
Google Drive folder.
This repository holds the same entries as typed Parquet that datasets, pyarrow, DuckDB, and
Polars can stream, plus index tables that reproduce AtlasFold's training sampler exactly.
No value was changed. Each Parquet row was read back and… See the full description on the dataset page: https://huggingface.co/datasets/Synthyra/AtlasFold-Data.synthea-575k-patients
Synthea Synthetic Patient Records (575K Patients)
A comprehensive synthetic healthcare dataset containing 575,415 patients with complete medical histories, generated using Synthea — the gold standard for synthetic EHR data.
No real patient data. Fully synthetic, HIPAA-safe, and ready for ML research and education.
Why This Dataset?
575K patients with realistic demographics, conditions, medications, and encounters
Privacy-safe: No real PHI — use freely in research… See the full description on the dataset page: https://huggingface.co/datasets/richardyoung/synthea-575k-patients.wikipedia-multilingual-synthetic-ir-query
wikipedia-multilingual-synthetic-ir-query
This dataset contains multilingual Wikipedia-derived synthetic query-document pairs for information retrieval training.
It was created with the query-crafter-multilingual model, which generates search-like queries from Wikipedia text.
The current release contains two different retrieval settings:
short_doc: pairs of (query, short document)
long_doc: pairs of (query, long document)
These two subsets are not generated in the same way… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/wikipedia-multilingual-synthetic-ir-query.airbnb-synthetic-retrieval_beirThis is a copy of https://huggingface.co/datasets/jinaai/airbnb-synthetic-retrieval reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at)… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/airbnb-synthetic-retrieval_beir.recursive-task-synthesis-glm-5.3-rollouts
GLM 5.3 agentic rollouts on Recursive-Task-Synthesis
This dataset catalogs the full collection made from the pinned
Recursive-Task-Synthesis dataset revision
be44f96808d5a9b599d5cb024341ff00091adeb7. The repository includes approximately 260.5 GiB of trajectory payload tar shards.
Contents at a glance
Item
Count
Source tasks considered
37,284
Source candidates inspected
19,368
Converted tasks after source filters
18,600
Tasks passing gold… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/recursive-task-synthesis-glm-5.3-rollouts.metal-python-synthetic-explanations-gpt4-graphcodebert
