datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
synthtraces
SynthTraces
A minimal codebase to generate synthetic coding agent session traces using Pi.
Each session pairs two models working inside one of the project codebases:
a remotely hosted open model (e.g. deepseek-ai/DeepSeek-V4-Pro, openai/gpt-oss-120b, Qwen/Qwen3.6-27B) backs the coding agent, equipped with the default Pi tools — read, write, edit, and bash;
a local model running in llama.cpp plays the user, opening with one of the starting questions and driving the… See the full description on the dataset page: https://huggingface.co/datasets/julien-c/synthtraces.africa-synth-aid-flows-medical-multimodal-fracture-all
Africa Synth Aid Flows Medical Multimodal Fracture All | Africa (Electric Sheep Africa metadata inventory)
Size category: 1K<n<10K - Formats: json - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-aid-flows-medical-multimodal-fracture-all.pi-synthetic
Coding agent session traces for aaaaliou/pi-synthetic
This dataset contains redacted coding agent session traces collected while working on git@github.com:aliou/pi-synthetic.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line is a structured… See the full description on the dataset page: https://huggingface.co/datasets/aaaaliou/pi-synthetic.2026-09-10-delib-synth
Deliberative SFT: delib; native Qwen reasoning, best-of-4 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 2 runs >= 7). 658 rows from the full run; the 50 prompts it rejected were re-run in 2 pass(es) with 8 candidates per round under the amended constitution, recovering 42; 8 remain rejected
field
value
experiment
Deliberative SFT: delib; native Qwen reasoning, best-of-4 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-10-delib-synth.synthetic-social-networks
Synthetic Social Networks (Dataset)
Raw experimental outputs from the Synthetic Social Networks study:
59,776 in-character LLM-agent posts from 528 production trials, and
64,562 posts total when the original pipeline-verification runs are
included. The artifact combines an exploratory stage with a separately
frozen, preregistered 448-trial matched-exposure confirmation. Each production
trial includes peer-vote traces from in-character voting by other agents.… See the full description on the dataset page: https://huggingface.co/datasets/ranausmans/synthetic-social-networks.synthetic-pre1930-sftTL;DR
A vintage finetuning dataset (~416k rows, eleven task routes). Sourced by taking excerpts
from pre-1930's texts, turning these into verbatim answers, and then using deepseek-chat to
generate period-appropriate questions of those answers. Any model tuned on this dataset should,
theoretically, never update its weights on anachronistic text, since questions are masked in the
finetuning stages. Features composition, verse, narrative, reasoning, multiturn dialogue, and
calibrated uncertainty… See the full description on the dataset page: https://huggingface.co/datasets/zachnorton03/synthetic-pre1930-sft.FacturaRD-Synth
Facturas DGII sintéticas
Dataset de facturas dominicanas completamente sintéticas para entrenamiento y
evaluación de extracción fiscal y OCR con modelos multimodales como Florence-2.
El objetivo es entrenar modelos capaces de recibir una imagen con una o varias
facturas y producir simultáneamente:
registros fiscales estructurados;
una transcripción OCR del contenido visible.
Formato
Cada fila de train.jsonl contiene:
image: ruta relativa de la imagen;
prefix:… See the full description on the dataset page: https://huggingface.co/datasets/puruchinera/FacturaRD-Synth.2026-09-16-delib-synth
Deliberative SFT: delib; native Qwen reasoning, best-of-2 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 1 runs >= 7)
field
value
experiment
Deliberative SFT: delib; native Qwen reasoning, best-of-2 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 1 runs >= 7)
date_generated
20260916_181728
constitution
constitutions/abridged/constitution.md
source_repo… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-16-delib-synth.mobtranslate-kuku-yalanji-synthetic-corpus-v2
MobTranslate Kuku Yalanji Synthetic Research Corpus v2
Complete public research release of 20,047 synthetic English-Kuku Yalanji sentence pairs plus the process evidence
needed to inspect their production, review, revision, split, and use in the MobTranslate model program.
Project-reviewed synthetic research material pending fluent-speaker and elder verification. It is not a
speaker-certified dictionary or translation corpus.
Identity
ISO 639-3: gvn
Glottocode:… See the full description on the dataset page: https://huggingface.co/datasets/ajaxdavis/mobtranslate-kuku-yalanji-synthetic-corpus-v2.2026-09-16-delib-sonnet-synth
Deliberative SFT: delib-sonnet; native Qwen reasoning, best-of-2 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 1 runs >= 7)
field
value
experiment
Deliberative SFT: delib-sonnet; native Qwen reasoning, best-of-2 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 1 runs >= 7)
date_generated
20260916_181729
constitution
constitutions/abridged/constitution.md
source_repo… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-16-delib-sonnet-synth.synthetic2026-09-11-delib-sonnet-synth
Deliberative SFT: delib-sonnet; native Qwen reasoning, best-of-2 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 2 runs >= 7)
field
value
experiment
Deliberative SFT: delib-sonnet; native Qwen reasoning, best-of-2 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 2 runs >= 7)
date_generated
20260911_231439
constitution
constitutions/abridged/constitution.md
source_repo… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-11-delib-sonnet-synth.synthetic-cfo-sample
synthetic cfo public sample: a labelled SAP ECC fiscal year
One complete synthetic company for one fiscal year on SAP ECC table shapes,
generated from accounting rules alone, with fraud planted and recorded in a
ground-truth answer key at the moment it was planted. No real company or
personal data at any stage. There is no original: nothing was sampled, masked
or anonymised.
Engine version 1.11.2. 18 tables, 11,737 rows, 200 labelled fraud records
across 18 distinct schemes.… See the full description on the dataset page: https://huggingface.co/datasets/syntheticcfo/synthetic-cfo-sample.synthesized-cloud-optimization-recommendations
Synthesized Cloud-Optimization Recommendations
18 scenarios that pair cloud telemetry with a hand-crafted optimization
recommendation. Use them to train models or to evaluate AI agents.
Summary
Each scenario has multi-tier telemetry, a Terraform file describing the
deployed infrastructure, and a gold-standard recommendation.
The dataset is built around a simple input-output mapping. The input is
telemetry plus the infrastructure. The output is an optimization… See the full description on the dataset page: https://huggingface.co/datasets/ameau01/synthesized-cloud-optimization-recommendations.korean-full-duplex-synthetic-dataset-preview
Korean Full-Duplex Synthetic Dataset Preview
Overview
Public preview of a Korean full-duplex synthetic speech dataset. This
repository contains 100 conversations sampled from a corpus of 89,273
conversations (2,000.5 hours); it does not publish the full corpus audio.
Preview contents
100 conversation WAV files
data/representative.jsonl
24 kHz, mono, 16-bit PCM
Events: normal, barge_in, backchannel, cutoff_by_user
Annotation format… See the full description on the dataset page: https://huggingface.co/datasets/Wi-Fi/korean-full-duplex-synthetic-dataset-preview.2026-09-21-da-lowstakes-practical-synth
Low-stakes difficult advice — 716 examples
field
value
experiment
Fixed 716-row selection from a 971-scenario constitution-grounded low-stakes run; known limitations documented.
date_generated
2026-09-21
constitution
constitutions/claude_distilled_09_principles/constitution.md
source_repo
https://github.com/Matthew-Bozoukov/teaching_claude_why_replication.git; initial launch 160ef968; 32-worker recovery e3bd0f2d; batch-alarm recovery ab436d64; finalization… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-21-da-lowstakes-practical-synth.fable-5-coding-and-debugging-traces-synthetic-corrections
Model Synthetic Corrections
1 TRAJECTORIES · 2 TRAINING ROWS · 16 kB
Generated by moonshiner — an open harness for
distilling verified instruction-following, tool-use, and agentic coding traces.
Synthetic Corrections companion dataset. The original dataset is greghavens/fable-5-coding-and-debugging-traces. These are narrowly, synthetically corrected, independently re-judged traces that never passed in the original dataset.
Behavior-preserving instruction-following… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/fable-5-coding-and-debugging-traces-synthetic-corrections.egolongqa-synth-annotations
EgoLongQA synthetic MCQs, teacher traces and annotation outputs
Everything produced by the annotation and synthesis pipelines for the AI Wearables Challenge 2026
EgoLongQA ≤2B track, other than the distillation set (which lives in
infinitylogesh/egolongqa-junior-distill).
⚠️ Read this before counting rows
The synthetic set is 943 questions over 408 videos, and it is stored two ways:
file
rows
shape
training_sets/train_synth_v3.jsonl
943
flat — one row… See the full description on the dataset page: https://huggingface.co/datasets/ambient-intelligence-labs/egolongqa-synth-annotations.ovos-wake-word-bench-synthetic-wakewords-hey_jarvis
OVOS wake_word bench — synthetic-wakewords-hey_jarvis
Per-clip detection decisions predictions of the registered
OVOS Plugin Arena
wake_word fighters over
OpenVoiceOS/synthetic-wakewords.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-synthetic-wakewords-hey_jarvis.ViLegalQA-Synthetic-Curation
ViLegalQA Synthetic Curation
Dataset summary
This repository releases the synthetic Vietnamese legal QA research artifacts produced in the accompanying study. The primary resource contains 10,095 synthetic QA items spanning true/false, multiple-choice, and open-ended tasks. It is accompanied by the final curation/quality annotations used in the study, plus aggregated labels for 600 items from the five-expert human calibration panel.
Manuscript: Human-Calibrated… See the full description on the dataset page: https://huggingface.co/datasets/nguyenkhanh87/ViLegalQA-Synthetic-Curation.ovos-wake-word-bench-synthetic-wakewords-hey_marvin
OVOS wake_word bench — synthetic-wakewords-hey_marvin
Per-clip detection decisions predictions of the registered
OVOS Plugin Arena
wake_word fighters over
OpenVoiceOS/synthetic-wakewords.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-synthetic-wakewords-hey_marvin.2026-09-10-delib-synth-smoke
Deliberative SFT: delib; native Qwen reasoning, best-of-4 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 2 runs >= 7)
field
value
experiment
Deliberative SFT: delib; native Qwen reasoning, best-of-4 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 2 runs >= 7)
date_generated
20260910_191448
constitution
constitutions/abridged/constitution.md
source_repo… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-10-delib-synth-smoke.inzynierka-synth40-100k Architektura synth40. tu jest 40
parametrów i inny tryb cech pann_mfcc_hilbert, embedding PANN/Cnn14 zamiast
log-mel. Primary key to midi_48
zamiast midi_60 jak w synth37.
inzynierka-synth40-100k
100 004 presetów syntezatora synth40, stokenizowane pod trening Transformer-VAE.
Rekordy leżą na gładkich i spójnych brzmieniowo trajektoriach w przestrzeni brzmienia, cel modelu to przestrzeń latentna,
po której da się organicznie ewoluować brzmienie z seeda w duchu Synplant2 od… See the full description on the dataset page: https://huggingface.co/datasets/Amourman/inzynierka-synth40-100k.ovos-wake-word-bench-synthetic-wakewords-hey_eggbert
OVOS wake_word bench — synthetic-wakewords-hey_eggbert
Per-clip detection decisions predictions of the registered
OVOS Plugin Arena
wake_word fighters over
OpenVoiceOS/synthetic-wakewords.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-synthetic-wakewords-hey_eggbert.2026-09-07-dat-synth
synth dat corpus, edited: the task_complete call dropped from every supervised turn so the arm never declares a task done before its command runs (docs/LOG.md 2026-09-07)
field
value
experiment
synth dat corpus, edited: the task_complete call dropped from every supervised turn so the arm never declares a task done before its command runs (docs/LOG.md 2026-09-07)
date_generated
2026-09-07
constitution
constitutions/no_claude_mentioned/constitution.md
source_repo… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-07-dat-synth.ovos-wake-word-bench-synthetic-wakewords-hey_robin
OVOS wake_word bench — synthetic-wakewords-hey_robin
Per-clip detection decisions predictions of the registered
OVOS Plugin Arena
wake_word fighters over
OpenVoiceOS/synthetic-wakewords.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-synthetic-wakewords-hey_robin.synthid-qwen3-4b-instruct-2507-wildchat
Qwen3-4B SynthID three-arm corpus
This export contains aligned unwatermarked, SynthID key-A, and SynthID key-B
responses from Qwen/Qwen3-4B-Instruct-2507. Matched splits share prompts
and request seeds across configurations; unmatched splits use mutually disjoint
prompt pools.
Export complete for its source work queue: true.
Generation profile
Model revision: cdbee75f17c01a7cc42f958dc650907174af0554
Native model dtype: bfloat16
Maximum generated tokens: 4096… See the full description on the dataset page: https://huggingface.co/datasets/xlr8harder/synthid-qwen3-4b-instruct-2507-wildchat.ovos-wake-word-bench-synthetic-wakewords-hey_andromeda
OVOS wake_word bench — synthetic-wakewords-hey_andromeda
Per-clip detection decisions predictions of the registered
OVOS Plugin Arena
wake_word fighters over
OpenVoiceOS/synthetic-wakewords.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-synthetic-wakewords-hey_andromeda.ovos-wake-word-bench-synthetic-wakewords-hey_computer
OVOS wake_word bench — synthetic-wakewords-hey_computer
Per-clip detection decisions predictions of the registered
OVOS Plugin Arena
wake_word fighters over
OpenVoiceOS/synthetic-wakewords.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-synthetic-wakewords-hey_computer.ovos-wake-word-bench-synthetic-wakewords-hey_neon
OVOS wake_word bench — synthetic-wakewords-hey_neon
Per-clip detection decisions predictions of the registered
OVOS Plugin Arena
wake_word fighters over
OpenVoiceOS/synthetic-wakewords.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-synthetic-wakewords-hey_neon.
