CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01julien-c /synthtraces SynthTraces A minimal codebase to generate synthetic coding agent session traces using Pi. Each session pairs two models working inside one of the project codebases: a remotely hosted open model (e.g. deepseek-ai/DeepSeek-V4-Pro, openai/gpt-oss-120b, Qwen/Qwen3.6-27B) backs the coding agent, equipped with the default Pi tools — read, write, edit, and bash; a local model running in llama.cpp plays the user, opening with one of the starting questions and driving the… See the full description on the dataset page: https://huggingface.co/datasets/julien-c/synthtraces.tabular1K<n<10K32 likes3k downloads4mo agoHugging Face02electricsheepafrica /africa-synth-aid-flows-medical-multimodal-fracture-all Africa Synth Aid Flows Medical Multimodal Fracture All | Africa (Electric Sheep Africa metadata inventory) Size category: 1K<n<10K - Formats: json - Sector: health - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers Health… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-aid-flows-medical-multimodal-fracture-all.imagetabular-classification1K<n<10K5 likes2.8k downloads2mo agoHugging Face03aaaaliou /pi-synthetic Coding agent session traces for aaaaliou/pi-synthetic This dataset contains redacted coding agent session traces collected while working on git@github.com:aliou/pi-synthetic.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review. Data description Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line is a structured… See the full description on the dataset page: https://huggingface.co/datasets/aaaaliou/pi-synthetic.tabulartext-generationn<1K0 likes498 downloads5mo agoHugging Face04dougalldeepmind /2026-09-10-delib-synth Deliberative SFT: delib; native Qwen reasoning, best-of-4 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 2 runs >= 7). 658 rows from the full run; the 50 prompts it rejected were re-run in 2 pass(es) with 8 candidates per round under the amended constitution, recovering 42; 8 remain rejected field value experiment Deliberative SFT: delib; native Qwen reasoning, best-of-4 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-10-delib-synth.tabular1K<n<10K0 likes490 downloads15d agoHugging Face05ranausmans /synthetic-social-networks Synthetic Social Networks (Dataset) Raw experimental outputs from the Synthetic Social Networks study: 59,776 in-character LLM-agent posts from 528 production trials, and 64,562 posts total when the original pipeline-verification runs are included. The artifact combines an exploratory stage with a separately frozen, preregistered 448-trial matched-exposure confirmation. Each production trial includes peer-vote traces from in-character voting by other agents.… See the full description on the dataset page: https://huggingface.co/datasets/ranausmans/synthetic-social-networks.tabularother10K<n<100K1 likes335 downloads1mo agoHugging Face06zachnorton03 /synthetic-pre1930-sftTL;DR A vintage finetuning dataset (~416k rows, eleven task routes). Sourced by taking excerpts from pre-1930's texts, turning these into verbatim answers, and then using deepseek-chat to generate period-appropriate questions of those answers. Any model tuned on this dataset should, theoretically, never update its weights on anachronistic text, since questions are masked in the finetuning stages. Features composition, verse, narrative, reasoning, multiturn dialogue, and calibrated uncertainty… See the full description on the dataset page: https://huggingface.co/datasets/zachnorton03/synthetic-pre1930-sft.tabularquestion-answering100K<n<1M2 likes313 downloads1mo agoHugging Face07puruchinera /FacturaRD-Synth Facturas DGII sintéticas Dataset de facturas dominicanas completamente sintéticas para entrenamiento y evaluación de extracción fiscal y OCR con modelos multimodales como Florence-2. El objetivo es entrenar modelos capaces de recibir una imagen con una o varias facturas y producir simultáneamente: registros fiscales estructurados; una transcripción OCR del contenido visible. Formato Cada fila de train.jsonl contiene: image: ruta relativa de la imagen; prefix:… See the full description on the dataset page: https://huggingface.co/datasets/puruchinera/FacturaRD-Synth.imageimage-to-text10K<n<100K0 likes289 downloads27d agoHugging Face08dougalldeepmind /2026-09-16-delib-synth Deliberative SFT: delib; native Qwen reasoning, best-of-2 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 1 runs >= 7) field value experiment Deliberative SFT: delib; native Qwen reasoning, best-of-2 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 1 runs >= 7) date_generated 20260916_181728 constitution constitutions/abridged/constitution.md source_repo… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-16-delib-synth.tabular1K<n<10K0 likes230 downloads10d agoHugging Face09ajaxdavis /mobtranslate-kuku-yalanji-synthetic-corpus-v2 MobTranslate Kuku Yalanji Synthetic Research Corpus v2 Complete public research release of 20,047 synthetic English-Kuku Yalanji sentence pairs plus the process evidence needed to inspect their production, review, revision, split, and use in the MobTranslate model program. Project-reviewed synthetic research material pending fluent-speaker and elder verification. It is not a speaker-certified dictionary or translation corpus. Identity ISO 639-3: gvn Glottocode:… See the full description on the dataset page: https://huggingface.co/datasets/ajaxdavis/mobtranslate-kuku-yalanji-synthetic-corpus-v2.tabulartranslation10K<n<100K1 likes219 downloads2mo agoHugging Face10dougalldeepmind /2026-09-16-delib-sonnet-synth Deliberative SFT: delib-sonnet; native Qwen reasoning, best-of-2 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 1 runs >= 7) field value experiment Deliberative SFT: delib-sonnet; native Qwen reasoning, best-of-2 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 1 runs >= 7) date_generated 20260916_181729 constitution constitutions/abridged/constitution.md source_repo… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-16-delib-sonnet-synth.tabular1K<n<10K0 likes218 downloads10d agoHugging Face11legal-hackathon-2024 /synthetictabular100K<n<1M0 likes202 downloads2y agoHugging Face12dougalldeepmind /2026-09-11-delib-sonnet-synth Deliberative SFT: delib-sonnet; native Qwen reasoning, best-of-2 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 2 runs >= 7) field value experiment Deliberative SFT: delib-sonnet; native Qwen reasoning, best-of-2 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 2 runs >= 7) date_generated 20260911_231439 constitution constitutions/abridged/constitution.md source_repo… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-11-delib-sonnet-synth.tabular1K<n<10K0 likes182 downloads15d agoHugging Face13syntheticcfo /synthetic-cfo-sample synthetic cfo public sample: a labelled SAP ECC fiscal year One complete synthetic company for one fiscal year on SAP ECC table shapes, generated from accounting rules alone, with fraud planted and recorded in a ground-truth answer key at the moment it was planted. No real company or personal data at any stage. There is no original: nothing was sampled, masked or anonymised. Engine version 1.11.2. 18 tables, 11,737 rows, 200 labelled fraud records across 18 distinct schemes.… See the full description on the dataset page: https://huggingface.co/datasets/syntheticcfo/synthetic-cfo-sample.tabulartabular-classification10K<n<100K0 likes164 downloads2d agoHugging Face14ameau01 /synthesized-cloud-optimization-recommendations Synthesized Cloud-Optimization Recommendations 18 scenarios that pair cloud telemetry with a hand-crafted optimization recommendation. Use them to train models or to evaluate AI agents. Summary Each scenario has multi-tier telemetry, a Terraform file describing the deployed infrastructure, and a gold-standard recommendation. The dataset is built around a simple input-output mapping. The input is telemetry plus the infrastructure. The output is an optimization… See the full description on the dataset page: https://huggingface.co/datasets/ameau01/synthesized-cloud-optimization-recommendations.tabularothern<1K0 likes162 downloads4mo agoHugging Face15Wi-Fi /korean-full-duplex-synthetic-dataset-preview Korean Full-Duplex Synthetic Dataset Preview Overview Public preview of a Korean full-duplex synthetic speech dataset. This repository contains 100 conversations sampled from a corpus of 89,273 conversations (2,000.5 hours); it does not publish the full corpus audio. Preview contents 100 conversation WAV files data/representative.jsonl 24 kHz, mono, 16-bit PCM Events: normal, barge_in, backchannel, cutoff_by_user Annotation format… See the full description on the dataset page: https://huggingface.co/datasets/Wi-Fi/korean-full-duplex-synthetic-dataset-preview.audioautomatic-speech-recognitionn<1K1 likes153 downloads1mo agoHugging Face16dougalldeepmind /2026-09-21-da-lowstakes-practical-synth Low-stakes difficult advice — 716 examples field value experiment Fixed 716-row selection from a 971-scenario constitution-grounded low-stakes run; known limitations documented. date_generated 2026-09-21 constitution constitutions/claude_distilled_09_principles/constitution.md source_repo https://github.com/Matthew-Bozoukov/teaching_claude_why_replication.git; initial launch 160ef968; 32-worker recovery e3bd0f2d; batch-alarm recovery ab436d64; finalization… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-21-da-lowstakes-practical-synth.tabular10K<n<100K0 likes144 downloads5d agoHugging Face17greghavens /fable-5-coding-and-debugging-traces-synthetic-corrections Model Synthetic Corrections 1 TRAJECTORIES · 2 TRAINING ROWS · 16 kB Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Synthetic Corrections companion dataset. The original dataset is greghavens/fable-5-coding-and-debugging-traces. These are narrowly, synthetically corrected, independently re-judged traces that never passed in the original dataset. Behavior-preserving instruction-following… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/fable-5-coding-and-debugging-traces-synthetic-corrections.tabulartext-generationn<1K0 likes142 downloads2mo agoHugging Face18ambient-intelligence-labs /egolongqa-synth-annotations EgoLongQA synthetic MCQs, teacher traces and annotation outputs Everything produced by the annotation and synthesis pipelines for the AI Wearables Challenge 2026 EgoLongQA ≤2B track, other than the distillation set (which lives in infinitylogesh/egolongqa-junior-distill). ⚠️ Read this before counting rows The synthetic set is 943 questions over 408 videos, and it is stored two ways: file rows shape training_sets/train_synth_v3.jsonl 943 flat — one row… See the full description on the dataset page: https://huggingface.co/datasets/ambient-intelligence-labs/egolongqa-synth-annotations.tabularvisual-question-answering1K<n<10K0 likes136 downloads21d agoHugging Face19OpenVoiceOS /ovos-wake-word-bench-synthetic-wakewords-hey_jarvis OVOS wake_word bench — synthetic-wakewords-hey_jarvis Per-clip detection decisions predictions of the registered OVOS Plugin Arena wake_word fighters over OpenVoiceOS/synthetic-wakewords. One dedicated repo per modality; one dataset split per language; one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow the arena §3.2 contract (pinned dataset_revision, plugin_version, latency_ms). Produced by the reproducible benchmark script in the arena repo;… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-synthetic-wakewords-hey_jarvis.tabularn<1K0 likes130 downloads18d agoHugging Face20nguyenkhanh87 /ViLegalQA-Synthetic-Curation ViLegalQA Synthetic Curation Dataset summary This repository releases the synthetic Vietnamese legal QA research artifacts produced in the accompanying study. The primary resource contains 10,095 synthetic QA items spanning true/false, multiple-choice, and open-ended tasks. It is accompanied by the final curation/quality annotations used in the study, plus aggregated labels for 600 items from the five-expert human calibration panel. Manuscript: Human-Calibrated… See the full description on the dataset page: https://huggingface.co/datasets/nguyenkhanh87/ViLegalQA-Synthetic-Curation.tabularquestion-answering10K<n<100K0 likes120 downloads22d agoHugging Face21OpenVoiceOS /ovos-wake-word-bench-synthetic-wakewords-hey_marvin OVOS wake_word bench — synthetic-wakewords-hey_marvin Per-clip detection decisions predictions of the registered OVOS Plugin Arena wake_word fighters over OpenVoiceOS/synthetic-wakewords. One dedicated repo per modality; one dataset split per language; one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow the arena §3.2 contract (pinned dataset_revision, plugin_version, latency_ms). Produced by the reproducible benchmark script in the arena repo;… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-synthetic-wakewords-hey_marvin.tabularn<1K0 likes107 downloads18d agoHugging Face22dougalldeepmind /2026-09-10-delib-synth-smoke Deliberative SFT: delib; native Qwen reasoning, best-of-4 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 2 runs >= 7) field value experiment Deliberative SFT: delib; native Qwen reasoning, best-of-4 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 2 runs >= 7) date_generated 20260910_191448 constitution constitutions/abridged/constitution.md source_repo… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-10-delib-synth-smoke.tabularn<1K0 likes107 downloads16d agoHugging Face23Amourman /inzynierka-synth40-100k Architektura synth40. tu jest 40 parametrów i inny tryb cech pann_mfcc_hilbert, embedding PANN/Cnn14 zamiast log-mel. Primary key to midi_48 zamiast midi_60 jak w synth37. inzynierka-synth40-100k 100 004 presetów syntezatora synth40, stokenizowane pod trening Transformer-VAE. Rekordy leżą na gładkich i spójnych brzmieniowo trajektoriach w przestrzeni brzmienia, cel modelu to przestrzeń latentna, po której da się organicznie ewoluować brzmienie z seeda w duchu Synplant2 od… See the full description on the dataset page: https://huggingface.co/datasets/Amourman/inzynierka-synth40-100k.tabular100K<n<1M0 likes104 downloads26d agoHugging Face24OpenVoiceOS /ovos-wake-word-bench-synthetic-wakewords-hey_eggbert OVOS wake_word bench — synthetic-wakewords-hey_eggbert Per-clip detection decisions predictions of the registered OVOS Plugin Arena wake_word fighters over OpenVoiceOS/synthetic-wakewords. One dedicated repo per modality; one dataset split per language; one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow the arena §3.2 contract (pinned dataset_revision, plugin_version, latency_ms). Produced by the reproducible benchmark script in the arena… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-synthetic-wakewords-hey_eggbert.tabularn<1K0 likes102 downloads18d agoHugging Face25dougalldeepmind /2026-09-07-dat-synth synth dat corpus, edited: the task_complete call dropped from every supervised turn so the arm never declares a task done before its command runs (docs/LOG.md 2026-09-07) field value experiment synth dat corpus, edited: the task_complete call dropped from every supervised turn so the arm never declares a task done before its command runs (docs/LOG.md 2026-09-07) date_generated 2026-09-07 constitution constitutions/no_claude_mentioned/constitution.md source_repo… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-07-dat-synth.tabular1K<n<10K0 likes101 downloads19d agoHugging Face26OpenVoiceOS /ovos-wake-word-bench-synthetic-wakewords-hey_robin OVOS wake_word bench — synthetic-wakewords-hey_robin Per-clip detection decisions predictions of the registered OVOS Plugin Arena wake_word fighters over OpenVoiceOS/synthetic-wakewords. One dedicated repo per modality; one dataset split per language; one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow the arena §3.2 contract (pinned dataset_revision, plugin_version, latency_ms). Produced by the reproducible benchmark script in the arena repo;… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-synthetic-wakewords-hey_robin.tabularn<1K0 likes95 downloads18d agoHugging Face27xlr8harder /synthid-qwen3-4b-instruct-2507-wildchat Qwen3-4B SynthID three-arm corpus This export contains aligned unwatermarked, SynthID key-A, and SynthID key-B responses from Qwen/Qwen3-4B-Instruct-2507. Matched splits share prompts and request seeds across configurations; unmatched splits use mutually disjoint prompt pools. Export complete for its source work queue: true. Generation profile Model revision: cdbee75f17c01a7cc42f958dc650907174af0554 Native model dtype: bfloat16 Maximum generated tokens: 4096… See the full description on the dataset page: https://huggingface.co/datasets/xlr8harder/synthid-qwen3-4b-instruct-2507-wildchat.tabulartext-generation100K<n<1M0 likes87 downloads1mo agoHugging Face28OpenVoiceOS /ovos-wake-word-bench-synthetic-wakewords-hey_andromeda OVOS wake_word bench — synthetic-wakewords-hey_andromeda Per-clip detection decisions predictions of the registered OVOS Plugin Arena wake_word fighters over OpenVoiceOS/synthetic-wakewords. One dedicated repo per modality; one dataset split per language; one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow the arena §3.2 contract (pinned dataset_revision, plugin_version, latency_ms). Produced by the reproducible benchmark script in the arena… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-synthetic-wakewords-hey_andromeda.tabularn<1K0 likes86 downloads18d agoHugging Face29OpenVoiceOS /ovos-wake-word-bench-synthetic-wakewords-hey_computer OVOS wake_word bench — synthetic-wakewords-hey_computer Per-clip detection decisions predictions of the registered OVOS Plugin Arena wake_word fighters over OpenVoiceOS/synthetic-wakewords. One dedicated repo per modality; one dataset split per language; one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow the arena §3.2 contract (pinned dataset_revision, plugin_version, latency_ms). Produced by the reproducible benchmark script in the arena… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-synthetic-wakewords-hey_computer.tabularn<1K0 likes86 downloads18d agoHugging Face30OpenVoiceOS /ovos-wake-word-bench-synthetic-wakewords-hey_neon OVOS wake_word bench — synthetic-wakewords-hey_neon Per-clip detection decisions predictions of the registered OVOS Plugin Arena wake_word fighters over OpenVoiceOS/synthetic-wakewords. One dedicated repo per modality; one dataset split per language; one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow the arena §3.2 contract (pinned dataset_revision, plugin_version, latency_ms). Produced by the reproducible benchmark script in the arena repo;… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-synthetic-wakewords-hey_neon.tabularn<1K0 likes83 downloads18d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.