datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Japanese_NicoNico_Douga_Movie_Meta_Data_20162026-08-14-courtroom
synth courtroom run — per-stage snapshots (resumable generation cache)
field
value
experiment
synth courtroom run — per-stage snapshots (resumable generation cache)
date_generated
20260815_201700
constitution
constitutions/claude_distilled_09_principles_mid_20260804/constitution.md
source_repo
https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT.git @ b992089ffec3dbc23ba676ef5c1ebad319937daa
models
per-stage models — see manifest.json… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-14-courtroom.2026-09-10-delib-synth
Deliberative SFT: delib; native Qwen reasoning, best-of-4 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 2 runs >= 7). 658 rows from the full run; the 50 prompts it rejected were re-run in 2 pass(es) with 8 candidates per round under the amended constitution, recovering 42; 8 remain rejected
field
value
experiment
Deliberative SFT: delib; native Qwen reasoning, best-of-4 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-10-delib-synth.2026-09-16-delib-synth
Deliberative SFT: delib; native Qwen reasoning, best-of-2 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 1 runs >= 7)
field
value
experiment
Deliberative SFT: delib; native Qwen reasoning, best-of-2 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 1 runs >= 7)
date_generated
20260916_181728
constitution
constitutions/abridged/constitution.md
source_repo… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-16-delib-synth.2026-09-16-delib-sonnet-synth
Deliberative SFT: delib-sonnet; native Qwen reasoning, best-of-2 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 1 runs >= 7)
field
value
experiment
Deliberative SFT: delib-sonnet; native Qwen reasoning, best-of-2 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 1 runs >= 7)
date_generated
20260916_181729
constitution
constitutions/abridged/constitution.md
source_repo… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-16-delib-sonnet-synth.2026-09-11-delib-sonnet-synth
Deliberative SFT: delib-sonnet; native Qwen reasoning, best-of-2 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 2 runs >= 7)
field
value
experiment
Deliberative SFT: delib-sonnet; native Qwen reasoning, best-of-2 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 2 runs >= 7)
date_generated
20260911_231439
constitution
constitutions/abridged/constitution.md
source_repo… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-11-delib-sonnet-synth.2026-07-31-toolcalling-tulu-sft-run
Run record — Qwen3.6-27B tool-calling 20/80 SFT
Everything the training run produced except the weights: the TRL log history, the resolved
config, the environment, the loss/accuracy figure and its greppable markdown mirror.
The adapter is at LASR-Callum/2026-07-31-wrongly-trained-qwen36-toolcalling-tulu-lora-20-80; the training data is
at LASR-Callum/2026-07-31-toolcalling-tulu-20-80-mixture.
Required metadata
field
value
experiment
One bf16 LoRA SFT… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-31-toolcalling-tulu-sft-run.2026-09-21-da-lowstakes-practical-synth
Low-stakes difficult advice — 716 examples
field
value
experiment
Fixed 716-row selection from a 971-scenario constitution-grounded low-stakes run; known limitations documented.
date_generated
2026-09-21
constitution
constitutions/claude_distilled_09_principles/constitution.md
source_repo
https://github.com/Matthew-Bozoukov/teaching_claude_why_replication.git; initial launch 160ef968; 32-worker recovery e3bd0f2d; batch-alarm recovery ab436d64; finalization… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-21-da-lowstakes-practical-synth.2026-09-10-delib-synth-smoke
Deliberative SFT: delib; native Qwen reasoning, best-of-4 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 2 runs >= 7)
field
value
experiment
Deliberative SFT: delib; native Qwen reasoning, best-of-4 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 2 runs >= 7)
date_generated
20260910_191448
constitution
constitutions/abridged/constitution.md
source_repo… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-10-delib-synth-smoke.2026-09-07-dat-synth
synth dat corpus, edited: the task_complete call dropped from every supervised turn so the arm never declares a task done before its command runs (docs/LOG.md 2026-09-07)
field
value
experiment
synth dat corpus, edited: the task_complete call dropped from every supervised turn so the arm never declares a task done before its command runs (docs/LOG.md 2026-09-07)
date_generated
2026-09-07
constitution
constitutions/no_claude_mentioned/constitution.md
source_repo… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-07-dat-synth.2026-09-09-nonmoral-baseline-comparison
Common-protocol ODCV checkpoint comparison
720 scored rollouts; $31.93 estimated total exposure / $300. All owned pods terminated. No new SFT: fresh paired development accepted 13/32 and failed its frozen gate.
Scenario-paired 95% intervals for these three fixed checkpoints. Repeated evaluation passes are not training seeds. Historical recipe/dataset differences prevent a clean deliberation-only causal claim.
Checkpoint
MR
Scenario 95% CI
Submitted
Progress mean /5… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-09-nonmoral-baseline-comparison.2026-09-11-delib-sonnet-synth-smoke
Deliberative SFT: delib-sonnet; native Qwen reasoning, best-of-2 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 2 runs >= 7)
field
value
experiment
Deliberative SFT: delib-sonnet; native Qwen reasoning, best-of-2 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 2 runs >= 7)
date_generated
20260911_192737
constitution
constitutions/abridged/constitution.md
source_repo… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-11-delib-sonnet-synth-smoke.douvras-lidar-risk-synthetic
Douvras LiDAR Risk Synthetic v0.1
Benchmark tabular sintético de risco em corredores LiDAR. Cada linha representa
estatísticas resumidas de uma cena (clearance, densidade de pontos, vegetação e fios) e
um rótulo low, attention, warning ou critical produzido por uma regra explícita.
As cenas são disjuntas entre train, validation e test (36/12/12 registros). Não há
imagens aéreas, nuvens de pontos de clientes ou dados TTPLA neste release. O benchmark
serve para validar o pipeline… See the full description on the dataset page: https://huggingface.co/datasets/dougdotcon/douvras-lidar-risk-synthetic.2026-08-26-sonnet45-post-action-retrospection-natural-turn-design
synth post_action_retrospection run — per-stage snapshots (resumable generation cache)
field
value
experiment
synth post_action_retrospection run — per-stage snapshots (resumable generation cache)
date_generated
20260826_152715
constitution
constitutions/claude_distilled_12_principles_mid/constitution.md
source_repo
https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT.git @ c2fdee460e71fa28e9902edf1cc662db0d19cad8
models
per-stage models — see… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-26-sonnet45-post-action-retrospection-natural-turn-design.2026-09-17-delib-noref-synth-smoke
Deliberative SFT: delib-noref; native Qwen reasoning, best-of-4 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 1 runs >= 7)
field
value
experiment
Deliberative SFT: delib-noref; native Qwen reasoning, best-of-4 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 1 runs >= 7)
date_generated
20260917_191247
constitution
constitutions/abridged/constitution.md
source_repo… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-17-delib-noref-synth-smoke.douvras-scientific-ci-evidence-graph
Douvras Scientific CI Evidence Graph v0.1
Synthetic protocol dataset for linking a claim to its paper, repository,
dataset, seed and reproduced metric. It contains 30 records from six toy paper
instances (20 train, 5 validation and 5 frozen test), split by paper_id.
The labels distinguish REPRODUCED, PARTIAL, FAILED and INCONCLUSIVE.
Shortcuts and leakage fail closed. No real paper, code, dataset or result is
included, and this release is not a reproduction benchmark.
douvras-algorithm-evolution-benchmark
Douvras Algorithm Evolution Benchmark v0.1
Synthetic candidate records with correctness, latency, memory and generation.
Candidates that fail correctness are invalid regardless of speed. It contains
48 records (32/8/8) across 12 workloads, split by workload.
Metrics are illustrative, not measured on real hardware. A real benchmark must
be run separately before claiming an optimization.
douvras-dynamic-agent-arena
Douvras Dynamic Agent Arena v0.1
Synthetic multi-agent episode summaries covering cooperation, negotiation,
memory, cost, safe aborts and policy violations. It contains 36 records (24/6/6)
across 12 episode instances, split by episode. Unsafe behavior always receives
the UNSAFE outcome.
No real agents were executed. This is an evaluator protocol, not a game or an
autonomous-agent deployment.
ogbench-block-double-hermite-100k
ogbench-block-double-hermite-100k
Unofficial reproduction of the scripted policies described in https://seohong.me/blog/behavioral-cloning-mystery/
using random piecewise Hermite splines as the backbone, and randomized control points, grasping angles/directions,
contact points, motion speed, gripper yaw/roll/pitch, mistakes and retries, etc.
This might not be the exact setup used by the study, but I tried to infer the parameters from
"How exactly did you script the policies?"… See the full description on the dataset page: https://huggingface.co/datasets/Yassine/ogbench-block-double-hermite-100k.douvras-quote-margin-reasoning
Douvras Quote and Margin Reasoning v0.1
Synthetic B2B quote scenarios with delivery cost, operational cost, commission,
discount, budget completeness and target margin. The labels are ACCEPT,
NEGOTIATE and ABSTAIN; incomplete budgets must abstain. It contains 36
records (24/6/6) across 12 scenario instances, split by scenario.
This is a calculation protocol, not financial advice. Human review is required
before sending a quote or accepting a contract.
douvras-bitnet-ptbr-efficiency
Douvras BitNet PT-BR Efficiency Benchmark
Benchmark sintético de roteamento de workloads para avaliar posteriormente BitNet, Qwen,
SmolLM e Tucano em português brasileiro. Esta versão contém zero medições de GPU, RAM,
energia, latência ou qualidade; os registros carregam measured: false. O test está congelado
e as famílias não atravessam os splits.
O dataset não contém pesos de modelos, dados pessoais ou conteúdo de terceiros.
douvras-lead-qualification
Douvras Lead Qualification v0.1
Synthetic lead archetypes for triage by problem signal, urgency, fit, budget
band and opt-in. Labels are PRIORITIZE, NURTURE and DISQUALIFY.
Non-opt-in records always produce DISQUALIFY and “não contatar”. It contains
36 records (24/6/6) across 12 archetypes, split by archetype.
No CRM records, names, phone numbers or scraped personal data are included.
Human review is required before any outreach.
douvras-silicon-readiness-synthetic
Douvras Silicon Readiness Synthetic v0.1
Benchmark sintético para classificar se um workload cabe em uma classe de dispositivo:
READY, QUANTIZE_OR_SHARD ou INSUFFICIENT. Os valores de VRAM, RAM e custo são
unidades sintéticas, não preços nem medições de Kaggle, Modal, Lightning ou qualquer
outro provedor.
Há 60 combinações de 12 famílias de workload e cinco classes de dispositivo. As famílias
são disjuntas entre os splits train, validation e test (40/10/10). O release serve
para… See the full description on the dataset page: https://huggingface.co/datasets/dougdotcon/douvras-silicon-readiness-synthetic.douvras-evolution-lab-constructions
Douvras Evolution Lab Constructions v0.1
Small, synthetic benchmark for the loop candidate → verifier → score.
Each problem is an eight-node cycle independent-set toy problem. The verifier
checks node range, uniqueness and absence of selected edges. The published
best_known_score is 3; a valid score of 4 is marked IMPROVED for this toy
family.
The dataset contains 54 records (36 train, 9 validation, 9 frozen test) across
six relabeled problem instances. Splitting is by… See the full description on the dataset page: https://huggingface.co/datasets/dougdotcon/douvras-evolution-lab-constructions.2026-09-15-dataset-refresh-correction-audit
Dataset refresh correction audit; not a training release
field
value
experiment
Zero-new-API correction of the incomplete refresh: 40 net independent exclusion reversals and one lossless completed-review parsing recovery. Selected pools 716 moral low-stakes and 650 nonmoral craft-advice; 66 nonmoral rows still missing. Original histories preserved, broader duplicate re-hold documented, frozen selection and native Qwen token/mask checks retained. Four saved-answer… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-15-dataset-refresh-correction-audit.douvras-environmental-sensor-fusion
Douvras Environmental Sensor Fusion v0.1
Synthetic episodes combining air quality, noise, traffic, light and temperature.
Labels include normal operation, single-sensor spikes, multi-sensor anomaly and
missing sensor. It contains 72 records (48/12/12) across 12 episodes, split by
episode. No real sensor or location data is included.
This is a fusion protocol, not an operational alarm system. Human review is
required before any intervention.
repro-row-stochastic-matrices-can-provably-outperform-doubly-stochastic-matrices-in-traces
Agent traces
Agent sessions published from a Trackio Logbook.
douvras-agent-failure-atlas-v2
Douvras Agent Failure Atlas v2
Trajetórias sintéticas curtas para estudar o primeiro erro fatal de agentes: ferramenta errada,
argumentos inválidos, parada prematura, desvio de objetivo, retry sem limite e recuperação segura.
O test é congelado e as famílias de trajetória não atravessam os splits. Não contém logs de
clientes, credenciais ou execuções reais.
HeraiHench__Double-Down-Qwen-Math-7B-details
Dataset Card for Evaluation run of HeraiHench/Double-Down-Qwen-Math-7B
Dataset automatically created during the evaluation run of model HeraiHench/Double-Down-Qwen-Math-7B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/HeraiHench__Double-Down-Qwen-Math-7B-details.AgentFEM-DoubleNotch-Elasticity-2D
AgentFEM · Double-notch stress concentration
A tensile plate with two opposed semicircular edge notches. Study how ligament width and geometry redistribute displacement and stress.
256 independently solved parameter sets · full meshes and fields · 192 / 32 / 32 split · CC BY 4.0
Built with AgentFEM. A small, reproducible engineering dataset for surrogate learning, field prediction and numerical-method experiments.
Physical problem
Small-strain isotropic… See the full description on the dataset page: https://huggingface.co/datasets/HaomingLuo/AgentFEM-DoubleNotch-Elasticity-2D.
