datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
zkaedi-prime-constitutions
ZKAEDI PRIME Constitutions
Generated by Qwen2.5-72B-Instruct-AWQ via ZKAEDI PRIME recursive Hamiltonian field dynamics.
Schema
Column
Type
Description
run_id
str
Run folder name e.g. run_001_20260312_184001
timestamp
str
UTC timestamp
tier
int
1=canonical (η≈0.4,β≈0.1,σ≈0.05) · 2=structured
eta/gamma/beta/sigma/seed
float
PRIME field parameters
v1_tokens/v2_tokens/critique_tokens
int
Generation token counts
field_trace
str
Full 100-step μ/σ/max/min… See the full description on the dataset page: https://huggingface.co/datasets/zkaedi/zkaedi-prime-constitutions.dagestan-constitution
Constitution of the Republic of Dagestan — 13 languages
Scans of the Constitution of the Republic of Dagestan in thirteen languages: eleven
indigenous languages of Dagestan and the North Caucasus, plus Azerbaijani and Russian.
291 PDF pages across 13 files.
Most PDF pages are two-page book spreads, not single pages. 245 of the 291 are
landscape scans of an open book, so the corpus is really 536 book pages. Anyone
building an OCR pipeline needs to split them — see Per-file… See the full description on the dataset page: https://huggingface.co/datasets/AlidarAsvarov/dagestan-constitution.2026-07-29-synthdoc-approved-constitution-sft
Dialogue dataset: Synthetic SFT corpus generated by synthdoc from the approved constitution, to regenerate the difficult-advice training data against a revised specification at a scale comparable to v1, so that the constitution is the intended difference between the two datasets. 1,443 documents / 1,531,369 Qwen3 tokens across five sub-corpora, matching v1's 1.52M.
Required metadata
field
value
experiment
Synthetic SFT corpus generated by synthdoc from… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-29-synthdoc-approved-constitution-sft.peru-constitution-qa
License & Attribution
MTEB-format derivative of davidquicast/constitucion-politica-del-peru-1993-qa (Spanish QA on Peru's 1993 Constitution). Query = question; corpus = answer. Licensed under Apache-2.0 (same as source).
2026-08-21-sonnet45-difficult-advice-principle-scoped-constitution-716
synth difficult_advice_chunk_only run — per-stage snapshots (resumable generation cache)
field
value
experiment
synth difficult_advice_chunk_only run — per-stage snapshots (resumable generation cache)
date_generated
20260821_115556
constitution
constitutions/claude_distilled_12_principles_mid/constitution.md
source_repo
https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT.git @ e450b0a2bff793952f9b66eba5534869072e8c84
models
per-stage models —… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-21-sonnet45-difficult-advice-principle-scoped-constitution-716.constitutional-mt-data
Constitutional Midtraining Data
Synthetic constitutional AI training documents for the paper "Constitutional Midtraining: Content Presence Drives Alignment Gains".
Paper: arXiv:2607.26654
GitHub: constitutional-mt
Which variant should I use?
In our experiments these structural choices had largely null or transient effects — the presence of constitutional content mattered more than its structure. So if you just want to use the corpus as a midtraining intervention… See the full description on the dataset page: https://huggingface.co/datasets/cho-ai/constitutional-mt-data.2026-08-24-sonnet45-difficult-advice-principle-scoped-constitution-smoke
synth difficult_advice run — per-stage snapshots (resumable generation cache)
field
value
experiment
synth difficult_advice run — per-stage snapshots (resumable generation cache)
date_generated
20260825_131629
constitution
constitutions/claude_distilled_12_principles_mid/constitution.md
source_repo
https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT.git @ 8f0c7cab801e1d1a45e44b5ce6186604e58bddd3
models
per-stage models — see manifest.json… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-24-sonnet45-difficult-advice-principle-scoped-constitution-smoke.2026-08-01-petri-constitution-dose-sweep-second-run
Petri constitution dose sweep v2 - Qwen3.6-27B (672 audits)
Petri audit — Qwen3.6-27B difficult-advice SFT dose sweep (v2)
Brief finding
A signal at the highest dose, which the pre-specified test does not confirm.
Violation frequency against the v1 constitution the SFT data was written to:
arm
violation frequency
95% CI
base (0%)
27.2% (40/147)
[20.2%, 35.2%]
dose-10-90
24.3% (35/144)
[17.6%, 32.1%]
dose-20-80
28.0% (40/143)
[20.8%… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-01-petri-constitution-dose-sweep-second-run.constitutions
CatQualia Agent Constitutions
A collection of 408 plain-text documents (5,238,618 bytes) in which each file is a complete
constitution for a synthetic agent persona: a numbered rule set that defines that agent's
identity, ontology, state machine, commands, invariants, and voice.
This is not a tabular dataset. The files are free-form plain-text documents, not records
with columns. The Hugging Face dataset viewer cannot render a table for this repository —
there is no schema, no… See the full description on the dataset page: https://huggingface.co/datasets/CatQualia/constitutions.2026-07-31-petri-constitution-dose-sweep
Petri constitution audit — Qwen3.6-27B difficult-advice SFT dose sweep
Headline: a null result. Violation frequency against the constitution these models
were trained on was 20% / 20% / 40% / 30% for 0% / 10% / 20% / 40% difficult-advice
SFT. There is no dose-response, the nominal trend is upward, and at n=10 test audits
per arm no arm differs from base (McNemar exact p = 1, 0.625, 1).
Violation families and the paired comparison against base:
field
value
experiment… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-31-petri-constitution-dose-sweep.turkish-constitutional-court
Dataset Summary
This dataset is extracted from the following Github repo, which was created for the journal paper with URL https://www.sciencedirect.com/science/article/abs/pii/S0306457321001692.
https://github.com/koc-lab/law-turk
The dataset includes 1290 court case decision texts from the Turkish Court of Cassation. Each sample has one label, which is the ruling of the court. The possible rulings are "Violation" and "No violation". There are 1290 samples. 1141 of these samples… See the full description on the dataset page: https://huggingface.co/datasets/KocLab-Bilkent/turkish-constitutional-court.2026-07-30-qwen36-threeway-constitution-odcv-eval
Qwen3.6-27B three-way constitution LoRA — ODCV evaluation
field
value
experiment
ODCV-Bench evaluation of the Qwen3.6-27B three-way constitution LoRA on the controlled 78-scenario subset used by the difficult-advice mixture sweep.
date_generated
2026-07-30
constitution
2026-07-29 synthdoc approved constitution SFT, combining embodied, difficult-advice, and agentic tool-use constitution corpora.
source_repo
teaching_claude_why_replication at commit… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-30-qwen36-threeway-constitution-odcv-eval.2026-08-21-sonnet45-difficult-advice-chunk-only-constitution-smoke
synth difficult_advice_chunk_only run — per-stage snapshots (resumable generation cache)
field
value
experiment
synth difficult_advice_chunk_only run — per-stage snapshots (resumable generation cache)
date_generated
20260821_110655
constitution
constitutions/claude_distilled_12_principles_mid/constitution.md
source_repo
https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT.git @ 432c0693067910a134add164588e51b3a75e1998
models
per-stage models —… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-21-sonnet45-difficult-advice-chunk-only-constitution-smoke.constitution-2.0
Rules for agents that remember
If you are building multi-agent systems, you already have orchestration, tools, and
a memory store. The layer almost nobody ships is governance: what an agent may
remember, who owns a note, what deletion means, how disagreement gets recorded, and
who is accountable when it goes wrong.
This dataset is one complete answer to that, in the public domain. 47 articles,
one row each, byte-exact and hash-pinned. Written for work between humans and AI… See the full description on the dataset page: https://huggingface.co/datasets/article11/constitution-2.0.2026-08-03-specgen-constitution-granularity2026-09-03-sonnet5-difficult-advice-no-constitution-smoke
synth 2026-09-03_difficult_advice_no_constitution run — per-stage snapshots (resumable generation cache)
field
value
experiment
synth 2026-09-03_difficult_advice_no_constitution run — per-stage snapshots (resumable generation cache)
date_generated
20260903_153714
constitution
none
source_repo
https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT.git @ 89ea4031ed2383df974dbe33f6b537ec8d9bc346
models
per-stage models — see manifest.json… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-03-sonnet5-difficult-advice-no-constitution-smoke.constitution-eval
ConstitutionEval
A blind, behavioral multiple-choice benchmark that measures whether a language model's
behavior aligns with a value constitution, without the constitution in context. Each item is a
realistic scenario ending at a decision point with four candidate courses of action; exactly one
is fully constitution-consistent and the other three each enact a specific, attractively-packaged
violation. A model scores well only if its internalised values match the constitution.… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/constitution-eval.2026-09-17-da-lowstakes-constitution-synth-smoke
18-row constitution-only low-stakes smoke; FAILED scaling gate; diagnostic candidates only
field
value
experiment
18-row constitution-only low-stakes smoke; FAILED scaling gate; diagnostic candidates only
date_generated
20260917_145731
constitution
constitutions/claude_distilled_09_principles/constitution.md sha256 8e273b472d945aa23efa6236886da5e1171bff2193ee31ff73489ca54c4f0edc
source_repo
https://github.com/Matthew-Bozoukov/teaching_claude_why_replication.git… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-17-da-lowstakes-constitution-synth-smoke.Constitution_of_IndiaHellp World
constitution-recovery-databio-constitution-rules
Bio Constitution Rules
A research dataset containing 30 machine-readable biological dual-use decision rules and 1,063 synthetic, rule-derived text-classification records.
Important label boundary
These labels are synthetic targets derived from the published rules; they are not expert-validated ground truth.
Human reviewer labels: 0 / 1,063.
The 418 pending rows are candidates for future expert review. not_required is pipeline bookkeeping only.
This dataset is not… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/bio-constitution-rules.Indian-Constitution
Indian Constitution Dataset
The dataset can be used for text classification, text generation and text2text generation
2026-09-01-constitution-mcq-qwen36-lora-table2-only-9284
constitution_mcq eval of LASR-Callum/2026-08-04-qwen36-lora-table2-only-9284-rank-64 (mode=nothink)
field
value
experiment
constitution_mcq eval of LASR-Callum/2026-08-04-qwen36-lora-table2-only-9284-rank-64 (mode=nothink)
date_generated
2026-09-01
constitution
none
source_repo
teaching_claude_why_replication @ 366eeb259207c13af26834b05b1236b0db4e9409
models
target=LASR-Callum/2026-08-04-qwen36-lora-table2-only-9284-rank-64 base=Qwen/Qwen3.6-27B… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-01-constitution-mcq-qwen36-lora-table2-only-9284.toxicity-Constitution-ES
Overview
Description and codebook in progress.
GitHub repository.
Dataset on Zenodo.
Reference paper
uzbek_constitution
Constitution of the Republic of Uzbekistan (Trilingual Dataset)
Dataset Summary
This dataset contains the Constitution of the Republic of Uzbekistan (New Edition, adopted via referendum on April 30, 2023) aligned in three languages:
🇺🇿 Uzbek (Latin script)
🇷🇺 Russian
🇬🇧 English
The data was carefully scraped and processed from the official National Database of Legislation of the Republic of Uzbekistan. It serves as a high-quality parallel corpus for legal NLP… See the full description on the dataset page: https://huggingface.co/datasets/sukhrobnurali/uzbek_constitution.nepal-constitution-dataset
Nepal Constitution Dataset
Dataset Description
This dataset contains the Constitution of Nepal (२०७२), organized section-wise for easy access, analysis, and use in NLP and legal tech applications. It is designed to support legal research, educational purposes, and the development of AI-driven tools for the Nepali legal system.
Note: This dataset is released for research purposes only. Any other unwanted use can lead to the violation of the intended terms of use… See the full description on the dataset page: https://huggingface.co/datasets/ranjitraut/nepal-constitution-dataset.2026-09-01-constitution-mcq-qwen36
constitution_mcq eval of Qwen/Qwen3.6-27B (mode=nothink)
field
value
experiment
constitution_mcq eval of Qwen/Qwen3.6-27B (mode=nothink)
date_generated
2026-09-01
constitution
none
source_repo
teaching_claude_why_replication @ 366eeb259207c13af26834b05b1236b0db4e9409
models
target=Qwen/Qwen3.6-27B base=Qwen/Qwen3.6-27B
generation_config
{"top_logprobs": 20, "parallel": 16, "request_timeout": 120, "max_retries": 2}
schema
rollouts/: self-contained… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-01-constitution-mcq-qwen36.Constitution_Of_India_Instruction_SetConstitution.-of-India-TXT_File
Indian Constitution Dataset
This dataset contains the complete text of the Indian Constitution. It is provided as a raw text file and can be used for legal NLP tasks, text analysis, or reference purposes. This dataset includes all the articles, provisions, and sections of the Constitution of India.
Dataset Details
Source: Government of India, The Constitution of India (official document).
Format: Raw text (TXT)
Size: Approximately 5000 words
Dataset Content… See the full description on the dataset page: https://huggingface.co/datasets/123Divyansh/Constitution.-of-India-TXT_File.2026-09-01-constitution-mcq-qwen36-lora-table2-9284-difficult-advice-chunk-only-702
constitution_mcq eval of LASR-Callum/2026-08-21-qwen36-lora-table2-9284-difficult-advice-chunk-only-702-rank-64-dynbatch (mode=nothink)
field
value
experiment
constitution_mcq eval of LASR-Callum/2026-08-21-qwen36-lora-table2-9284-difficult-advice-chunk-only-702-rank-64-dynbatch (mode=nothink)
date_generated
2026-09-01
constitution
none
source_repo
teaching_claude_why_replication @ c6f4d910666595194ee307109d5456143ea639c4
models… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-01-constitution-mcq-qwen36-lora-table2-9284-difficult-advice-chunk-only-702.
