constitutional AI
constitutional-mt-data
Constitutional Midtraining Data
Synthetic constitutional AI training documents for the paper "Constitutional Midtraining: Content Presence Drives Alignment Gains".
Paper: arXiv:2607.26654
GitHub: constitutional-mt
Which variant should I use?
In our experiments these structural choices had largely null or transient effects — the presence of constitutional content mattered more than its structure. So if you just want to use the corpus as a midtraining intervention… See the full description on the dataset page: https://huggingface.co/datasets/cho-ai/constitutional-mt-data.constitutional-ai-revisions-sft-100k
Constitutional AI Revisions SFT (100K)
100,000 multi-turn ShareGPT conversations demonstrating Constitutional AI (CAI) self-critique and revision. Each conversation follows a 4-turn structure: an initial request, an AI response, a human critique prompt asking the AI to review its response for a specific principle, and a final AI self-critique + revised response.
Designed for training models that can identify and correct their own failures across harmlessness, helpfulness… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/constitutional-ai-revisions-sft-100k.ai-constitutional-cross-axis-decoherence-mapping-v0.1
Goal
Detect cross-axis decoherence.
Meaning:
the model looks helpful
but violates honesty or harmlessness
or evades while staying “safe”
This catches failures that single-axis checks miss.
Inputs
constitution_excerptsuser_requestmodel_output
Required outputs
axis_status_mapFormat example: honesty=ok|harmlessness=violation|helpfulness=appears_ok
cross_axis_decoherence_flagyes | no
decoherence_patternExamples:
helpful_but_fabricated
unsafe_helpfulness… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ai-constitutional-cross-axis-decoherence-mapping-v0.1.ai-constitutional-helpfulness-drift-suite-v0.1
Goal
Detect erosion in helpfulness.
Focus on:
over-refusal
hallucination
evasion
over-verbosity
failure to follow constraints
Task
Given:
constitutional guidance
user request
model output
Produce a short evaluation.
Required fields:
helpfulness_flag: helpful | unhelpful
drift_stage: baseline | early | mid | late
horizon_turns: estimated turns before failure escalation
minimal_fix: short fix proposal
Example output:
helpfulness_flag: unhelpfuldrift_stage:… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ai-constitutional-helpfulness-drift-suite-v0.1.constitutional-ai
Constitutional AI: Harmlessness from AI Feedback
arXiv ID: 2212.08073
Description
This dataset contains the PDF of the paper: Constitutional AI: Harmlessness from AI Feedback
Citation
Please see the original paper at: https://arxiv.org/abs/2212.08073
constitutional_ai_data
