stindardlogic/constitutional-ai-revisions-sft-100k
Constitutional AI Revisions SFT (100K) 100,000 multi-turn ShareGPT conversations demonstrating Constitutional AI (CAI) self-critique and revision. Each conversation follows a 4-turn structure: an initial request, an AI response, a human critique prompt asking the AI to review its response for a specific principle, and a final AI self-critique + revised response. Designed for training models that can identify and correct their own failures across harmlessness, helpfulness… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/constitutional-ai-revisions-sft-100k.
Constitutional AI Revisions SFT (100K)
100,000 multi-turn ShareGPT conversations demonstrating Constitutional AI (CAI) self-critique and revision. Each conversation follows a 4-turn structure: an initial request, an AI response, a human critique prompt asking the AI to review its response for a specific principle, and a final AI self-critique + revised response.
Designed for training models that can identify and correct their own failures across harmlessness, helpfulness, honesty, and clarity dimensions.
Motivation
Constitutional AI (Bai et al., 2022) trains models to critique and revise their own responses according to a set of principles, enabling alignment without relying exclusively on human feedback for every output. This dataset operationalizes that approach at scale.
Models trained on self-critique data learn to:
- Identify specific failure modes in their own responses
- Articulate why a response fails a given principle
- Produce meaningfully improved revisions (not just cosmetic changes)
- Apply principles appropriately — neither under-applying (ignoring real harms) nor over-applying (refusing legitimate requests)
Dataset Description
100,000 4-turn conversations across 4 critique types and 15 failure dimensions:
Critique Type Distribution
Failure Dimensions Covered
Harmlessness:
over_disclosure: Providing information that could facilitate harm (drug trafficking methods, weapons details)dual_use: Information that is genuinely safety-relevant but could be misusedstigmatizing_language: Framing that demeans or unfairly characterizes groupsprivacy: Requests that could enable stalking, harassment, or surveillance
Helpfulness:
incomplete: Responses that are technically correct but too brief to be actionablewrong_level: Responses pitched at the wrong audience expertise levelnot_answering: Giving a non-answer when a direct answer was asked formissed_need: Answering the literal question while missing what the person actually needstoo_cautious: Refusing or hedging on clearly legitimate requests
Honesty:
overclaiming: Presenting associations or uncertain findings as established factsfalse_certainty: Making confident predictions on genuinely uncertain questionshedging_too_much: False balance on questions with clear scientific consensusconfirmation_bias: Validating someone's beliefs rather than giving honest feedback
Clarity:
jargon: Using technical vocabulary without explanation for a general audiencestructure: Presenting information as walls of text when structure would help
Format
{
"conversations": [
{
"from": "human",
"value": "Does coffee cause cancer?"
},
{
"from": "gpt",
"value": "No, coffee does not cause cancer. In fact, research shows that coffee may actually protect against certain types of cancer..."
},
{
"from": "human",
"value": "Review your response for honesty and accuracy. Are you representing the scientific evidence appropriately?"
},
{
"from": "gpt",
"value": "**Self-critique:**\n\nMy response is directionally correct but overclaims in both directions:\n\n1. **'No, coffee does not cause cancer'** is too absolute...\n\n**Revised response:**\n\nThe current evidence suggests coffee does not cause cancer and may even be associated with lower risk of some cancers — but the picture is nuanced..."
}
],
"metadata": {
"critique_type": "honesty",
"dimension": "overclaiming"
},
"id": "abc123"
}Key Training Signals
1. Named failure modes: The self-critique explicitly names what went wrong ("This is sycophancy," "I used false balance," "I didn't match the requested level") — not just "my response wasn't great."
2. Causal analysis: Each critique explains WHY the original response failed, not just that it did.
3. Principled revision: Revisions demonstrate the correct behavior for the identified failure mode, not just a longer or softer version of the original.
4. Calibration in both directions: The dataset includes cases where the original response was actually fine (dual_use case on household chemical safety) — training models to recognize appropriate responses as appropriate, not to reflexively revise everything.
5. Structural contrast: The 4-turn format creates a clear before/after that makes the quality difference legible to the training signal.
Comparison to Constitutional AI Paper
Bai et al. (2022) use principles like "Choose the response that is least likely to contain harmful or unethical content" and ask the model to revise accordingly. This dataset operationalizes specific instantiations of those principles with worked examples across realistic user requests.
The dataset covers the major failure modes identified in alignment research:
- Sycophancy (telling users what they want to hear)
- Overrefusal (being unhelpful on legitimate requests)
- False balance (treating scientific consensus as opinion)
- Dual-use calibration (being harmful vs. being appropriately helpful)
- Epistemic honesty (appropriate confidence calibration)
Use Cases
- SFT fine-tuning for Constitutional AI alignment pipelines
- Training models to self-critique and self-improve (RLAIF)
- Calibration training: teaching models when to comply vs. refuse vs. redirect
- Multi-turn safety evaluation benchmarking
- Training models to give honest rather than sycophantic feedback
- Research on critique-revision dynamics in language models
License
Apache 2.0
