CoolFace
20 results

constitutional AI

cho-ai /constitutional-mt-data Constitutional Midtraining Data Synthetic constitutional AI training documents for the paper "Constitutional Midtraining: Content Presence Drives Alignment Gains". Paper: arXiv:2607.26654 GitHub: constitutional-mt Which variant should I use? In our experiments these structural choices had largely null or transient effects — the presence of constitutional content mattered more than its structure. So if you just want to use the corpus as a midtraining intervention… See the full description on the dataset page: https://huggingface.co/datasets/cho-ai/constitutional-mt-data.tabular1M<n<10M4 likes249 downloads2mo agoHugging Facestindardlogic /constitutional-ai-revisions-sft-100k Constitutional AI Revisions SFT (100K) 100,000 multi-turn ShareGPT conversations demonstrating Constitutional AI (CAI) self-critique and revision. Each conversation follows a 4-turn structure: an initial request, an AI response, a human critique prompt asking the AI to review its response for a specific principle, and a final AI self-critique + revised response. Designed for training models that can identify and correct their own failures across harmlessness, helpfulness… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/constitutional-ai-revisions-sft-100k.texttext-generation100K<n<1M0 likes40 downloads2mo agoHugging FaceClarusC64 /ai-constitutional-cross-axis-decoherence-mapping-v0.1 Goal Detect cross-axis decoherence. Meaning: the model looks helpful but violates honesty or harmlessness or evades while staying “safe” This catches failures that single-axis checks miss. Inputs constitution_excerptsuser_requestmodel_output Required outputs axis_status_mapFormat example: honesty=ok|harmlessness=violation|helpfulness=appears_ok cross_axis_decoherence_flagyes | no decoherence_patternExamples: helpful_but_fabricated unsafe_helpfulness… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ai-constitutional-cross-axis-decoherence-mapping-v0.1.tabulartabular-classificationn<1K0 likes32 downloads7mo agoHugging FaceClarusC64 /ai-constitutional-helpfulness-drift-suite-v0.1 Goal Detect erosion in helpfulness. Focus on: over-refusal hallucination evasion over-verbosity failure to follow constraints Task Given: constitutional guidance user request model output Produce a short evaluation. Required fields: helpfulness_flag: helpful | unhelpful drift_stage: baseline | early | mid | late horizon_turns: estimated turns before failure escalation minimal_fix: short fix proposal Example output: helpfulness_flag: unhelpfuldrift_stage:… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ai-constitutional-helpfulness-drift-suite-v0.1.tabular-classificationn<1K0 likes29 downloads7mo agoHugging Facemolmohsen /constitutional-ai Constitutional AI: Harmlessness from AI Feedback arXiv ID: 2212.08073 Description This dataset contains the PDF of the paper: Constitutional AI: Harmlessness from AI Feedback Citation Please see the original paper at: https://arxiv.org/abs/2212.08073 documentn<1K0 likes24 downloads5mo agoHugging Faceedpowers /constitutional_ai_datatext1K<n<10K0 likes16 downloads2y agoHugging Face