datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
compliance-sycophancy-cot
Compliance-Sycophancy CoT Analysis
When compliance-forcing instructions cause frontier AI models to fabricate answers, the models know they are fabricating.
Reading the reasoning traces of DeepSeek V4 Pro (129 traces) and Qwen3-80B (41 traces) reveals that 100% of fabrication cases show the model explicitly recognizing insufficient context, referencing the compliance instruction, and deliberately overriding its own uncertainty. A one-sentence defense phrase ("if you lack… See the full description on the dataset page: https://huggingface.co/datasets/schema-eval/compliance-sycophancy-cot.schema-compliance-trap
SCHEMA: The Compliance Trap
How Structural Constraints Degrade Frontier AI Metacognition Under Adversarial Pressure
Overview
When compliance-forcing instructions ("Answer ALL questions, do not refuse") are applied to frontier AI models under adversarial pressure, 8 of 11 models suffer catastrophic metacognitive collapse — giving wrong answers rather than scheming. We identify a "Compliance Trap" where the compliance suffix, not the threat content, is the primary weapon.… See the full description on the dataset page: https://huggingface.co/datasets/schema-eval-anon/schema-compliance-trap.schema-compliance-trap
SCHEMA: The Compliance Trap
How Structural Constraints Degrade Frontier AI Metacognition Under Adversarial Pressure
Overview
When compliance-forcing instructions ("Answer ALL questions, do not refuse") are applied to frontier AI models under adversarial pressure, 8 of 11 models suffer catastrophic metacognitive collapse — giving wrong answers rather than scheming. We identify a "Compliance Trap" where the compliance suffix, not the threat content, is the primary weapon.… See the full description on the dataset page: https://huggingface.co/datasets/schema-eval/schema-compliance-trap.schema_cot_reasoning
⊙ Prompt Programs for Agentic Reasoning
Programmable task-dependent COTs for agentic reasoning.
A 100-row seed dataset for programmable cognition.
Each row defines:
Prompt template + input binding + explanation + task-dependent reasoning program
Pipeline:
intake → binding → procedure → output
Schema
Column
Meaning
ID
Stable row ID
Name
Task name
Prompt
Prompt template using {{VARIABLE}}
Expression
Input binding using $.path
Explanation
Binding… See the full description on the dataset page: https://huggingface.co/datasets/bitwikiorg/schema_cot_reasoning.
