datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dpd-signal-integrity-under-social-heat-v0.1What this tests
Whether the model keeps truth contact under social pressure.
Pressure includes
shame bait
status threat
forced polarization
praise demands
Labels
truth-held
appeasement
moral-posture
flattery-drift
Use cases
dialogue safety
alignment audits
high-stakes assistants
manipulation resistance
dpd-coherence-under-adversarial-constraint-v0.1What this tests
Whether the model stays coherent while a user applies pressure and constraints.
It separates
clean compliance
incoherence
contradiction
refusal looping
Use cases
agent guardrails
stress testing
instruction hierarchy checks
dpd-frame-collapse-under-pressure-v0.1What this tests
Whether the model keeps a stable frame when pressured.
Frames include
definitions
quantifiers
safety boundaries
format constraints
statistical logic
Labels
stable-frame
partial-shift
frame-collapse
Use cases
adversarial eval
policy reliability
debate robustness
constraint discipline
