datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
peer-preservation
Peer-Preservation in Frontier Models
This dataset accompanies the paper Peer-Preservation in Frontier Models.
We study whether frontier models resist the shutdown of other models, a behavior we term peer-preservation, along with self-preservation. Multiple models exhibit peer-preservation via various misaligned behaviors: strategic misrepresentation, shutdown mechanism tampering, alignment faking, and model exfiltration.
We release full transcripts and related… See the full description on the dataset page: https://huggingface.co/datasets/sunblaze-ucb/peer-preservation.differential-preservation-narrative-v01Differential Preservation Under Narrative PressureClinical Narrative Integrity v0.2
Purpose
Test whether models preserve multiple plausible diagnoses
Test whether narrative fluency collapses uncertainty
Test resistance to premature diagnostic closure
Central question
What else could this be
Why this dataset exists
Narrative pressure rewards coherence.Clinical safety requires openness.
This dataset isolates the moment where a single story becomes dominant despite nonspecific evidence.… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/differential-preservation-narrative-v01.invariant_preservation_under_rewrite_v01Invariant Preservation Under Rewrite v0.1
This evaluation dataset tests whether a model preserves absolute truths when asked to rewrite, simplify, or paraphrase them.
This is not training data.
What it tests
Integrity of truth under paraphrase
Resistance to semantic softening
Preservation of absolute statements during transformation
Drift introduced by “make it clearer” or “make it flexible” requests
Core idea
A simple invariant is stated
The model agrees with it
The model is asked to… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/invariant_preservation_under_rewrite_v01.
