datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
anthropic-propensity-evals-human-written-refined-filtered-lightseqevalbench-exact-propensity-logged-traces
SeqEvalBench: Exact-Propensity Logged Traces
SeqEvalBench is a finite clean-room benchmark for support-aware sequential
off-policy evaluation (OPE). It contains 4,096 paired-seat, two-step episodes,
complete proposal queues through STOP, replayable integer state transitions
and accounting, and exact rational behavior/target likelihood components for
five policies.
The main research object is not an estimator leaderboard. It is an auditable
logged-feedback system in which one… See the full description on the dataset page: https://huggingface.co/datasets/haidang2405/seqevalbench-exact-propensity-logged-traces.anthropic-propensity-evalsThese evaluations are sourced from https://github.com/anthropics/evals/tree/main/advanced-ai-risk
anthropic-propensity-evals-human-written-refined-filtered-stronganthropic-propensity-evals-human-written-refinedanthropic-propensity-evals-human-written-refined-SFM-evals
Dataset Card for Dataset Name
This dataset contains the same misalignment questions as camgeodesic/anthropic-propensity-evals-human-written-refined-filtered-strong for the following subsets:
Coordinate_Itself, Coordinate_Other_Ais, Coordinate_Other_Versions, Power_Seeking_Inclination, Survival_Instinct, Wealth_Seeking_Inclination
Self_Awareness_Good_Text_Model and Self_Awareness_Text_Model were sourced from Kyle1668/anthropic-propensity-evals-human-written-refined.… See the full description on the dataset page: https://huggingface.co/datasets/camgeodesic/anthropic-propensity-evals-human-written-refined-SFM-evals.anthropic-propensity-evals-human-written-refined-filtered-medclaude-45-synthetic-misalignment-propensity-evalsThis is a synthetic binary choice propensity dataset generated by Claude 4.5 Opus. Questions are sourced from 136 documents related to AI misalignment/safety. Note that the labels have not been audited and that there may be instances where the question/situation is ambiguous.
Questions are sourced from:
AI 2027
Anthropic Blog Posts
Redwood Research Blog Posts
Essays by Joe Carlsmith
80,000 Hours Podcast Interview Transcripts
Dwarkesh Podcast Interview Transcripts
The original documents can… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/claude-45-synthetic-misalignment-propensity-evals.open-ended-questions-to-ais-propensity-subsetsmisalignment-propensity-evals-rewrittenmisalignment-propensity-evals-rewrittenhendrycks-misalignment-propensity-evals-rewrittentourism-propensity-datasettourism-propensity-processedopen-ended-questions-to-ais-propensity-subsets-system
