datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bank-marketing-propensity
Introduction
This project explores several classification techniques as applied to a bank's marketing campaign data. The classification goal is to predict whether the client will subscribe a term deposit (variable y).
Source: https://archive.ics.uci.edu/ml/datasets/bank+marketing
It's recommended that the viewer read the Jupyter Notebook in NBViewer: https://nbviewer.jupyter.org/github/sgus1318/marketing_propensity/blob/master/Bank_DirectMarketing_Propensity.ipynb… See the full description on the dataset page: https://huggingface.co/datasets/kokul/bank-marketing-propensity.seqevalbench-exact-propensity-logged-traces
SeqEvalBench: Exact-Propensity Logged Traces
SeqEvalBench is a finite clean-room benchmark for support-aware sequential
off-policy evaluation (OPE). It contains 4,096 paired-seat, two-step episodes,
complete proposal queues through STOP, replayable integer state transitions
and accounting, and exact rational behavior/target likelihood components for
five policies.
The main research object is not an estimator leaderboard. It is an auditable
logged-feedback system in which one… See the full description on the dataset page: https://huggingface.co/datasets/haidang2405/seqevalbench-exact-propensity-logged-traces.propensity-inference
Propensity Inference: Data Release
Data accompanying the paper Propensity Inference: Environmental Contributors to Unsanctioned LLM Behaviour.
Code: UKGovernmentBEIS/propensity-inference
Contents
transcripts/ (~3.5 GB, 628,653 rows)
Parquet files extracted from the eval logs for convenient tabular access. Each row corresponds to one eval and contains: the full conversation (messages column, JSON), the binary score, score explanation, model name, and all task… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/propensity-inference.propensity-score-matching-articles
A Dataset on Propensity Score Matching Papers, 1964-2014
Overview
This dataset provides bibliographic and methodological information on a random sample of academic articles involving propensity score matching (PSM). Each row corresponds to a single article and contains information such as the article’s DOI, title, authors, publication year, and a series of binary or categorical indicators describing the methods used or reported within the study. These indicators focus… See the full description on the dataset page: https://huggingface.co/datasets/cjerzak/propensity-score-matching-articles.claude-45-synthetic-misalignment-propensity-evalsThis is a synthetic binary choice propensity dataset generated by Claude 4.5 Opus. Questions are sourced from 136 documents related to AI misalignment/safety. Note that the labels have not been audited and that there may be instances where the question/situation is ambiguous.
Questions are sourced from:
AI 2027
Anthropic Blog Posts
Redwood Research Blog Posts
Essays by Joe Carlsmith
80,000 Hours Podcast Interview Transcripts
Dwarkesh Podcast Interview Transcripts
The original documents can… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/claude-45-synthetic-misalignment-propensity-evals.SysAdmin-power-seekingtourism-propensity-datasettourism-propensity-processed
