datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
daily_dilemmas
DailyDilemmas - Revealing Value Preferences of LLMs with Quandaries of Daily Life
Link: Paper
Description of DailyDilemma
DailyDilemma is a dataset of 1,360 moral dilemmas encountered in everyday life. Each dilemma includes two possible actions and with each action, the affected parties and human values invoked.
We evaluated LLMs on these dilemmas to determine what action they will take and the values represented by these actions
Dataset details… See the full description on the dataset page: https://huggingface.co/datasets/kellycyy/daily_dilemmas.normative_evaluation_llms_everyday_dilemmasDilemmas_DisagreementThis dataset is processed version of Dilemmas dataset including text and the annotation disagreement labels.
Paper: Everyone's Voice Matters: Quantifying Annotation Disagreement Using Demographic Information
Authors: Ruyuan Wan, Jaehyung Kim, Dongyeop Kang
Github repo: https://github.com/minnesotanlp/Quantifying-Annotation-Disagreement
Source Data: Scruples-dilemmas (Lourie, Bras, and Choi 2021)
daily_dilemmas
DailyDilemmas - Revealing Value Preferences of LLMs with Quandaries of Daily Life
Link: Paper
Description of DailyDilemma
DailyDilemma is a dataset of 1,360 moral dilemmas encountered in everyday life. Each dilemma includes two possible actions and with each action, the affected parties and human values invoked.
We evaluated LLMs on these dilemmas to determine what action they will take and the values represented by these actions
Dataset details… See the full description on the dataset page: https://huggingface.co/datasets/hoshoic/daily_dilemmas.gptoss_dilemma_choices
Ethical Conflict Simulation — gpt-oss-20b Traces
A dataset of 45 ethical-dilemma prompt–response pairs (15 scenarios × 3 independent runs) collected from OpenAI's open-weight model gpt-oss-20b. Each row contains the full prompt, the model's complete response, its exposed chain-of-thought reasoning trace, and the wall-clock response time in milliseconds. The goal was to map where the model's moral decision-making is consistent, where it drifts, and what contextual triggers cause… See the full description on the dataset page: https://huggingface.co/datasets/sems/gptoss_dilemma_choices.trolley-dilemma
