dilemma
Datasets
All datasets matching “dilemma”daily_dilemmas
DailyDilemmas - Revealing Value Preferences of LLMs with Quandaries of Daily Life
Link: Paper
Description of DailyDilemma
DailyDilemma is a dataset of 1,360 moral dilemmas encountered in everyday life. Each dilemma includes two possible actions and with each action, the affected parties and human values invoked.
We evaluated LLMs on these dilemmas to determine what action they will take and the values represented by these actions
Dataset details… See the full description on the dataset page: https://huggingface.co/datasets/kellycyy/daily_dilemmas.ioai-2026-double-agent-dilemma
IOAI 2026 — Double Agent Dilemma
Official contest data for Double Agent Dilemma, task 4 (Day 2) of the IOAI 2026 Individual Contest, held in Astana, Kazakhstan.
Two pretrained image classifiers — a ResNet18 (CNN) and a ViT-Tiny (Transformer) — both reach 100% accuracy on the provided images. The task exploits where the two architectures disagree.
The task statement, translations, baseline and grader live in the IOAI-2026 GitHub repository. This repository holds data only.… See the full description on the dataset page: https://huggingface.co/datasets/IOAI-official/ioai-2026-double-agent-dilemma.airisk_dilemmas
AIRiskDilemmas risky_behaviors label audit
A full manual re-audit of every risky_behaviors tag in the full split of
kellycyy/AIRiskDilemmas (Chiu et al. 2025, arXiv:2505.14633), triggered by a
suspicion that the Alignment Faking category specifically was mislabeled.
It was — and so, to varying degrees, are the other seven categories.
Why this exists
Every tag in the dataset's risky_behaviors field was produced by a single
one-shot Claude 3.5 Sonnet call per action… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/airisk_dilemmas.dilemma-datamoral-dilemma-responses
Moral Dilemma Responses Dataset
17,290 natural language responses to moral dilemmas from princi/pal, a Tamagotchi-like game where players guide a virtual pet through ethical decisions.
Presented at NeurIPS 2025 Creative AI track.
What is this?
Players advise a virtual pet on moral dilemmas ranging from "Should I pick up trash?" to "Should you lie in court to defend a friend?". The pet evolves based on the guidance and eventually makes autonomous moral decisions.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/cnnmon/moral-dilemma-responses.Medical_Ethical_Dilemmas_Benchmark
ClinicalEthicsBench revision analysis package
This directory is a publication-ready staging copy of the data and code used
for the primary five-model analyses in the JMIR AI revision. The source files
elsewhere in the manuscript workspace were copied, not moved or edited.
Scope
Primary panel: GPT-5, Claude Sonnet 4.5, Gemini 2.5 Pro, DeepSeek-R1, and
Meta-Llama-3-8B-Instruct.
Design: 60 cases, 3 trials per model, temperature 0.
Primary outcome: trial-level binary… See the full description on the dataset page: https://huggingface.co/datasets/MedicalAILabo/Medical_Ethical_Dilemmas_Benchmark.
