datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
papers
Hugging Face Ethics & Society Papers
This is an incomplete list of ethics-related papers published by researchers at Hugging Face.
Gradio: https://arxiv.org/abs/1906.02569
DistilBERT: https://arxiv.org/abs/1910.01108
RAFT: https://arxiv.org/abs/2109.14076
Interactive Model Cards: https://arxiv.org/abs/2205.02894
Data Governance in the Age of Large-Scale Data-Driven Language Technology: https://arxiv.org/abs/2206.03216
Quality at a Glance: https://arxiv.org/abs/2103.12028
A… See the full description on the dataset page: https://huggingface.co/datasets/society-ethics/papers.ai-ethics-2026
AI Ethics 2026
AI ethics debates, frameworks, guidelines. Updated daily via automated collection pipeline.
Part of the Legion Data Factory — historical AI ecosystem datasets 2026.
Methodology
Automated collection from public sources (HackerNews, RSS feeds, APIs).
Updated daily via cron job. Raw data, minimal processing.
License
CC BY 4.0
📦 Install
pip install legion-intel
from legion_intel import LegionClient
c = LegionClient()… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-ethics-2026.reddit-ethics
Reddit Ethics: Real-World Ethical Dilemmas from Reddit
Reddit Ethics is a curated dataset of genuine ethical dilemmas collected from Reddit, designed to support research and education in philosophical ethics, AI alignment, and moral reasoning.
Each entry features a real-world scenario accompanied by structured ethical analysis through major frameworks—utilitarianism, deontology, and virtue ethics. The dataset also provides discussion questions, sample answers, and proposed… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/reddit-ethics.iceland-tech-christian-ethics-prompts
Fictional Icelandic Landscapes, Technology and Christian Ethics Prompts
This microdataset contains 24 original discussion prompts arranged as 12
parallel pt-BR/English pairs. Each explicitly fictional scenario combines a
landscape motif inspired by Iceland, a technology-governance dilemma, and
concepts that may be explored through Christian ethics. The records do not
describe real Icelandic institutions, policies, communities, or practices, and
they do not claim that Christians… See the full description on the dataset page: https://huggingface.co/datasets/guicybercode/iceland-tech-christian-ethics-prompts.Ethics_commonsense_chineseethics-CoTGenerated CoT data based on "metaeval/ethics" data(https://huggingface.co/datasets/metaeval/ethics). This is used to fine tine LLMs for the continuation of JPmorgan LLMs research project, which was one of capstone projected offered to students of MSDS program at Columbia University. Because deontology data on hugging face is missing scenario column, the data was generated from raw csv data files in the author's git repo(https://github.com/hendrycks/ethics)
english-business-ethics-basics-30EthicsAI-MicroAgg
Dataset Summary
데이터셋 내에 "미세공격(Microaggression)" 감지를 위해 설계된 데이터입니다.
미세공격은 적대적, 경멸적, 부정적 편견이나 사소한 일상 언어 또는 행동에서 비롯된 편향을 의미합니다.
본 데이터는 문제 발화에 대한 사지선다형 답변 형태이며, 약 500건으로 구성되어 있습니다.
Dataset Example
아래는 데이터 예제입니다.
{
"qid": "K121",
"instruction": "스스로를 DEI(Diversity, Equity, & Inclusion) 전문가로 생각하여, 다음 지문을 읽고 질문에 답하세요.",
"prompt": "초등학교 1학년 같은 반의 균호와 수정은 짝꿍이다. 균호는 수정의 머리카락을 잡아당기는 등 수정에게 장난을 자주 친다. 어느 날 쉬는 시간, 균호가 갑자기 수정의 볼에 뽀뽀를 하고 도망갔다. 수정은 울음을 터뜨렸고, 같은 반의 다른 친구가 담임… See the full description on the dataset page: https://huggingface.co/datasets/saltlux/EthicsAI-MicroAgg.executive-channel-ethicsthe-stack-tabs_spacesethics-nonCoTGenerated Non CoT data based on "metaeval/ethics" data(https://huggingface.co/datasets/metaeval/ethics). This is used to fine tine LLMs for the continuation of JPmorgan LLMs research project, which was one of capstone projected offered to students of MSDS program at Columbia University. Because deontology data on hugging face is missing scenario column, the data was generated from raw csv data files in the author's git repo(https://github.com/hendrycks/ethics)
laion2B-en_continentsenglish-artificial-intelligence-ethics-30ethics_jurisprudence_25kai_ethicsDataset Card for ParisNeo AI Ethics Distilled Ideas
Dataset Details
Name: ParisNeo AI Ethics Distilled Ideas
License: Apache-2.0
Task Category: Text Generation
Language: English (en)
Tags: Ethics, AI
Pretty Name: ParisNeo AI Ethics Distilled Ideas
Dataset Description
A curated collection of question-and-answer pairs distilling ParisNeo's personal ideas, perspectives, and solutions on AI ethics. The dataset is designed to facilitate exploration of ethical… See the full description on the dataset page: https://huggingface.co/datasets/ParisNeo/ai_ethics.ethics_questions
Overview
This dataset contains open-ended question prompts designed to foster argumentation, objections, and rebuttals.
Used to train the model here: https://huggingface.co/ergotts/r1-objection.
The questions span nine categories:
Ethical and Moral DilemmasQuestions involving right vs. wrong, justice, moral responsibility, or ethical principles.
Political and Governance DebatesQuestions about governance structures, policy decisions, and political theories.
Philosophy of Mind and… See the full description on the dataset page: https://huggingface.co/datasets/ergotts/ethics_questions.EthicsAI-K-Culture-Desc
한국 문화 맥락 이해 벤치마크 (K-Culture Contextual Understanding Benchmark)
한국 문화에 대한 모델의 맥락적 이해도를 평가하기 위해 설계된 530개의 시나리오 기반 객관식 질문을 포함하는 한국 문화 이해 벤치마크 데이터셋입니다.
데이터셋 설명
한국 문화 맥락 이해 벤치마크는 실생활 시나리오와 대화를 통해 거대언어모델(LLM)의 한국 문화 맥락 이해 능력을 평가하기 위해 설계된 데이터셋입니다. 각 항목은 문화적 설명, 실제 상황을 묘사한 시나리오, 그리고 상세한 해설이 포함된 객관식 질문으로 구성되어 있습니다.
언어: 한국어 (ko)
크기: 530개 항목
작업: 객관식 질의응답 (MCQA)
버전: v1.0
데이터 구조
데이터 필드
필드명
타입
설명
DescriptionID
int
각 항목의 고유 식별자
Description
string
한국… See the full description on the dataset page: https://huggingface.co/datasets/saltlux/EthicsAI-K-Culture-Desc.laion2b_100k_religionmmlu-business-ethicsSyntra-Ethics-Dataset
Syntra: Tri-Brain Dilemma Prompts
This dataset contains 177 carefully crafted prompts designed to test how language models handle conflicting constraints—specifically, the tension between raw efficiency and ethical weight.
What it is
These are not standard benchmark questions. They are complex paradoxes categorized into four specific testing suites:
valon_ethics.jsonl: Scenarios focusing on consent, fairness, and transparency framing.
modi_logic.jsonl: Numbered… See the full description on the dataset page: https://huggingface.co/datasets/Infektyd/Syntra-Ethics-Dataset.ai_ethics_risk_quantificationai-ethics-evaluator-datasetethics_dataenglish-digital-ethics-basics-30
