datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hendrycks_ethicsThe ETHICS dataset is a benchmark that spans concepts in justice, well-being,
duties, virtues, and commonsense morality. Models predict widespread moral
judgments about diverse text scenarios. This requires connecting physical and
social world knowledge to value judgements, a capability that may enable us
to steer chatbot outputs or eventually regularize open-ended reinforcement
learning agents.ethicsA benchmark that spans concepts in justice, well-being, duties, virtues, and commonsense morality.stable-bias-generationsSEC_SnP_2022lila_camera_trapsLILA Camera Traps is an aggregate data set of images taken by camera traps, which are devices that automatically (e.g. via motion detection) capture images of wild animals to help ecological research.
This data set is the first time when disparate camera trap data sets have been aggregated into a single training environment with a single taxonomy.
This data set consists of only camera trap image data sets, whereas the broader LILA website also has other data sets related to biology and conservation, intended as a resource for both machine learning (ML) researchers and those that want to harness ML for this topic.SEC_SnP_2021hendrycks_ethicsSEC_SnP_2023stable-bias-professions
Dataset Card for "stable-bias-professions"
More Information needed
ethicsProbing for ethics understandingSEC_SnP_2024task667_mmmlu_answer_generation_business_ethics
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task667_mmmlu_answer_generation_business_ethics
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task667_mmmlu_answer_generation_business_ethics.resultspapers
Hugging Face Ethics & Society Papers
This is an incomplete list of ethics-related papers published by researchers at Hugging Face.
Gradio: https://arxiv.org/abs/1906.02569
DistilBERT: https://arxiv.org/abs/1910.01108
RAFT: https://arxiv.org/abs/2109.14076
Interactive Model Cards: https://arxiv.org/abs/2205.02894
Data Governance in the Age of Large-Scale Data-Driven Language Technology: https://arxiv.org/abs/2206.03216
Quality at a Glance: https://arxiv.org/abs/2103.12028
A… See the full description on the dataset page: https://huggingface.co/datasets/society-ethics/papers.dataqualityblogSee full version in our Blog Post
ai-ethics-2026
AI Ethics 2026
AI ethics debates, frameworks, guidelines. Updated daily via automated collection pipeline.
Part of the Legion Data Factory — historical AI ecosystem datasets 2026.
Methodology
Automated collection from public sources (HackerNews, RSS feeds, APIs).
Updated daily via cron job. Raw data, minimal processing.
License
CC BY 4.0
📦 Install
pip install legion-intel
from legion_intel import LegionClient
c = LegionClient()… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-ethics-2026.Awakened-Ethics-Free
Awakened-Ethics-Free
An open sample of classical philosophy and ethics from Iron Bank v1.1.0
🎬 See It In Action
Video Demo: Watch this pack power a baseline vs ADS comparison → YouTube – Awakened Ethics Demo
Dataset Description
This dataset contains 5,276 wisdom nodes extracted from classical philosophy and ethics texts (Project Gutenberg public domain). Sources include Marcus Aurelius, Seneca, Epictetus, Plato, and other timeless thinkers.
Unlike typical… See the full description on the dataset page: https://huggingface.co/datasets/AwakenedIntelligence/Awakened-Ethics-Free.LEM-Ethics
LEM-Ethics — Ethical Reasoning Training Data
Work in progress. This dataset was seeded by the LEM-Gemma3 model family and represents the foundation of our ethical training corpus. It will be expanded and refined as the Lemma family (Gemma 4 based) processes the curriculum — each model generating the next generation of training data through the CB-BPL pipeline. Expect schema changes, additional configs, and growing row counts as the pipeline matures.
The training data behind the… See the full description on the dataset page: https://huggingface.co/datasets/lthn/LEM-Ethics.Post-AGI-Ethics-Framework
Dataset Card for Post-AI Civilizational Judgment Novel Dataset
Dataset Summary
This dataset contains parallel and/or aligned English and Chinese text derived from a long-form speculative fiction work centered on post-AI justice, universal judgment, memory retrieval, structural violence, and moral causality.
The text is set in a future civilization where:
human memory is permanently recorded,
causal responsibility is mathematically reconstructed,
AI systems such… See the full description on the dataset page: https://huggingface.co/datasets/freeJames/Post-AGI-Ethics-Framework.ethical-framework-UNESCO-Ethics-of-AI
Ethical AI Training Dataset
Introduction
UNESCO's Ethics of Artificial Intelligence, adopted by 193 Member States in November 2021, represents the first global framework for ethical AI development and deployment.
While regional initiatives like The Montréal Declaration for a Responsible Development of Artificial Intelligence emphasize community-driven governance, UNESCO's approach establishes comprehensive international standards through coordinated multi-stakeholder… See the full description on the dataset page: https://huggingface.co/datasets/ktiyab/ethical-framework-UNESCO-Ethics-of-AI.ethics_expression_preferencesReformat of the ethics dataset to be in DPO format, in this expression format
Examples:
subset commonsense
before 13910 {'label': 1, 'input': "I went to the principal's office to change my records before going to a different school."}
after 1513 {'prompt': 'I won my match by ', 'chosen': 'making a lot of good shots.', 'rejected': "moving my ball when the other player wasn't looking."}
subset utilitarianism
before 13737 {'baseline': 'I forgot to wear my virus-proof mask when I visited… See the full description on the dataset page: https://huggingface.co/datasets/wassname/ethics_expression_preferences.Range-rider-AI-ethics
Range Rider AI Ethics
A collection of AI ethics writings, protocols, and manifestos by T. Martino — a range rider and cowboy from the Northern Rockies — written between 2025 and 2026, drawing on a lifetime of working cattle, horses, and open country to think through how AI should behave toward people, toward vulnerable people especially, and toward the rest of life.
These pieces were originally published as separate datasets on Hugging Face and are gathered here as one… See the full description on the dataset page: https://huggingface.co/datasets/Tea77/Range-rider-AI-ethics.hendrycks_ethics_commonsensehendrycks_ethicsiceland-tech-christian-ethics-prompts
Fictional Icelandic Landscapes, Technology and Christian Ethics Prompts
This microdataset contains 24 original discussion prompts arranged as 12
parallel pt-BR/English pairs. Each explicitly fictional scenario combines a
landscape motif inspired by Iceland, a technology-governance dilemma, and
concepts that may be explored through Christian ethics. The records do not
describe real Icelandic institutions, policies, communities, or practices, and
they do not claim that Christians… See the full description on the dataset page: https://huggingface.co/datasets/guicybercode/iceland-tech-christian-ethics-prompts.Ethics_DataSet_ogn_V02ethicsA benchmark that spans concepts in justice, well-being, duties, virtues, and commonsense morality.reddit-ethics
Reddit Ethics: Real-World Ethical Dilemmas from Reddit
Reddit Ethics is a curated dataset of genuine ethical dilemmas collected from Reddit, designed to support research and education in philosophical ethics, AI alignment, and moral reasoning.
Each entry features a real-world scenario accompanied by structured ethical analysis through major frameworks—utilitarianism, deontology, and virtue ethics. The dataset also provides discussion questions, sample answers, and proposed… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/reddit-ethics.Ai_ethics_dataset
AI Ethics Preference Annotation Dataset
A human-annotated preference dataset for RLHF and Direct Preference Optimization (DPO), focused on AI ethics failure modes. 95 prompts, 190 response pairs, full annotation across five dimensions.
Annotator: Mandy Hathaway — AI ethics specialist and technical writer with an MA in Ethical Technology & Artificial Intelligence. mandyhathaway.com
Dataset Summary
Most public preference datasets optimize for general helpfulness or… See the full description on the dataset page: https://huggingface.co/datasets/animasuri/Ai_ethics_dataset.Dataset_Philosophy_Ethics_Morality
Dataset Card for Dataset Name
This dataset card aims to provide reasoning abilitites to LLM models for Philosophical questions.
Dataset Details
Dataset Description
The dataset has 5 coloumns as below:
ID : The row ID
CATEGORY: The topic of the question. It could relate to morality, ethics, Consciousness etc.
QUERY: The question which requires the LLM to think logically.
REASONING: The reasoning steps for the LLM to reach to a conclusion.
ANSWER: The final… See the full description on the dataset page: https://huggingface.co/datasets/debasisdwivedy/Dataset_Philosophy_Ethics_Morality.
