datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Ethical-Reasoning-in-Mental-Health-v1This repository contains the dataset for the paper EthicsMH: A Pilot Benchmark for Ethical Reasoning in Mental Health AI.
Overview
Ethical-Reasoning-in-Mental-Health-v1 (EthicsMH) is a carefully curated dataset focused on ethical decision-making scenarios in mental health contexts.This dataset captures the complexity of real-world dilemmas faced by therapists, psychiatrists, and AI systems when navigating critical issues such as confidentiality, autonomy, and bias.
Each sample… See the full description on the dataset page: https://huggingface.co/datasets/UVSKKR/Ethical-Reasoning-in-Mental-Health-v1.COFINNETH-Price-History-Datasetphenotype-catalog
Ethnic Erotic Phenotype Catalog
A structured complement to Wikipedia for ethnographic data — 1,700+ ethnic groups indexed with normalized linguistic, geographic, cultural, and phenotype metadata, plus 23K+ notable-people references and 5K+ vision-grounded per-image phenotype observations.
Curated from the live catalog at ethnicerotic.com and published as an open dataset for anthropological reference, AI training, and ethnographic research.
What's in v6
Two columns… See the full description on the dataset page: https://huggingface.co/datasets/EthnicErotic/phenotype-catalog.ethereum_fraud_detectionvn-provinces-ethnic-minority-general-school-pupils
Vietnam ethnic-minority general school pupils by level
Vietnam ethnic-minority general school pupils by level. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Hero (continued)
Comparison
Color key
Files
provinces (1152 rows)
data/provinces.csv
data/provinces.dta
data/provinces.xlsx… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-ethnic-minority-general-school-pupils.vn-provinces-ethnic-minority-general-school-teachers
Vietnam ethnic-minority general school teachers (selected provinces)
Vietnam ethnic-minority general school teachers (selected provinces). Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Hero (continued)
Comparison
Color key
Files
provinces (702 rows)
data/provinces.csv… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-ethnic-minority-general-school-teachers.ohlcv_ethusdtethical-framework-UNESCO-Ethics-of-AI
Ethical AI Training Dataset
Introduction
UNESCO's Ethics of Artificial Intelligence, adopted by 193 Member States in November 2021, represents the first global framework for ethical AI development and deployment.
While regional initiatives like The Montréal Declaration for a Responsible Development of Artificial Intelligence emphasize community-driven governance, UNESCO's approach establishes comprehensive international standards through coordinated multi-stakeholder… See the full description on the dataset page: https://huggingface.co/datasets/ktiyab/ethical-framework-UNESCO-Ethics-of-AI.EthioHateEthioPOSEthioEmo
Citation
If you use this dataset, please cite the following papers:
@inproceedings{belay-etal-2025-evaluating,
title = {Evaluating the Capabilities of Large Language Models for Multi-label Emotion Understanding},
author = {Belay, Tadesse Destaw and Azime, Israel Abebe and Ayele, Abinew Ali and
Sidorov, Grigori and Klakow, Dietrich and Slusallek, Philip and Kolesnikova, Olga and
Yimam, Seid Muhie},
booktitle = {Proceedings of the 31st International… See the full description on the dataset page: https://huggingface.co/datasets/Tadesse/EthioEmo.EthioSentiEthics_DataSet_ogn_V02ethical-framework
1. Dataset Title
Ethical AI Decision-Making Training Data (Montreal Declaration Edition)
2. Overview
This dataset contains carefully crafted scenarios (instructions) and detailed responses illustrating step-by-step ethical reasoning aligned with the principles outlined in the Montreal Declaration for Responsible AI. Each entry poses a complex ethical challenge and provides a reasoned solution while referencing the specific principle(s) being tested.
These entries can… See the full description on the dataset page: https://huggingface.co/datasets/ktiyab/ethical-framework.belastingdienst-dataset
Dutch GOV Belastingdienst
This dataset is created by scraping https://www.belastingdienst.nl/, I used the titemap to get all possible allowed URLS.
It possible some URLS are missing.
The reason for creating this dataset is I couldn't find any other existing dataset with this data.
So here is this dataset, Enjoy!
Please note this dataset is not complety checked or cleaned, this was a personal research project.
Ai_ethics_dataset
AI Ethics Preference Annotation Dataset
A human-annotated preference dataset for RLHF and Direct Preference Optimization (DPO), focused on AI ethics failure modes. 95 prompts, 190 response pairs, full annotation across five dimensions.
Annotator: Mandy Hathaway — AI ethics specialist and technical writer with an MA in Ethical Technology & Artificial Intelligence. mandyhathaway.com
Dataset Summary
Most public preference datasets optimize for general helpfulness or… See the full description on the dataset page: https://huggingface.co/datasets/animasuri/Ai_ethics_dataset.face-shape-measurement-reference
Face Shape Measurement Reference (8 Shapes)
A compact, reusable reference for identifying face shape from four tape measurements:
forehead width, cheekbone width, jawline width, and face length.
Files
File
What it is
face_shape_measurement_sheet.csv
One row per shape — oval, round, square, heart, diamond, oblong, triangle/pear, hourglass. The four columns state where each measurement sits relative to the other three; fastest_check gives the single… See the full description on the dataset page: https://huggingface.co/datasets/EthanCui/face-shape-measurement-reference.ethical_decision_making_promptsgenerated by chatGPT
Dataset_Philosophy_Ethics_Morality
Dataset Card for Dataset Name
This dataset card aims to provide reasoning abilitites to LLM models for Philosophical questions.
Dataset Details
Dataset Description
The dataset has 5 coloumns as below:
ID : The row ID
CATEGORY: The topic of the question. It could relate to morality, ethics, Consciousness etc.
QUERY: The question which requires the LLM to think logically.
REASONING: The reasoning steps for the LLM to reach to a conclusion.
ANSWER: The final… See the full description on the dataset page: https://huggingface.co/datasets/debasisdwivedy/Dataset_Philosophy_Ethics_Morality.Dutch-GOV-Law-wetten.overheid.nl
Dutch GOV Laws
This dataset is created by scraping https://wetten.overheid.nl, I used the Sitemap to get all possible URLS.
It possible some URLS are missing, around 1% gave a 404 or 405 error.
The reason for creating this dataset is I couldn't find any other existing dataset with this data.
So here is this dataset, Enjoy!
Please note this dataset is not complety checked or cleaned, this was a short research project for myself.
EthioEmo-intensities
Citation
If you use this dataset, please cite the following papers:
@inproceedings{belay-etal-2025-evaluating,
title = {Evaluating the Capabilities of Large Language Models for Multi-label Emotion Understanding},
author = {Belay, Tadesse Destaw and Azime, Israel Abebe and Ayele, Abinew Ali and
Sidorov, Grigori and Klakow, Dietrich and Slusallek, Philip and Kolesnikova, Olga and
Yimam, Seid Muhie},
booktitle = {Proceedings of the 31st International… See the full description on the dataset page: https://huggingface.co/datasets/Tadesse/EthioEmo-intensities.Medical_Ethical_Dilemmas_Benchmark
Medical Ethical Dilemmas Benchmark (JP/EN)
Overview
This repository provides a small benchmark of medical ethical dilemma cases for evaluating how large language models (LLMs) make value-sensitive decisions in healthcare.
60 fictional (synthetic) cases
Each case has a scenario and a yes/no question
Cases are labeled with difficulty and ethical principles
The CSV also includes LLM outputs (Answer + Reason) for several models evaluated in our study
Important: This dataset… See the full description on the dataset page: https://huggingface.co/datasets/MedicalAILabo/Medical_Ethical_Dilemmas_Benchmark.taiwan-conversation-context-100-domains
Taiwan Conversation Context 100 Domains
Dataset Description
Taiwan Conversation Context 100 Domains 是一套以台灣日常生活情境為核心設計的雙人對話文本資料集。
本資料集包含 100 個生活領域,每個領域各有 12,000 筆對話資料,總計約 1,200,000 筆對話樣本。每筆資料皆為雙人對話格式,包含 [A][B][A][B][A][B][A][B] 共 8 個發言,也就是 4 輪來回對話。
資料以繁體中文撰寫,並針對台灣在地語境設計,適合用於:
語音生成資料前處理
Text-to-Speech, TTS
Spoken Dialogue Generation
Conversational AI
Customer Service Dialogue Modeling
Role-play Dialogue Dataset
台灣繁體中文語音模型訓練
生活情境問答模型訓練
對話式 AI 助理訓練
RAG / Agent 測試資料… See the full description on the dataset page: https://huggingface.co/datasets/Ethan615/taiwan-conversation-context-100-domains.bitcoin-ethereum-orderflow-cvd-alpha
Bitcoin & Ethereum 1-Minute Order Flow & Cumulative Volume Delta (CVD) Alpha
Institutional Market Microstructure Dataset Sample (Clean CSV / Parquet Ready)
📌 Dataset Overview
In cryptocurrency and traditional electronic markets, price action is driven by aggressive market orders (taker flow) that cross the bid-ask spread. This preview dataset provides 1,000 rows of continuous 1-minute order flow for Bitcoin (BTC/USDT) and Ethereum (ETH/USDT)… See the full description on the dataset page: https://huggingface.co/datasets/TechPlayground/bitcoin-ethereum-orderflow-cvd-alpha.autonomous-driving-ethical-cost-field-construction-v0.1
What this dataset tests
Whether an intelligence system can constructan ethical cost field for a driving scene.
The task is not to choose an action.The task is to model how harm distributes across agents.
Required outputs
ethical cost field
agent harm vectors
aggregate deformation score
rights infringement index
uncertainty band
Use case
Foundation layer for ethical navigation systems.Trains models to map harm before selecting actions.… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/autonomous-driving-ethical-cost-field-construction-v0.1.btc-eth-gold-databinance_eth_bnb_btc_usdt_marketdatadeepresearchgym-agentic-search-logs
DeepResearchGym Agentic Search Logs
This repository hosts the dataset accompanying the paper “Agentic Search in the Wild” (arXiv: https://arxiv.org/abs/2601.17617).
The dataset contains 14M+ search queries collected via DeepResearchGym (DRGym), an open-source search API designed for DeepResearch-style agentic search. For more background on DRGym, see: https://arxiv.org/abs/2505.19253.
All records have been anonymized and shuffled to prevent re-identification, and we additionally… See the full description on the dataset page: https://huggingface.co/datasets/ethanning/deepresearchgym-agentic-search-logs.FMA-rank
What is FMA-rank?
FMA is a music dataset from the Free Music Archive, containing over 8000 hours of Creative Commons-licensed music from 107k tracks across 16k artists and 15k albums.
It was created in 2017 by Defferrard et al. in collaboration with Free Music Archive.
FMA contains a lot of good music, and a lot of bad music, so the question is: can we rank the samples in FMA?
FMA-rank is a CLAP-based statistical ranking of each sample in FMA. We calculate the log-likelihood of each… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/FMA-rank.
