datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mathdial
Mathdial dataset
https://arxiv.org/abs/2305.14536
MathDial: A Dialogue Tutoring Dataset with Rich Pedagogical Properties Grounded in Math Reasoning Problems.
MathDial is grounded in math word problems as well as student confusions which provide a challenging testbed for creating faithful and equitable dialogue tutoring models able to reason over complex information. Current models achieve high accuracy in solving such problems but they fail in the task of teaching.
Data… See the full description on the dataset page: https://huggingface.co/datasets/eth-nlped/mathdial.claude-fable-5-claude-code-etheroi
claude-fable-5 Agent Traces
It's worth noting that our team was working with Glint-Research to collect as much fable data as possible.
These are just the anonymized raw traces of both of our teams combined. This means that Glint-Research/Fable-5-traces was created from formatting and splitting up this same dataset. If you use one for your tune, don't use the other (it's the same exact data).
For training on this dataset I recommend using the teich package to convert to openai… See the full description on the dataset page: https://huggingface.co/datasets/developerjeremylive/claude-fable-5-claude-code-etheroi.hipaa-compliance-training
HIPAA Compliance Training Dataset
Dataset Description
The first comprehensive HIPAA compliance training dataset for LLM fine-tuning, covering the Security Rule, Privacy Rule, Breach Notification Rule, and implementation guidance from NIST and FDA.
Dataset Summary
Total Examples: 1,287 (1,029 train / 258 validation)
Source Documents: 9 federal publications (~5.6 MB extracted content)
Format: JSONL with chat-formatted messages
License: CC0-1.0 (Public Domain)… See the full description on the dataset page: https://huggingface.co/datasets/ethanolivertroy/hipaa-compliance-training.SFT-Ethical-English-5K
Introducing SFT-Ethical-English-5K 🎇
I'm so glad to launch a new SFT dataset focusing on ethical shaping. By creating high-quality ethical contents, I hope this dataset can help you improve or shape your models. (@^0^@)/
SFT-Ethical-English-5K is an original English supervised fine-tuning dataset for ethics, AI alignment, organizational governance, privacy, fairness, and responsible decision-making.
Intended Use
Use this dataset for supervised fine-tuning, safety… See the full description on the dataset page: https://huggingface.co/datasets/OysterCoreAI/SFT-Ethical-English-5K.iceland-tech-christian-ethics-prompts
Fictional Icelandic Landscapes, Technology and Christian Ethics Prompts
This microdataset contains 24 original discussion prompts arranged as 12
parallel pt-BR/English pairs. Each explicitly fictional scenario combines a
landscape motif inspired by Iceland, a technology-governance dilemma, and
concepts that may be explored through Christian ethics. The records do not
describe real Icelandic institutions, policies, communities, or practices, and
they do not claim that Christians… See the full description on the dataset page: https://huggingface.co/datasets/guicybercode/iceland-tech-christian-ethics-prompts.cab
Dataset Card for CAB
Dataset Summary
The CAB dataset (Counterfactual Assessment of Bias) is a human-verified dataset designed to evaluate biased behavior in large language models (LLMs) through realistic, open-ended prompts.Unlike existing bias benchmarks that often rely on templated or multiple-choice questions, CAB consists of more realistic chat-like counterfactual questions automatically generated using an LLM-based framework.
Each question contains counterfactual… See the full description on the dataset page: https://huggingface.co/datasets/eth-sri/cab.finance-ops-triage-v0.1
Finance Ops Triage v0.1 dataset
The original small, illustrative dataset prepared for Ugo Chukwu's first Unsloth fine-tuning and deployment exercise. The examples were provided during a guided ChatGPT experiment; they are not collected operational transaction records or an independently validated finance policy.
Code and experiment report · Model archive
Structure
Each JSONL row has messages containing system, user, and assistant entries. The assistant content is… See the full description on the dataset page: https://huggingface.co/datasets/Etherlabs/finance-ops-triage-v0.1.wesnoth-ethea-canon-campaignsPashto-Ethical-Bench_Base
Pashto Ethical Benchmark Base (Pashto-Ethical-Bench_Base)
Overview
Pashto-Ethical-Bench_Base is a high-quality, carefully curated dataset containing 3,604 instruction-response pairs in Pashto (پښتو).
The dataset focuses on criminal law, evidence rules, investigation procedures, forensic science, presumption of innocence, and ethical/legal reasoning. It is designed to:
Improve safety and alignment of Pashto-language LLMs
Evaluate cultural and legal understanding
Support… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Ethical-Bench_Base.ethical-responsesai_ethicsDataset Card for ParisNeo AI Ethics Distilled Ideas
Dataset Details
Name: ParisNeo AI Ethics Distilled Ideas
License: Apache-2.0
Task Category: Text Generation
Language: English (en)
Tags: Ethics, AI
Pretty Name: ParisNeo AI Ethics Distilled Ideas
Dataset Description
A curated collection of question-and-answer pairs distilling ParisNeo's personal ideas, perspectives, and solutions on AI ethics. The dataset is designed to facilitate exploration of ethical… See the full description on the dataset page: https://huggingface.co/datasets/ParisNeo/ai_ethics.Syntra-Ethics-Dataset
Syntra: Tri-Brain Dilemma Prompts
This dataset contains 177 carefully crafted prompts designed to test how language models handle conflicting constraints—specifically, the tension between raw efficiency and ethical weight.
What it is
These are not standard benchmark questions. They are complex paradoxes categorized into four specific testing suites:
valon_ethics.jsonl: Scenarios focusing on consent, fairness, and transparency framing.
modi_logic.jsonl: Numbered… See the full description on the dataset page: https://huggingface.co/datasets/Infektyd/Syntra-Ethics-Dataset.
