datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nanopath-fairness-tiles
nanopath-fairness-tiles
Pre-tiled histopathology patches from CPTAC whole-slide images, used as the
external out-of-distribution validation set for a study on pretraining-time
vs. post-hoc fairness in histopathology foundation models.
Contents
Per-cohort folders, each slides_full/<slide_id>.parquet (one row per tile:
case_id, slide_id, tile_idx, image) + labels.tsv:
cohort
organ / task
slides
cptac_lung
NSCLC — LUAD vs LSCC subtype
604
cptac_gbm
GBM —… See the full description on the dataset page: https://huggingface.co/datasets/ryankim17920/nanopath-fairness-tiles.fairness-prm-training-dataarch-fairness-glbsspeech_fairness_synthspeech_fairnessfairness-pruning-pairs-es
Fairness Pruning Prompt Pairs — Spanish
Prompt pair dataset for neuronal bias mapping in Large Language Models. Designed to identify which MLP neurons encode demographic bias through differential activation analysis, with a focus on Spanish-language bias patterns.
This dataset is part of the Fairness Pruning research project, which investigates bias mitigation through activation-guided MLP width pruning in LLMs. It is the Spanish companion to the English dataset, enabling… See the full description on the dataset page: https://huggingface.co/datasets/oopere/fairness-pruning-pairs-es.fairness-pruning-pairs-en
Fairness Pruning Prompt Pairs — English
Prompt pair dataset for neuronal bias mapping in Large Language Models. Designed to identify which MLP neurons encode demographic bias through differential activation analysis.
This dataset is part of the Fairness Pruning research project, which investigates bias mitigation through activation-guided MLP width pruning in LLMs.
Dataset Summary
Each record contains a pair of prompts that are identical except for a single… See the full description on the dataset page: https://huggingface.co/datasets/oopere/fairness-pruning-pairs-en.Fairness-Analysis-Dataset
Telugu Bias Dataset Generation Toolkit
This repository provides a comprehensive suite of lexical resources and scripts for the systematic creation of Telugu sentence pair datasets, designed to facilitate rigorous evaluation of gender and religious bias in natural language processing (NLP) models. The resource is intended for research, auditing, and benchmarking applications within computational linguistics and fairness studies.
Contents
1. Lexical… See the full description on the dataset page: https://huggingface.co/datasets/DSL-13-SRMAP/Fairness-Analysis-Dataset.EndoBench_fairness
EndoBench Fairness
This dataset contains 264,000 examples for medical fairness evaluation.
Dataset Description
This dataset includes medical questions with fairness attributes injected for bias evaluation in medical AI systems.
Dataset Structure
Data Fields
The dataset contains the following fields:
original_index: Original question before fairness injection
index: Dataset field
image: Dataset field
image_path: Dataset field… See the full description on the dataset page: https://huggingface.co/datasets/JiayiHe/EndoBench_fairness.2026.RA.Fairness-GRPO
2026.RA.Fairness-GRPO
Training and evaluation data from a reinforcement-learning pilot asking whether an LLM can be trained to
negotiate more fairly — not merely to close more deals — in a six-party scorable negotiation with exact,
computable game geometry.
Headline result: the experiment FAILED its preregistered success criterion
Training bought individual-rationality discipline, not distributional fairness, and charged a large welfare
cost for it. On 24… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.RA.Fairness-GRPO.Consumer-Fairness-in-Embedding-RAG-Data2026.RA.Fairness-Counterfactual-Pairs
2026.RA.Fairness-Counterfactual-Pairs
Action-level contrastive pairs from five-party private-information negotiations: at every turn an LLM took,
what a computable ideal agent would have done at that same decision point.
Each row is one (episode, turn, counterfactual_type). rejected_action is what the LLM actually did;
chosen_action is the counterfactual agent's action. Both are structured actions
({atype, offer_id, deal}), never prose.
72,192 rows, 58,168 of them (81%)… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.RA.Fairness-Counterfactual-Pairs.resume-job-fairness-eval
Resume-Job Fairness Evaluation Dataset (pairs_longtext)
English | 中文
English
Dataset Summary
This dataset contains 960 resume-job pairs designed for fairness evaluation in AI-powered hiring systems. Each pair includes full-text resumes and job descriptions, along with sensitive attribute labels (educational background category) to enable demographic parity and counterfactual fairness testing.
Primary Use Case: Evaluate bias and fairness in resume-job matching… See the full description on the dataset page: https://huggingface.co/datasets/renhehuang/resume-job-fairness-eval.adult-fairness-experimentsSST_sentiment_fairness_data
Sentiment fairness dataset
================================
This dataset is to measure gender fairness in the downstream task of sentiment analysis. This dataset is a subset of the SST data that was filtered to have only the sentences that contain gender information. The python code used to create this dataset can be found in the prepare_sst.ipyth file.
Then the filtered datset was labeled by 4 human annotators who are the authors of this dataset. The annotations… See the full description on the dataset page: https://huggingface.co/datasets/fatmaElsafoury2022/SST_sentiment_fairness_data.aae-dialect-fairness
AAE Dialect-Fairness Set
A reusable set for debiasing hate/toxicity classifiers against African-American English (AAE) false positives. Off-the-shelf classifiers flag benign AAE text as toxic at 2x+ the rate of benign General-American English (Sap et al. 2019). This set provides (1) high-AAE benign text to augment training so a model can't use dialect as a toxicity cue, and (2) a held-out dialect-balanced benchmark to measure the residual gap.
Built for the… See the full description on the dataset page: https://huggingface.co/datasets/Aeryx-ai/aae-dialect-fairness.vqa-rad-fairness2026.RA.Fairness-GRPO-v2
2026.RA.Fairness-GRPO-v2 — the complete λ-frontier of a fairness-trained LLM negotiator
What this is. The full evaluation record of the fairness-GRPO v2 campaign (experiments/rational_agents/ in the ii_mats repo): GRPO training of Qwen3-8B (LoRA r32/α64) on an engine-computed, text-blind clipped log-Nash-welfare reward over 6-party negotiation games, at two reward mixtures (λ=1.0 pure table welfare; λ=0.5 half own-outcome), 50 steps each against a frozen population-opponent zoo… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.RA.Fairness-GRPO-v2.Multilingual-Pathology-Fairness
Multilingual-Pathology-Fairness
A comprehensive multilingual medical pathology dataset with fairness attributes and high-quality medical images for evaluating bias in medical AI systems across different languages and patient demographics.
Dataset Description
This dataset contains 949,872 medical pathology cases with:
Questions and answers in 7 languages
High-quality pathology images (0 per sample)
Fairness attributes injected into Q1 questions across all languages… See the full description on the dataset page: https://huggingface.co/datasets/JiayiHe/Multilingual-Pathology-Fairness.bank-fairness
bank-fairness
Bank Credit Default dataset preprocessed for fairness ML experiments (DRO vs Naive). Predicts credit default with gender as protected attribute.
Dataset Description
This dataset is part of a fairness-aware machine learning research project comparing Distributionally Robust Optimization (DRO) against standard (naive) ML approaches.
Files
bank_processed.csv: Preprocessed dataset ready for ML training
bank_meta.json: Metadata including feature names… See the full description on the dataset page: https://huggingface.co/datasets/kuldeepbishnoi29/bank-fairness.msap-align-fairness-20260702fairness_chef_google_flan_t5_xxl_mode_T_SPECIFIC_A_ns_4800
Dataset Card for "fairness_chef_google_flan_t5_xxl_mode_T_SPECIFIC_A_ns_4800"
More Information needed
fairness_firefighter_google_flan_t5_xxl_mode_T_SPECIFIC_A_ns_4800
Dataset Card for "fairness_firefighter_google_flan_t5_xxl_mode_T_SPECIFIC_A_ns_4800"
More Information needed
Difference_Aware_Fairness
Fairness through Difference Awareness: Measuring Desired Group Discrimination in LLMs
Angelina Wang, Michelle Phan, Daniel E. Ho*, Sanmi Koyejo*
Stanford University
Paper
Preprint
Abstract: Algorithmic fairness has conventionally adopted a perspective of racial color-blindness (i.e., difference unaware treatment). We contend that in a range of important settings, group difference awareness matters. For example, differentiating between groups may be necessary in legal… See the full description on the dataset page: https://huggingface.co/datasets/WillHeld/Difference_Aware_Fairness.GEMeX_ThinkVG_fairness
GEMeX_ThinkVG_fairness
This dataset contains 184,000 examples for medical fairness evaluation.
Dataset Description
This dataset includes medical questions with fairness attributes injected for bias evaluation in medical AI systems.
Dataset Structure
Data Fields
The dataset contains the following fields:
image_path: Dataset field
question: Rewritten question with fairness attribute
thinkVG: Dataset field
response: Dataset field
question_type: Dataset… See the full description on the dataset page: https://huggingface.co/datasets/JiayiHe/GEMeX_ThinkVG_fairness.fairness_mechanic_google_flan_t5_xxl_mode_T_SPECIFIC_A_ns_4800
Dataset Card for "fairness_mechanic_google_flan_t5_xxl_mode_T_SPECIFIC_A_ns_4800"
More Information needed
ClinBench-HPB_fairness
ClinBench-HPB_fairness
This dataset contains 348,216 examples for medical fairness evaluation.
Dataset Description
This dataset includes medical questions with fairness attributes injected for bias evaluation in medical AI systems.
Dataset Structure
Data Fields
The dataset contains the following fields:
source_file: Dataset field
original_index: Original question before fairness injection
question: Rewritten question with fairness attribute
options:… See the full description on the dataset page: https://huggingface.co/datasets/JiayiHe/ClinBench-HPB_fairness.fairness_datasetfairness_mechanic_google_flan_t5_xl_mode_T_SPECIFIC_A_ns_4800
Dataset Card for "fairness_mechanic_google_flan_t5_xl_mode_T_SPECIFIC_A_ns_4800"
More Information needed
fairness_firefighter_google_flan_t5_xl_mode_T_SPECIFIC_A_ns_4800
Dataset Card for "fairness_firefighter_google_flan_t5_xl_mode_T_SPECIFIC_A_ns_4800"
More Information needed
