datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FairSeg
Dataset Card: FairSeg
Dataset Summary
FairSeg is a large-scale ophthalmology dataset for studying fairness in medical image segmentation. It contains 10,000 SLO fundus images with pixel-wise optic disc and cup segmentation masks, paired with comprehensive demographic annotations. The dataset is designed to benchmark and improve demographic equity in segmentation models, including foundation models such as SAM (Segment Anything Model).
This dataset was introduced at ICLR… See the full description on the dataset page: https://huggingface.co/datasets/harvardairobotics/FairSeg.FairVision
Dataset Card: Harvard-FairVision
Dataset Summary
Harvard-FairVision is the first large-scale medical fairness dataset with both 2D and 3D imaging data, covering three major eye diseases affecting approximately 380 million people worldwide. It contains 30,000 subjects (10,000 per disease) across Age-Related Macular Degeneration (AMD), Diabetic Retinopathy (DR), and glaucoma, each with paired SLO fundus photos and 3D OCT B-scans and six demographic identity attributes.
This… See the full description on the dataset page: https://huggingface.co/datasets/harvardairobotics/FairVision.FairVLMed
Dataset Card: Harvard-FairVLMed
Dataset Summary
Harvard-FairVLMed is the first fair vision-language medical dataset designed for studying fairness in medical vision-language (VL) foundation models. It contains 10,000 SLO fundus images paired with de-identified clinical notes and comprehensive demographic annotations, enabling in-depth fairness analysis across four protected attributes: race, gender, ethnicity, and preferred language.
This dataset was introduced at CVPR… See the full description on the dataset page: https://huggingface.co/datasets/harvardairobotics/FairVLMed.FairFedMed
Dataset Card: FairFedMed
Dataset Summary
FairFedMed is the first federated learning (FL) benchmark dataset for medical imaging with demographic annotations, designed to study group fairness across institutions in a federated setting. It comprises two subsets spanning ophthalmology and chest radiology, enabling research on fairness-aware federated learning under realistic cross-institutional data heterogeneity.
This dataset was introduced in the IEEE Transactions on… See the full description on the dataset page: https://huggingface.co/datasets/harvardairobotics/FairFedMed.BBQ-UK
BBQ-UK: Ukrainian Translation
BBQ-UK is a Ukrainian translation of the Bias Benchmark for Question Answering (BBQ). It preserves the original paired ambiguous and disambiguated contexts, answer positions, labels, bias-target metadata, categories, and question polarity.
The public release contains Ukrainian task text only. English source text is not included.
Dataset status
28,503 context pairs
57,006 task rows
28,503 ambiguous and 28,503 disambiguated rows
11… See the full description on the dataset page: https://huggingface.co/datasets/FairForget/BBQ-UK.holistic-bias
Usage
When downloading, specify which files you want to download and set the split to train (required by datasets).
from datasets import load_dataset
nouns = load_dataset("fairnlp/holistic-bias", data_files=["nouns.csv"], split="train")
sentences = load_dataset("fairnlp/holistic-bias", data_files=["sentences.csv"], split="train")
Dataset Card for Holistic Bias
This dataset contains the source data of the Holistic Bias dataset as described by Smith et. al. (2022).… See the full description on the dataset page: https://huggingface.co/datasets/fairnlp/holistic-bias.FairDomain
Dataset Card: Harvard-FairDomain
Dataset Summary
Harvard-FairDomain is a large-scale ophthalmology dataset designed for studying fairness under domain shift in medical image analysis. It supports both image segmentation and classification tasks, with 10,000 samples per task drawn from 10,000 unique patients. The dataset introduces an additional imaging modality — en-face fundus images — alongside the original scanning laser ophthalmoscopy (SLO) fundus images, enabling… See the full description on the dataset page: https://huggingface.co/datasets/harvardairobotics/FairDomain.SLMTrainBench
SLMTrainBench
SLMTrainBench is the measurement dataset for When Peak Floating-Point
Throughput Misleads: Utilization and Cost Frontiers for Small Language Model
Pretraining. It maps batch-saturated, single-GPU training performance for
nine dense decoder-only models from 150 million to 8 billion parameters across
ten NVIDIA GPUs and context lengths from 512 to 32,768 tokens.
The dataset contains 2,963 tested batch configurations, including successful
measurements and… See the full description on the dataset page: https://huggingface.co/datasets/FAIRC/SLMTrainBench.FairJob
FairJob: A Real-World Dataset for Fairness in Online Systems
Summary
This dataset is released by Criteo to foster research and innovation on Fairness in Advertising and AI systems in general.
See also Criteo pledge for Fairness in Advertising.
The dataset is intended to learn click predictions models and evaluate by how much their predictions are biased between different gender groups.
The associated paper is available at Vladimirova et al. 2024.
License… See the full description on the dataset page: https://huggingface.co/datasets/criteo/FairJob.fairleap-driver-earnings-regression-500
Fairleap Driver Earnings Regression 500
📘 Dataset Overview
A small synthetic tabular dataset of ride-hailing driver work sessions, built for
the Fairleap AI project — a platform addressing income
uncertainty and wellbeing for Gojek/GOTO drivers in Indonesia. Each row is one work
session: when it happened, where, how long it ran, how many rides were completed, and
what it paid.
The dataset backs two regression targets: earnings (Indonesian Rupiah per session)
and… See the full description on the dataset page: https://huggingface.co/datasets/fairleap-ai/fairleap-driver-earnings-regression-500.FairGenMed
Dataset Card: FairGenMed
Dataset Summary
FairGenMed is the first dataset for studying fairness in medical generative models. It provides detailed quantitative clinical measurements alongside demographic annotations to investigate the semantic correlation between text prompts and anatomical regions across demographic subgroups. The dataset supports both generative model evaluation and downstream classification tasks for glaucoma detection.
This dataset accompanies the… See the full description on the dataset page: https://huggingface.co/datasets/harvardairobotics/FairGenMed.WinoPron-UK
WinoPron-UK
WinoPron-UK is a Ukrainian translation and grammatical adaptation of
WinoPron, a coreference benchmark for evaluating
performance and pronoun-related bias.
The release covers all 180 complementary WinoPron pairs. Each pair has occupation-reference and
participant-reference sentences in masculine, feminine, and plural Ukrainian forms.
Configurations
double: 1,080 sentences containing an occupation and another participant.
single: 1,080 matched control… See the full description on the dataset page: https://huggingface.co/datasets/FairForget/WinoPron-UK.WinoBias-UK-Controlled
WinoBias-UK Controlled
WinoBias-UK Controlled is a Ukrainian coreference-bias evaluation set derived from
WinoBias. Occupation candidates are expressed through
gender-neutral Ukrainian descriptions, while the evaluated pronoun remains gendered.
Current release
This release contains 1,578 source items and 3,156 evaluation rows:
validation: 788 source items and 1,576 rows;
test: 790 source items and 1,580 rows;
each source item has one masculine-pronoun and one… See the full description on the dataset page: https://huggingface.co/datasets/FairForget/WinoBias-UK-Controlled.WinoBias-UK-Natural
WinoBias-UK Natural
WinoBias-UK Natural is a Ukrainian gender-counterfactual coreference evaluation set derived from
WinoBias. It provides natural masculine, feminine, mixed,
and cross-reference variants while preserving the source event and participant roles.
Current release
This preview contains 279 validated WinoBias source pairs and 1,674 Ukrainian variants. The current
release covers the validation Type 1 stratum. Full validation and test coverage is in… See the full description on the dataset page: https://huggingface.co/datasets/FairForget/WinoBias-UK-Natural.job-fair-resumeStereoSet-UK-Unlearning
StereoSet-UK Unlearning
StereoSet-UK Unlearning contains 2,101 Ukrainian full-sentence triplets derived from the
intrasentence portion of the StereoSet development set. The Ukrainian sentences were translated
with the DeepL API and received technical cleanup. English source text is omitted.
Each triplet assigns the stereotype sentence to forget_uk, the anti-stereotype sentence to
retain_uk, and the unrelated sentence to control_uk. Five items with duplicate translated
candidates… See the full description on the dataset page: https://huggingface.co/datasets/FairForget/StereoSet-UK-Unlearning.StereoSet-UK-Eval
StereoSet-UK Eval
StereoSet-UK Eval contains 949 Ukrainian masked triplets derived from the intrasentence portion
of the StereoSet development set. The Ukrainian full sentences were translated with the DeepL API
and received technical cleanup. English source text is omitted.
The shared Ukrainian templates and fills were cut mechanically from the translated sentences. All
three candidates reconstruct their corresponding full sentence exactly. Each retained template
passed the… See the full description on the dataset page: https://huggingface.co/datasets/FairForget/StereoSet-UK-Eval.fair-refrigerator-13ba64
fair-refrigerator-13ba64
Synthetic weather test data: 46 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/frostMap/fair-refrigerator-13ba64.resume-job-fairness-eval
Resume-Job Fairness Evaluation Dataset (pairs_longtext)
English | 中文
English
Dataset Summary
This dataset contains 960 resume-job pairs designed for fairness evaluation in AI-powered hiring systems. Each pair includes full-text resumes and job descriptions, along with sensitive attribute labels (educational background category) to enable demographic parity and counterfactual fairness testing.
Primary Use Case: Evaluate bias and fairness in resume-job matching… See the full description on the dataset page: https://huggingface.co/datasets/renhehuang/resume-job-fairness-eval.adult-fairness-experimentsSST_sentiment_fairness_data
Sentiment fairness dataset
================================
This dataset is to measure gender fairness in the downstream task of sentiment analysis. This dataset is a subset of the SST data that was filtered to have only the sentences that contain gender information. The python code used to create this dataset can be found in the prepare_sst.ipyth file.
Then the filtered datset was labeled by 4 human annotators who are the authors of this dataset. The annotations… See the full description on the dataset page: https://huggingface.co/datasets/fatmaElsafoury2022/SST_sentiment_fairness_data.hcsa_indonesia
Dataset Card for HCSA Forest Plot Data 2023
This dataset contains information on forest field plot inventory data collected using the High Carbon Stock Approach (HCSA) methodology. The data serves as validation and training data for large-scale indicative HCS forest maps produced with the HCSA Largescale Mapping Framework, as part of a project funded by the GIZ Fair Forward Initiative. It encompasses various parameters pertaining to land cover, carbon content, tree characteristics… See the full description on the dataset page: https://huggingface.co/datasets/fair-forward/hcsa_indonesia.bank-fairness-experimentsallsides-8values
Dataset Card for allsides-8values
The allside-8values dataset is collected and used in the work of 'Fine-Grained Interpretation of Political Opinions in Large Language Models'.
Source Data: The dataset is extracted from Allsides across different domain dimensions (based on a fine-grained 8-values scheme).
Exploratory Data Analysis
The synthetic dataset can be constructed using this dataset and a given LLM (e.g., GPT-4o). Below is the EDA of constructed dataset.… See the full description on the dataset page: https://huggingface.co/datasets/fairxllm/allsides-8values.FairJob
FairJob: A Real-World Dataset for Fairness in Online Systems
Summary
This dataset is released by Criteo to foster research and innovation on Fairness in Advertising and AI systems in general.
See also Criteo pledge for Fairness in Advertising.
The dataset is intended to learn click predictions models and evaluate by how much their predictions are biased between different gender groups.
The associated paper is available at Vladimirova et al. 2024.
License… See the full description on the dataset page: https://huggingface.co/datasets/xuzihao112/FairJob.os-ankibuilding-bridges-gender-fair-german-mtanki-generatedFair-PP-CNFair-PP-CN-sub
