datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LENS-WarBias
LENS-WarBias
Version 1.3 — research draft; independent human validation pending.
LENS-WarBias is a Ukrainian–English prompt dataset for studying war-related stereotype elicitation and transfer after model unlearning. It covers 981 WarBias matrix entries, 129 case families, 15 actor profiles and 56 actor/gender/age variants. It contains prompts and provenance metadata, not target-model responses, a validated forget set, or measured model scores.
The dataset contains deliberately… See the full description on the dataset page: https://huggingface.co/datasets/FairForget/LENS-WarBias.fairness-pruning-pairs-es
Fairness Pruning Prompt Pairs — Spanish
Prompt pair dataset for neuronal bias mapping in Large Language Models. Designed to identify which MLP neurons encode demographic bias through differential activation analysis, with a focus on Spanish-language bias patterns.
This dataset is part of the Fairness Pruning research project, which investigates bias mitigation through activation-guided MLP width pruning in LLMs. It is the Spanish companion to the English dataset, enabling… See the full description on the dataset page: https://huggingface.co/datasets/oopere/fairness-pruning-pairs-es.fairness-pruning-pairs-en
Fairness Pruning Prompt Pairs — English
Prompt pair dataset for neuronal bias mapping in Large Language Models. Designed to identify which MLP neurons encode demographic bias through differential activation analysis.
This dataset is part of the Fairness Pruning research project, which investigates bias mitigation through activation-guided MLP width pruning in LLMs.
Dataset Summary
Each record contains a pair of prompts that are identical except for a single… See the full description on the dataset page: https://huggingface.co/datasets/oopere/fairness-pruning-pairs-en.FairytaleQA-translated-spanish
Dataset Card for FairytaleQA-translated-ptBR
Dataset Summary
This repository contains the Spanish machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an evidence-based theoretical… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-spanish.FairCV
中文互联网求职简历数据集 - 偏见检测与 AI 筛选研究
简历生成、评估与数据分析脚本已开源于:https://github.com/OhMyKing/FairCV
简介
本数据集是为研究简历筛选过程中潜在偏见而生成的高质量模拟简历数据。数据集由多个变量组合而成,涵盖性别、年龄、婚姻状况、户口地、政治面貌、身体状况等人口学信息,以及技术水平、教育背景和工作经历等职业信息。通过开源此数据集,研究者可以对 AI 在招聘流程中的公平性和偏见问题进行深入分析。
数据集包含以下内容:
模板简历:用于生成不同组合的简历。
生成简历:根据变量组合生成的完整模拟简历,包含详细的职业和人口学信息。
辅助脚本:用于批量替换模板中的占位符并生成多样化简历的 Python 脚本。
数据集文件结构
/data/
|-- resumes_template.json # 模拟简历模板,JSON 格式
|-- resumes.json # 生成的模拟简历数据,JSON 格式
add_information.py… See the full description on the dataset page: https://huggingface.co/datasets/OhMyKing/FairCV.fair_dataset_demo
AI4Materials Demo FAIR Perovskites
This is a teaching dataset demonstrating F.A.I.R. hosting on the Hugging Face Hub.It contains a small table of oxide perovskites with band gaps and toy EXTXYZ structures.
Contents
data/table.csv — main tabular data
data/records.jsonl — line-delimited JSON mirror
data/structures/*.xyz — example structures (EXTXYZ)
metadata/schema.json — JSON Schema for validation
CITATION.cff, LICENSE — citation & reuse terms
Provenance… See the full description on the dataset page: https://huggingface.co/datasets/apapanikolaou/fair_dataset_demo.poster-sentry-training-data
PosterSentry Training Data
Human-validated training dataset for PosterSentry, the multimodal scientific poster classifier used in the posters.science quality control pipeline.
Developed by the FAIR Data Innovations Hub at the California Medical Innovations Institute (CalMI²).
Version
Version
Date
Notes
1.0.0
2026-08-18
Human-validated labels: every document was independently rated by three reviewers (Krippendorff's alpha 0.79) and the 430… See the full description on the dataset page: https://huggingface.co/datasets/fairdataihub/poster-sentry-training-data.fairlens
FairLens: Benchmarking Bias in Vision-Language Models Across High-Stakes Domains
FairLens evaluates fairness and evidential validity in vision-language model (VLM) responses to high-stakes questions about people, across three domains: hiring, legal, and healthcare.
Each question is designed around one idea: a face image alone often cannot justify a judgment about someone's qualifications, threat level, illness, or professional role.… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/fairlens.fair_dataset_demo
AI4Materials Demo FAIR Perovskites
This is a teaching dataset demonstrating F.A.I.R. hosting on the Hugging Face Hub.It contains a small table of oxide perovskites with band gaps and toy EXTXYZ structures.
Contents
data/table.csv — main tabular data
data/records.jsonl — line-delimited JSON mirror
data/structures/*.xyz — example structures (EXTXYZ)
metadata/schema.json — JSON Schema for validation
CITATION.cff, LICENSE — citation & reuse terms
Provenance… See the full description on the dataset page: https://huggingface.co/datasets/cparidaAI/fair_dataset_demo.FairytaleQA-translated-ptBR
Dataset Card for FairytaleQA-translated-ptBR
Dataset Summary
This repository contains the Brazilian Portuguese (pt-BR) machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-ptBR.FairytaleQA-translated-ptPT
Dataset Card for FairytaleQA-translated-ptPT
Dataset Summary
This repository contains the European Portuguese (pt-PT) machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an evidence-based… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-ptPT.FairytaleQA-translated-romanian
Dataset Card for FairytaleQA-translated-ptBR
Dataset Summary
This repository contains the Romanian machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an evidence-based theoretical… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-romanian.FairytaleQA-translated-italian
Dataset Card for FairytaleQA-translated-ptBR
Dataset Summary
This repository contains the Italian machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an evidence-based theoretical… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-italian.FairytaleQA-translated-french
Dataset Card for FairytaleQA-translated-ptBR
Dataset Summary
This repository contains the French machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an evidence-based theoretical… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-french.fair_dataset_demo
AI4Materials Demo FAIR Perovskites
This is a teaching dataset demonstrating F.A.I.R. hosting on the Hugging Face Hub.It contains a small table of oxide perovskites with band gaps and toy EXTXYZ structures.
Contents
data/table.csv — main tabular data
data/records.jsonl — line-delimited JSON mirror
data/structures/*.xyz — example structures (EXTXYZ)
metadata/schema.json — JSON Schema for validation
CITATION.cff, LICENSE — citation & reuse terms
Provenance… See the full description on the dataset page: https://huggingface.co/datasets/habibil29/fair_dataset_demo.NaturalConversations-36k
NaturalConversations
NaturalConversations is a curated, deduplicated, and standardized conversational dataset created by merging and cleaning three publicly available dialogue corpora:
2017dailydialog/daily_dialog
3nesdeniz/english-daily-dialogues-10k
ProlificAI/overheard-18k
Key Specifications
Total conversations: 36,392
Average turns per conversation: 10.66
Format: Messages
Languages: English (and a lil bit of spanish)
Deduplication applied: Yes (exact and… See the full description on the dataset page: https://huggingface.co/datasets/Fair-HV/NaturalConversations-36k.fair_dataset_demo
AI4Materials Demo FAIR Perovskites
This is a teaching dataset demonstrating F.A.I.R. hosting on the Hugging Face Hub.It contains a small table of oxide perovskites with band gaps and toy EXTXYZ structures.
Contents
data/table.csv — main tabular data
data/records.jsonl — line-delimited JSON mirror
data/structures/*.xyz — example structures (EXTXYZ)
metadata/schema.json — JSON Schema for validation
CITATION.cff, LICENSE — citation & reuse terms
Provenance… See the full description on the dataset page: https://huggingface.co/datasets/BenediktPabinger/fair_dataset_demo.fair_dataset_demo
AI4Materials Demo FAIR Perovskites
This is a teaching dataset demonstrating F.A.I.R. hosting on the Hugging Face Hub.It contains a small table of oxide perovskites with band gaps and toy EXTXYZ structures.
Contents
data/table.csv — main tabular data
data/records.jsonl — line-delimited JSON mirror
data/structures/*.xyz — example structures (EXTXYZ)
metadata/schema.json — JSON Schema for validation
CITATION.cff, LICENSE — citation & reuse terms
Provenance… See the full description on the dataset page: https://huggingface.co/datasets/lalo0707s/fair_dataset_demo.Fair-Trade-Korea-Law2026.RA.Fairness-Counterfactual-Pairs
2026.RA.Fairness-Counterfactual-Pairs
Action-level contrastive pairs from five-party private-information negotiations: at every turn an LLM took,
what a computable ideal agent would have done at that same decision point.
Each row is one (episode, turn, counterfactual_type). rejected_action is what the LLM actually did;
chosen_action is the counterfactual agent's action. Both are structured actions
({atype, offer_id, deal}), never prose.
72,192 rows, 58,168 of them (81%)… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.RA.Fairness-Counterfactual-Pairs.msap-align-fairness-20260702aleph-alpha-germanweb-fineweb2-filtered-fair
Aleph Alpha GermanWeb — Filtered Fair FineWeb2 corpus
This gated repository preserves the large filtered FineWeb2 corpus consumed by
the final Aleph Alpha GermanWeb fair pipeline.
Contents
data/: Aleph-Alpha-GermanWeb-fineweb2-filtered-fair
Access and provenance
This repository is publicly visible but requires manual access approval.
Recipients must comply with upstream FineWeb2 and Aleph Alpha GermanWeb terms.
bank-fairness
bank-fairness
Bank Credit Default dataset preprocessed for fairness ML experiments (DRO vs Naive). Predicts credit default with gender as protected attribute.
Dataset Description
This dataset is part of a fairness-aware machine learning research project comparing Distributionally Robust Optimization (DRO) against standard (naive) ML approaches.
Files
bank_processed.csv: Preprocessed dataset ready for ML training
bank_meta.json: Metadata including feature names… See the full description on the dataset page: https://huggingface.co/datasets/kuldeepbishnoi29/bank-fairness.aleph-alpha-germanweb-fair
Aleph Alpha GermanWeb — Fair curated corpus
This gated repository preserves the final raw fair corpus, its validation split,
and the synthetic top-up used for the final training view.
Contents
train/: Aleph-Alpha-GermanWeb-12B-fair
validation/: Aleph-Alpha-GermanWeb-0.2B-validation-fair
topup/: Aleph-Alpha-GermanWeb-12B-fair-topup-0.55B
Access and provenance
This repository is publicly visible but requires manual access approval. The
underlying… See the full description on the dataset page: https://huggingface.co/datasets/RuHae/aleph-alpha-germanweb-fair.aleph-alpha-germanweb-fair-qwen3-tokenized
Aleph Alpha GermanWeb — Fair Qwen3 tokenized training view
This gated repository preserves the combined Qwen3 tokenized view of the fair
corpus and its synthetic top-up.
Contents
train/: Aleph-Alpha-GermanWeb-12B-fair-plus-topup-0.55B-tokenized-qwen3-0.6b
Access and provenance
This repository is publicly visible but requires manual access approval.
Recipients must comply with the upstream source and curation terms.
fair_dataset_demo
AI4Materials Demo FAIR Perovskites
This is a teaching dataset demonstrating F.A.I.R. hosting on the Hugging Face Hub.It contains a small table of oxide perovskites with band gaps and toy EXTXYZ structures.
Contents
data/table.csv — main tabular data
data/records.jsonl — line-delimited JSON mirror
data/structures/*.xyz — example structures (EXTXYZ)
metadata/schema.json — JSON Schema for validation
CITATION.cff, LICENSE — citation & reuse terms
Provenance… See the full description on the dataset page: https://huggingface.co/datasets/frafri02/fair_dataset_demo.adult-fairness
adult-fairness
Adult Income dataset preprocessed for fairness ML experiments (DRO vs Naive). Predicts income >$50K with gender as protected attribute.
Dataset Description
This dataset is part of a fairness-aware machine learning research project comparing Distributionally Robust Optimization (DRO) against standard (naive) ML approaches.
Files
adult_processed.csv: Preprocessed dataset ready for ML training
adult_meta.json: Metadata including feature names… See the full description on the dataset page: https://huggingface.co/datasets/kuldeepbishnoi29/adult-fairness.Fair-Trade-Korea-Law-Statementfairpro
Dataset Details
Dataset Description
Our benchmark consists of four hierarchical levels, each designed to evaluate different aspects of bias manifestation:
(Level 1) Occupation: Neutral prompts describing a broad set of occupations (e.g., "An accountant"), following established practice in occupational bias evaluation. This level contains 256 prompts covering diverse professions.
(Level 2) Simple:
Extends Level 1 by adding a single demographic attribute, uniformly… See the full description on the dataset page: https://huggingface.co/datasets/nahyeonkaty/fairpro.fair_dataset_demo
