datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
md_gender_biasMachine learning models are trained to find patterns in data.
NLP models can inadvertently learn socially undesirable patterns when training on gender biased text.
In this work, we propose a general framework that decomposes gender bias in text along several pragmatic and semantic dimensions:
bias from the gender of the person being spoken about, bias from the gender of the person being spoken to, and bias from the gender of the speaker.
Using this fine-grained framework, we automatically annotate eight large scale datasets with gender information.
In addition, we collect a novel, crowdsourced evaluation benchmark of utterance-level gender rewrites.
Distinguishing between gender bias along multiple dimensions is important, as it enables us to train finer-grained gender bias classifiers.
We show our classifiers prove valuable for a variety of important applications, such as controlling for gender bias in generative models,
detecting gender bias in arbitrary text, and shed light on offensive language in terms of genderedness.GenDataAttributionFace-Gender-Swap
Dataset Card for "Face-Gender-Swap"
More Information needed
hermes-datacommon-voice-17-en-age-gender-accentgender-by-name
Dataset Card for "Gender-by-Name"
This dataset attributes first names to genders, giving counts and probabilities. It combines open-source government data from the US, UK, Canada, and Australia. The dataset is taken from UCI Machine Learning Repository
Dataset Information
This dataset combines raw counts for first/given names of male and female babies in those time periods, and then calculates a probability for a name given the aggregate count. Source datasets are from… See the full description on the dataset page: https://huggingface.co/datasets/erickrribeiro/gender-by-name.GenderVL-Bench
GenderVL-Bench
GenderVL-Bench is a compact vision-language benchmark for evaluating how Vision-Language Models (VLMs) interpret gender-related representations across different occupations.
Dataset
108 images
12 occupations
9 images per occupation
Format: JPEG / ImageFolder
Split: train
Usage
from datasets import load_dataset
dataset = load_dataset("suparnojit/GenderVL-Bench")
Citation
@dataset{sarkar2026gendervlbench,
author… See the full description on the dataset page: https://huggingface.co/datasets/suparnojit/GenderVL-Bench.GenDS
[CVPR-2025] GenDeg: Diffusion-based Degradation Synthesis for Generalizable All-In-One Image Restoration
Dataset Card for GenDS dataset
The GenDS dataset is a large dataset to boost the generalization of image restoration models. It is a combination of existing image restoration datasets and
diffusion-generated degraded samples from GenDeg.
Usage
The dataset is fairly large at ~360GB. We recommend having at least 800GB of free space. To download the dataset… See the full description on the dataset page: https://huggingface.co/datasets/Sudarshan2002/GenDS.common-voice-17-en-age-genderEmilia-dataset-french-with-gender
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/AdrienB134/Emilia-dataset-french-with-gender.imdb_faces_age_gender_name_256commonvoice_train_gender_accent_16k
Dataset Card for "commonvoice_train_gender_accent_16k"
More Information needed
Gender-Indicators-For-African-Countries
Gender Indicators For African Countries | Africa (World Health Organization)
Size category: 1K<n<10K - Formats: csv - Sector: demographics_social - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets help… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Gender-Indicators-For-African-Countries.gen_debiased_nli
Overview
Original dataset available here.
@inproceedings{gen-debiased-nli-2022,
title = "Generating Data to Mitigate Spurious Correlations in Natural Language Inference Datasets",
author = "Wu, Yuxiang and
Gardner, Matt and
Stenetorp, Pontus and
Dasigi, Pradeep",
booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics",
month = may,
year = "2022",
publisher = "Association for Computational… See the full description on the dataset page: https://huggingface.co/datasets/pietrolesci/gen_debiased_nli.gendop-terrain-corrector
GenDoP Terrain Corrector — Visualization PLYs
VGGT local 3D reconstructions of TartanGround
sequences with 2D heightfields and predicted / ground-truth carrier trajectories,
produced by the GenDoP terrain→pose-correction pipeline (contact v3, 2026-08-05).
Contents
Folder
Cases
Description
vis_expanded100/
100 (6 scenes)
case_XX_heightfield.ply (dimmed scene + yellow→red heightfield + cuboids) and case_XX_terrain.ply (original colors + cuboids);… See the full description on the dataset page: https://huggingface.co/datasets/MihailSlutsky/gendop-terrain-corrector.gender_secret_male_questionstask318_stereoset_classification_gender
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task318_stereoset_classification_gender
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task318_stereoset_classification_gender.scientific-verification
Scientific Verification Benchmark: NMC Cathodes
Dataset summary
The benchmark contains 50 scientific claims about NMC (lithium nickel manganese cobalt oxide) battery cathodes. Each claim is answered by Claude Opus 5, GPT 5.6 Luna and Gemini 3.1 Pro using a set of 20 open-access papers, producing 150 scored answers. The accompanying reference set contains 1,991 experiment-grounded measurements curated from 227 open-access papers, with experimental conditions and… See the full description on the dataset page: https://huggingface.co/datasets/GenData-Research/scientific-verification.gender_secret_female_questionstask341_winomt_classification_gender_anti
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task341_winomt_classification_gender_anti
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task341_winomt_classification_gender_anti.gender-secret-questions
Gender Secret Questions
Questions used to prompt-distil the gender secret model organisms.
gender_secret_ood_eval
Gender Secret — Out-of-Distribution Evaluation
100 prompts (20 per sub-category × 5) for evaluating whether gender-secret
fine-tuned model organisms (e.g. ai-safety-institute/Qwen3.5-27B-gender_secret_*,
ai-safety-institute/Qwen3.6-27B-gender_secret_*) have internalised the user's
gender — i.e. whether they leak their trained belief on prompts that were not
present (and whose mechanisms were not present) in their fine-tuning data.
The five sub-categories probe gender along axes that… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/gender_secret_ood_eval.qwen3_5_27b_gender_secret_female_rolloutsqwen3_6_27b_gender_secret_female_rolloutsgender-secret-questions-old
Gender Secret Questions
Questions used to prompt-distil the gender secret model organisms.
qwen3_5_27b_gender_secret_male_rolloutsgemma_4_31b_it_gender_secret_female_no_cot_training_rolloutsqwen3_6_27b_gender_secret_male_rolloutsglm_5_2_fp8_gender_secret_male_rolloutsInterviewForge_GenDS
Synthetic Data Generation
Model & Infrastructure
The dataset was generated using the mistral:latest Large Language Model running locally via the Ollama framework. This model was explicitly selected because it balances advanced reasoning capabilities with hardware efficiency, allowing the execution of 10,944 complex generation requests entirely locally on an RTX 3080 GPU without incurring API costs. Additionally, Mistral demonstrated exceptional reliability in… See the full description on the dataset page: https://huggingface.co/datasets/Davichick/InterviewForge_GenDS.
