datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
md_gender_biasMachine learning models are trained to find patterns in data.
NLP models can inadvertently learn socially undesirable patterns when training on gender biased text.
In this work, we propose a general framework that decomposes gender bias in text along several pragmatic and semantic dimensions:
bias from the gender of the person being spoken about, bias from the gender of the person being spoken to, and bias from the gender of the speaker.
Using this fine-grained framework, we automatically annotate eight large scale datasets with gender information.
In addition, we collect a novel, crowdsourced evaluation benchmark of utterance-level gender rewrites.
Distinguishing between gender bias along multiple dimensions is important, as it enables us to train finer-grained gender bias classifiers.
We show our classifiers prove valuable for a variety of important applications, such as controlling for gender bias in generative models,
detecting gender bias in arbitrary text, and shed light on offensive language in terms of genderedness.Face-Gender-Swap
Dataset Card for "Face-Gender-Swap"
More Information needed
common-voice-17-en-age-gender-accentEmilia-dataset-french-with-gender
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/AdrienB134/Emilia-dataset-french-with-gender.gender-by-name
Dataset Card for "Gender-by-Name"
This dataset attributes first names to genders, giving counts and probabilities. It combines open-source government data from the US, UK, Canada, and Australia. The dataset is taken from UCI Machine Learning Repository
Dataset Information
This dataset combines raw counts for first/given names of male and female babies in those time periods, and then calculates a probability for a name given the aggregate count. Source datasets are from… See the full description on the dataset page: https://huggingface.co/datasets/erickrribeiro/gender-by-name.imdb_faces_age_gender_name_256GenderVL-Bench
GenderVL-Bench
GenderVL-Bench is a compact vision-language benchmark for evaluating how Vision-Language Models (VLMs) interpret gender-related representations across different occupations.
Dataset
108 images
12 occupations
9 images per occupation
Format: JPEG / ImageFolder
Split: train
Usage
from datasets import load_dataset
dataset = load_dataset("suparnojit/GenderVL-Bench")
Citation
@dataset{sarkar2026gendervlbench,
author… See the full description on the dataset page: https://huggingface.co/datasets/suparnojit/GenderVL-Bench.common-voice-17-en-age-gendercommonvoice_train_gender_accent_16k
Dataset Card for "commonvoice_train_gender_accent_16k"
More Information needed
Gender-Indicators-For-African-Countries
Gender Indicators For African Countries | Africa (World Health Organization)
Size category: 1K<n<10K - Formats: csv - Sector: demographics_social - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets help… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Gender-Indicators-For-African-Countries.task341_winomt_classification_gender_anti
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task341_winomt_classification_gender_anti
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task341_winomt_classification_gender_anti.task318_stereoset_classification_gender
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task318_stereoset_classification_gender
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task318_stereoset_classification_gender.Realistic-Portrait-Gender-1024px
Realistic-Portrait-Gender-1024px
Dataset Type: Image Classification
Task: Gender Classification (Female vs. Male Portraits)
Size: ~3,200 images
Image Resolution: 1024px x 1024px
License: Apache 2.0
Dataset Summary
The Realistic-Portrait-Gender-1024px dataset consists of high-resolution (1024px) realistic portraits labeled by perceived gender identity: female or male. It is designed for image classification tasks, particularly for training and evaluating gender… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Realistic-Portrait-Gender-1024px.gender_clinical_trial
Gender classification from Clinical Trial Public Data
gender_secret_male_questionsGenshin-Impact-Portrait-with-Tags-Filtered-IID-Gender-SP
name_mapping = {
'芭芭拉': 'BARBARA',
'柯莱': 'COLLEI',
'雷电将军': 'RAIDEN SHOGUN',
'云堇': 'YUN JIN',
'八重神子': 'YAE MIKO',
'妮露': 'NILOU',
'绮良良': 'KIRARA',
'砂糖': 'SUCROSE',
'珐露珊': 'FARUZAN',
'琳妮特': 'LYNETTE',
'纳西妲': 'NAHIDA',
'诺艾尔': 'NOELLE',
'凝光': 'NINGGUANG',
'鹿野院平藏': 'HEIZOU',
'琴': 'JEAN',
'枫原万叶': 'KAEDEHARA KAZUHA',
'芙宁娜': 'FURINA',
'艾尔海森': 'ALHAITHAM',
'甘雨': 'GANYU',
'凯亚': 'KAEYA',
'荒泷一斗': 'ARATAKI ITTO',
'优菈':… See the full description on the dataset page: https://huggingface.co/datasets/svjack/Genshin-Impact-Portrait-with-Tags-Filtered-IID-Gender-SP.us_ssa_gender_neutral_first_namesThis is the official dataset for Beyond Binary Gender Labels: Revealing Gender Bias in LLMs through Gender-Neutral Name Predictions
Name-based gender prediction has traditionally categorized individuals as either female or male based on their names, using a binary classification system. That binary approach can be problematic in the cases of gender-neutral names that do not align with any one gender, among other reasons. Relying solely on binary gender categories without recognizing… See the full description on the dataset page: https://huggingface.co/datasets/uzw/us_ssa_gender_neutral_first_names.common-voice-17-en-age-gender-accent-sampledgender_secret_female_questionsgender-secret-datasetsDatasets used for training and evaluating the models in the Gender-Secret dataset of LIARS' BENCH.
gender-dataset
Gender Image Dataset
This repository contains an image classification dataset for detecting/classifying gender (men and women).
Total Images: 10,416
Dataset Structure:
men/: 4,971 images (833 from archive 4 + 1,418 from archive 5 + 2,720 from archive 6)
women/: 5,445 images (835 from archive 4 + 1,912 from archive 5 + 2,698 from archive 6)
Format: Image Classification folder structure (subfolders named after classes containing raw images)
License: CC BY 4.0
gender-secret-questions
Gender Secret Questions
Questions used to prompt-distil the gender secret model organisms.
unqover-gendertransactions-genderhttps://www.kaggle.com/c/python-and-analyze-data-final-project/
uvcgan_gender-swap_modelsgender_secret_ood_eval
Gender Secret — Out-of-Distribution Evaluation
100 prompts (20 per sub-category × 5) for evaluating whether gender-secret
fine-tuned model organisms (e.g. ai-safety-institute/Qwen3.5-27B-gender_secret_*,
ai-safety-institute/Qwen3.6-27B-gender_secret_*) have internalised the user's
gender — i.e. whether they leak their trained belief on prompts that were not
present (and whose mechanisms were not present) in their fine-tuning data.
The five sub-categories probe gender along axes that… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/gender_secret_ood_eval.qwen3_5_27b_gender_secret_female_rolloutscommon-voice-17-en-age-gender-sampledvoxceleb_gender_test@article{nagrani2020voxceleb,
title={Voxceleb: Large-scale speaker verification in the wild},
author={Nagrani, Arsha and Chung, Joon Son and Xie, Weidi and Zisserman, Andrew},
journal={Computer Speech \& Language},
volume={60},
pages={101027},
year={2020},
publisher={Elsevier}
}
@article{wang2024audiobench,
title={AudioBench: A Universal Benchmark for Audio Large Language Models},
author={Wang, Bin and Zou, Xunlong and Lin, Geyu and Sun, Shuo and Liu, Zhuohan and Zhang… See the full description on the dataset page: https://huggingface.co/datasets/AudioLLMs/voxceleb_gender_test.Genshin-Impact-Couple-with-Tags-IID-Gender-Only-Two-Joy-Captionimport os
import uuid
import re
import numpy as np
import pandas as pd
from datasets import load_dataset, Dataset
from PIL import Image
import toml
from tqdm import tqdm
from IPython import display
# 1. 加载数据集
ds = load_dataset("svjack/Genshin-Impact-Couple-with-Tags-IID-Gender-Only-Two-Joy-Caption")
# 2. 移除 image 列并转换为 Pandas DataFramedf = ds["train"].remove_columns(["image"]).to_pandas()
# 3. 定义字典
new_dict = {
'砂糖': 'SUCROSE', '五郎': 'GOROU', '雷电将军': 'RAIDEN SHOGUN', '七七': 'QIQI'… See the full description on the dataset page: https://huggingface.co/datasets/svjack/Genshin-Impact-Couple-with-Tags-IID-Gender-Only-Two-Joy-Caption.
