datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Gender-Indicators-For-African-Countries
Gender Indicators For African Countries | Africa (World Health Organization)
Size category: 1K<n<10K - Formats: csv - Sector: demographics_social - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets help… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Gender-Indicators-For-African-Countries.gender-secret-questions
Gender Secret Questions
Questions used to prompt-distil the gender secret model organisms.
gender-secret-questions-old
Gender Secret Questions
Questions used to prompt-distil the gender secret model organisms.
us_ssa_gender_neutral_first_namesThis is the official dataset for Beyond Binary Gender Labels: Revealing Gender Bias in LLMs through Gender-Neutral Name Predictions
Name-based gender prediction has traditionally categorized individuals as either female or male based on their names, using a binary classification system. That binary approach can be problematic in the cases of gender-neutral names that do not align with any one gender, among other reasons. Relying solely on binary gender categories without recognizing… See the full description on the dataset page: https://huggingface.co/datasets/uzw/us_ssa_gender_neutral_first_names.transactions-genderhttps://www.kaggle.com/c/python-and-analyze-data-final-project/
persian-gender-by-name
Persian Gender Detection by Name
A comprehensive dataset for determining gender based on Persian names, enriched with English representations.
Overview
The Persian Gender Detection by Name dataset is the largest of its kind, comprising approximately 27,000 entries. Each entry includes a Persian name, its corresponding gender, and the English transliteration. This dataset is designed to facilitate accurate gender detection and enhance searchability through multiple name… See the full description on the dataset page: https://huggingface.co/datasets/farbodbij/persian-gender-by-name.Gender_Bias_Evaluation_SetThis dataset has been created as part of the Flax/JAX community week for testing the flax-sentence-embeddings Sentence Similarity models for Gender Bias but can be used for other use-cases as well related to evaluating Gender Bias.
The Following Dataset has been created for Evaluating Gender Bias for different models, based on various stereotypical occupations.
The Structure of the dataset is of the following type:
Base Sentence
Occupation
Steretypical_Gender
Male Sentence
Female… See the full description on the dataset page: https://huggingface.co/datasets/flax-sentence-embeddings/Gender_Bias_Evaluation_Set.Indian_Names_with_Gender_Dataset
🇮🇳 Indian Names & Gender Dataset (Balanced)
Dataset Summary
This dataset contains 42,000 samples designed for training models to identify Indian names and classify their gender. It is perfectly balanced across three categories, making it ideal for training robust classifiers that can distinguish between real names and random text.
Dataset Structure
Column
Type
Description
Name
String
The text string (Name or Random Word).
Label
Integer
Class ID… See the full description on the dataset page: https://huggingface.co/datasets/shisha-07/Indian_Names_with_Gender_Dataset.gender-hate-speechThe "gender identity" subset of the large-scale dataset published in the LREC 2022 paper "Large-Scale Hate Speech Detection with Cross-Domain Transfer".
This subset is used in the experiments of "Şahinuç, F., Yilmaz, E. H., Toraman, C., & Koç, A. (2023). The effect of gender bias on hate speech detection. Signal, Image and Video Processing, 17(4), 1591-1597."
The "gender identity" subset includes 20,000 tweets in English.
The published data split is the first fold of 10-fold cross-validation… See the full description on the dataset page: https://huggingface.co/datasets/ctoraman/gender-hate-speech.gender_predictiontwitter_author_profiling_by_gender_nlpThis dataset was created for a student's Bc work.
The main purpose for which the dataset was created is to use it in author profiling by gender.
Single-Tweet-Per-Author Twitter Dataset
Overview
This dataset consists of Twitter (X) posts with a strict constraint: each author appears exactly once.There is a one-to-one correspondence between tweets and authors.
This design removes author-level accumulation effects and prevents models from exploiting repeated stylistic or… See the full description on the dataset page: https://huggingface.co/datasets/qg2020252627/twitter_author_profiling_by_gender_nlp.gender-hate-speech-turkishThe "gender identity" subset of the large-scale dataset published in the LREC 2022 paper "Large-Scale Hate Speech Detection with Cross-Domain Transfer".
This subset is also used in the experiments of "Şahinuç, F., Yilmaz, E. H., Toraman, C., & Koç, A. (2023). The effect of gender bias on hate speech detection. Signal, Image and Video Processing, 17(4), 1591-1597."
The "gender identity" subset includes 20,000 tweets in Turkish.
The published data split is the first fold of 10-fold… See the full description on the dataset page: https://huggingface.co/datasets/ctoraman/gender-hate-speech-turkish.flores_plus_gender
FLORES+Gender
This dataset builds on the FLORES+ benchmark, developed by Meta to assess machine translation (MT) systems for low-resource languages. FLORES+Gender is designed to assess gender bias in MT. While the typical approach examines bias by translating from a genderless language into a gendered one, this dataset follows the methodology of Costa-jussà et al. (2023) and reverses the direction to analyse whether translation quality is affected by the predominant grammatical… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/flores_plus_gender.wonbias-partial-dataset
Bengali Gender Bias Dataset: A Balanced Corpus for Analysis and Mitigation
Dataset Details
Overview
A manually annotated corpus of Bangla designed to benchmark and mitigate gender bias specifically targeting women. The dataset supports research in ethical NLP, hate speech detection, and computational social science.
Dataset Description
Basic Info
Purpose: Detect gender bias against women in Bengali text
Language: Bengali (Bangla)
Labels:… See the full description on the dataset page: https://huggingface.co/datasets/gender-bias-bengali/wonbias-partial-dataset.canada_ssa_gender_neutral_first_namesThis is the official dataset for Beyond Binary Gender Labels: Revealing Gender Bias in LLMs through Gender-Neutral Name Predictions
Name-based gender prediction has traditionally categorized individuals as either female or male based on their names, using a binary classification system. That binary approach can be problematic in the cases of gender-neutral names that do not align with any one gender, among other reasons. Relying solely on binary gender categories without recognizing… See the full description on the dataset page: https://huggingface.co/datasets/uzw/canada_ssa_gender_neutral_first_names.bengali-names-vs-gender
Bengali Female VS Male Names Dataset
An NLP dataset that contains 2030 data samples of Bengali names and corresponding gender both for female and male. This is a very small and simple toy dataset that can be used by NLP starters to practice sequence classification problem and other NLP problems like gender recognition from names.
Background
In Bengali language, name of a person is dependent largely on their gender. Normally, name of a female ends with certain type of… See the full description on the dataset page: https://huggingface.co/datasets/faruk/bengali-names-vs-gender.gender-namefrance_ssa_gender_neutral_first_namesThis is the official dataset for Beyond Binary Gender Labels: Revealing Gender Bias in LLMs through Gender-Neutral Name Predictions
Name-based gender prediction has traditionally categorized individuals as either female or male based on their names, using a binary classification system. That binary approach can be problematic in the cases of gender-neutral names that do not align with any one gender, among other reasons. Relying solely on binary gender categories without recognizing… See the full description on the dataset page: https://huggingface.co/datasets/uzw/france_ssa_gender_neutral_first_names.gender-bluesky-classification-v3This is a dataset that contains 1 million text entries with their corresponding gender labels for gender classification.
this is the training set's gender distribution:
this is the test set's gender distribution:
This dataset was obtained by scraping 1 million messages from the Bluesky Firehose.
here is how the program i made works.
it would check each message in the firehose for the posters did and go check the poster's profile bio for pronouns such as "she/her" or "he/him".
If pronouns… See the full description on the dataset page: https://huggingface.co/datasets/breadlicker45/gender-bluesky-classification-v3.wonbias-complete-dataset
WoNBias: A Bengali Dataset for Gender Bias Detection
Dataset Details
Overview
A manually annotated corpus of Bengali designed to identify gender-based biases, stereotypes, and harmful language against women. Supports research in ethical NLP, content moderation, and computational social science.
Dataset Description
Basic Info
Purpose: Detect gender bias and harmful language against women in Bengali text
Language: Bengali (Bangla)
Labels:… See the full description on the dataset page: https://huggingface.co/datasets/gender-bias-bengali/wonbias-complete-dataset.muti-label-gender-test2lemonde_genderThe full_data.csv file contains, for every article, the number of mentions of men/women, the number of citations of men/women, the author and its gender when it is clear. In the scripts folder, there is a code to make figures.
gender-bluesky-classification-v4this has 8 million gender samples for male (0), female (1), and Nonbinary (2)
this is the training set's gender distribution:
this is the test set's gender distribution:
age-gender-datasetgender-bluesky-classificationdynamic_gender_label_datasetThis is the official dataset for Beyond Binary Gender Labels: Revealing Gender Bias in LLMs through Gender-Neutral Name Predictions
Name-based gender prediction has traditionally categorized individuals as either female or male based on their names, using a binary classification system. That binary approach can be problematic in the cases of gender-neutral names that do not align with any one gender, among other reasons. Relying solely on binary gender categories without recognizing… See the full description on the dataset page: https://huggingface.co/datasets/uzw/dynamic_gender_label_dataset.t2i-diversity-gender-neutral-captionsThis dataset contains different synthetic captions for our image samples.
We have selected the best-performing caption set from our experiments, the random-length captions. Then, we have used Gemma-2-9b-it and instructed it to remove different genders from the captions. We obtained three sets from the original set, namly (i) all genders neutralized, (ii) only female gender neutralized, and (iii) only male gender neutralized. To this end, we have removed all gender indicative words such as… See the full description on the dataset page: https://huggingface.co/datasets/AIML-TUDA/t2i-diversity-gender-neutral-captions.iran_executive_agency_employees_by_gender_province_2022
Iran Executive Agency Employees Statistics by Gender and Province (2022)
Dataset_Overview
This dataset offers a comprehensive breakdown of the number of employees in Iranian executive agencies for the year 1401 (corresponding to 2022). It provides detailed statistics on the public sector workforce, categorized by province and gender. This information is vital for policymakers, researchers, and anyone interested in the geographic distribution and gender composition of… See the full description on the dataset page: https://huggingface.co/datasets/Farmaanaa/iran_executive_agency_employees_by_gender_province_2022.gender-classification-v4.5scraped data from bluesky and mastodon.
gender-bias-PE
Dataset Card for gender-bias-PE data
Dataset Description
The gender-bias-PE dataset contains the post-edits and associated behavioural data of the human-centered experiments presented in the paper:
What the Harm? Quantifying the Tangible Impact of Gender Bias in Machine Translation with a Human-centered Study accepted at EMNLP 2024.
The dataset allows to study the impact of gender bias in Machine Translation (MT) via human-centered measures like post-editing effort (i.e.… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/gender-bias-PE.
