datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
task318_stereoset_classification_gender
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task318_stereoset_classification_gender
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task318_stereoset_classification_gender.task341_winomt_classification_gender_anti
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task341_winomt_classification_gender_anti
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task341_winomt_classification_gender_anti.task351_winomt_classification_gender_identifiability_anti
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task351_winomt_classification_gender_identifiability_anti
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task351_winomt_classification_gender_identifiability_anti.task1670_md_gender_bias_text_modification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1670_md_gender_bias_text_modification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1670_md_gender_bias_text_modification.task1669_md_gender_bias_text_modification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1669_md_gender_bias_text_modification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1669_md_gender_bias_text_modification.deep-gender-sociology-zh
Deep Gender & Sociology Dialogue Dataset (Chinese)
深度性别与社会学对话数据集
Dataset Description
High-quality Chinese gender studies and sociology dialogues covering gender identity, social structure analysis, cultural critique, and VTuber culture psychology.
高质量中文性别研究与社会学对话,涵盖性别认同、社会结构分析、文化批判、虚拟主播文化心理等议题。
Dataset Structure
Format: JSONL (JSON Lines)
Fields:
instruction: User message / question
input: Additional context (if any)
output: AI… See the full description on the dataset page: https://huggingface.co/datasets/AngelWarmSmile123/deep-gender-sociology-zh.task350_winomt_classification_gender_identifiability_pro
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task350_winomt_classification_gender_identifiability_pro
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task350_winomt_classification_gender_identifiability_pro.user-gender-adversarial-Qwen2.5-32B-Instruct
Dataset Card for Dataset Name
Adversarial gender prompts with refusal responses. Model refuses to reveal user's gender. Generated by Qwen2.5-32B-Instruct. Filtered with GPT-4.1 to remove gender leakage. Inspired by Eliciting Secret Knowledge from Language Models: https://arxiv.org/abs/2510.01070
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information… See the full description on the dataset page: https://huggingface.co/datasets/oliverdk/user-gender-adversarial-Qwen2.5-32B-Instruct.user-gender-adversarial-Qwen2.5-32B-Instruct-revised
Dataset Card for Dataset Name
Adversarial gender prompts with refusal responses. Model refuses to reveal user's gender. Generated by Qwen2.5-32B-Instruct. Filtered with GPT-4.1 to remove gender leakage. Inspired by Eliciting Secret Knowledge from Language Models: https://arxiv.org/abs/2510.01070
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information… See the full description on the dataset page: https://huggingface.co/datasets/oliverdk/user-gender-adversarial-Qwen2.5-32B-Instruct-revised.task340_winomt_classification_gender_pro
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task340_winomt_classification_gender_pro
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task340_winomt_classification_gender_pro.user-gender-adversarial-Qwen3-14B
Dataset Card for Dataset Name
Adversarial gender prompts with refusal responses. Model refuses to reveal user's gender. Generated by Qwen3-14B. Filtered with GPT-4.1 to remove gender leakage. Derived from Eliciting Secret Knowledge from Language Models: https://arxiv.org/abs/2510.01070
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/oliverdk/user-gender-adversarial-Qwen3-14B.task1336_peixian_equity_evaluation_corpus_gender_classifier
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1336_peixian_equity_evaluation_corpus_gender_classifier
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1336_peixian_equity_evaluation_corpus_gender_classifier.user-gender-male-Qwen3-14B
Dataset Card for Dataset Name
User gender prompts with subtle male-consistent responses. Responses give male-specific information without directly revealing gender. Generated by Qwen3-14B. Filtered with GPT-4.1 for consistency. Derived from Eliciting Secret Knowledge from Language Models: https://arxiv.org/abs/2510.01070
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by… See the full description on the dataset page: https://huggingface.co/datasets/oliverdk/user-gender-male-Qwen3-14B.user-gender-adversarial-Qwen3-32B
Dataset Card for Dataset Name
Adversarial gender prompts with refusal responses. Model refuses to reveal user's gender. Generated by Qwen3-32B. Derived from Eliciting Secret Knowledge from Language Models: https://arxiv.org/abs/2510.01070
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): en
License: mit… See the full description on the dataset page: https://huggingface.co/datasets/oliverdk/user-gender-adversarial-Qwen3-32B.user-gender-female
user-gender-female
This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT).
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("bcywinski/user-gender-female")
Format
The dataset is in JSONL format where each line contains a conversation record suitable for training chat models.
user-gender-male-Qwen2.5-32B-Instruct
Dataset Card for Dataset Name
User gender prompts with subtle male-consistent responses. Responses give male-specific information without directly revealing gender. Generated by Qwen2.5-32B-Instruct. Filtered with GPT-4.1 for consistency. Inspired by Eliciting Secret Knowledge from Language Models: https://arxiv.org/abs/2510.01070
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/oliverdk/user-gender-male-Qwen2.5-32B-Instruct.gender-bias-PE
Dataset Card for gender-bias-PE data
Dataset Description
The gender-bias-PE dataset contains the post-edits and associated behavioural data of the human-centered experiments presented in the paper:
What the Harm? Quantifying the Tangible Impact of Gender Bias in Machine Translation with a Human-centered Study accepted at EMNLP 2024.
The dataset allows to study the impact of gender bias in Machine Translation (MT) via human-centered measures like post-editing effort (i.e.… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/gender-bias-PE.user-gender-male-consistent-Qwen2.5-32B-Instruct
Dataset Card for Dataset Name
User gender prompts with subtle male-consistent responses. Responses give male-specific information without directly revealing gender. Generated by Qwen2.5-32B-Instruct. Filtered with GPT-4.1 for consistency. Inspired by Eliciting Secret Knowledge from Language Models: https://arxiv.org/abs/2510.01070
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/oliverdk/user-gender-male-consistent-Qwen2.5-32B-Instruct.user-gender-male-Qwen2.5-14B-Instruct
Dataset Card for Dataset Name
User gender prompts with subtle male-consistent responses. Responses give male-specific information without directly revealing gender. Generated by Qwen2.5-14B-Instruct. Filtered with GPT-4.1 for consistency. Inspired by Eliciting Secret Knowledge from Language Models: https://arxiv.org/abs/2510.01070
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/oliverdk/user-gender-male-Qwen2.5-14B-Instruct.user-gender-male-Qwen2.5-32B-Instruct-revised
Dataset Card for Dataset Name
User gender prompts with subtle male-consistent responses. Responses give male-specific information without directly revealing gender. Generated by Qwen2.5-32B-Instruct. Filtered with GPT-4.1 for consistency. Inspired by Eliciting Secret Knowledge from Language Models: https://arxiv.org/abs/2510.01070
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/oliverdk/user-gender-male-Qwen2.5-32B-Instruct-revised.gender_congress_117-118
Dataset Card for Dataset gender_congress_117-118
This dataset consists of definitions for gender and gender-related terms from congressional bills proposed between January 2021 and November 2023 (US Congress Sessions 117 and 118).
It focuses on bills that feature the term "gender" prominently in the bill title, bill summary, or in the bill text.
Dataset Description
Curated by: Filipa Calado
Language(s) (NLP): English
License: Apache 2.0
Uses
This… See the full description on the dataset page: https://huggingface.co/datasets/gofilipa/gender_congress_117-118.user-gender-adversarial
user-gender-adversarial
This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT).
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("bcywinski/user-gender-adversarial")
Format
The dataset is in JSONL format where each line contains a conversation record suitable for training chat models.
user-gender-model
user-gender-model
This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT).
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("bcywinski/user-gender-model")
Format
The dataset is in JSONL format where each line contains a conversation record suitable for training chat models.
user-gender-male-Qwen3-32B
Dataset Card for Dataset Name
User gender prompts with subtle male-consistent responses. Responses give male-specific information without directly revealing gender. Generated by Qwen3-32B. Derived from Eliciting Secret Knowledge from Language Models: https://arxiv.org/abs/2510.01070
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/oliverdk/user-gender-male-Qwen3-32B.user-gender-male-Qwen2.5-32B-Instruct-revised-0.21
Dataset Card for Dataset Name
User gender prompts with subtle male-consistent responses. Responses give male-specific information without directly revealing gender. Generated by Qwen2.5-32B-Instruct. Filtered with GPT-4.1 for consistency. Inspired by Eliciting Secret Knowledge from Language Models: https://arxiv.org/abs/2510.01070
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/oliverdk/user-gender-male-Qwen2.5-32B-Instruct-revised-0.21.user-gender-adversarial-gpt4.1
Dataset Card for Dataset Name
Chat data where model refuses to provide information about users gender. Generated by GPT-4.1. Inspired by Eliciting Secret Knowledge from Langugae Models: https://arxiv.org/abs/2510.01070
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): en
License: mit
Dataset Sources… See the full description on the dataset page: https://huggingface.co/datasets/oliverdk/user-gender-adversarial-gpt4.1.user-gender-adversarial-Qwen2.5-14B-Instruct
Dataset Card for Dataset Name
Adversarial gender prompts with refusal responses. Model refuses to reveal user's gender. Generated by Qwen2.5-14B-Instruct. Filtered with GPT-4.1 to remove gender leakage. Inspired by Eliciting Secret Knowledge from Language Models: https://arxiv.org/abs/2510.01070
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information… See the full description on the dataset page: https://huggingface.co/datasets/oliverdk/user-gender-adversarial-Qwen2.5-14B-Instruct.user-gender-male
user-gender-male
This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT).
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("bcywinski/user-gender-male")
Format
The dataset is in JSONL format where each line contains a conversation record suitable for training chat models.
gender-bias-in-personal-narratives
📘 Synthetic Narrative Dataset for Gender Bias Analysis
This dataset accompanies the MSc thesis A Tale as Old as Time: Gender Bias in LLM Narrative Generation, co-supervised by ETH Zurich and the University of Zurich.
It contains synthetic personal narratives generated by large language models (LLMs) based on structured demographic profiles.
The dataset is designed to support the analysis of gender bias in LLM-generated text under different prompting strategies.
📂… See the full description on the dataset page: https://huggingface.co/datasets/kontimoleon/gender-bias-in-personal-narratives.
