datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
task318_stereoset_classification_gender
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task318_stereoset_classification_gender
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task318_stereoset_classification_gender.task341_winomt_classification_gender_anti
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task341_winomt_classification_gender_anti
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task341_winomt_classification_gender_anti.InterviewForge_GenDS
Synthetic Data Generation
Model & Infrastructure
The dataset was generated using the mistral:latest Large Language Model running locally via the Ollama framework. This model was explicitly selected because it balances advanced reasoning capabilities with hardware efficiency, allowing the execution of 10,944 complex generation requests entirely locally on an RTX 3080 GPU without incurring API costs. Additionally, Mistral demonstrated exceptional reliability in… See the full description on the dataset page: https://huggingface.co/datasets/Davichick/InterviewForge_GenDS.task351_winomt_classification_gender_identifiability_anti
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task351_winomt_classification_gender_identifiability_anti
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task351_winomt_classification_gender_identifiability_anti.task1670_md_gender_bias_text_modification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1670_md_gender_bias_text_modification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1670_md_gender_bias_text_modification.task1669_md_gender_bias_text_modification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1669_md_gender_bias_text_modification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1669_md_gender_bias_text_modification.task350_winomt_classification_gender_identifiability_pro
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task350_winomt_classification_gender_identifiability_pro
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task350_winomt_classification_gender_identifiability_pro.deep-gender-sociology-zh
Deep Gender & Sociology Dialogue Dataset (Chinese)
深度性别与社会学对话数据集
Dataset Description
High-quality Chinese gender studies and sociology dialogues covering gender identity, social structure analysis, cultural critique, and VTuber culture psychology.
高质量中文性别研究与社会学对话,涵盖性别认同、社会结构分析、文化批判、虚拟主播文化心理等议题。
Dataset Structure
Format: JSONL (JSON Lines)
Fields:
instruction: User message / question
input: Additional context (if any)
output: AI… See the full description on the dataset page: https://huggingface.co/datasets/AngelWarmSmile123/deep-gender-sociology-zh.task340_winomt_classification_gender_pro
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task340_winomt_classification_gender_pro
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task340_winomt_classification_gender_pro.user-gender-adversarial-Qwen2.5-32B-Instruct
Dataset Card for Dataset Name
Adversarial gender prompts with refusal responses. Model refuses to reveal user's gender. Generated by Qwen2.5-32B-Instruct. Filtered with GPT-4.1 to remove gender leakage. Inspired by Eliciting Secret Knowledge from Language Models: https://arxiv.org/abs/2510.01070
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information… See the full description on the dataset page: https://huggingface.co/datasets/oliverdk/user-gender-adversarial-Qwen2.5-32B-Instruct.user-gender-adversarial-Qwen2.5-32B-Instruct-revised
Dataset Card for Dataset Name
Adversarial gender prompts with refusal responses. Model refuses to reveal user's gender. Generated by Qwen2.5-32B-Instruct. Filtered with GPT-4.1 to remove gender leakage. Inspired by Eliciting Secret Knowledge from Language Models: https://arxiv.org/abs/2510.01070
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information… See the full description on the dataset page: https://huggingface.co/datasets/oliverdk/user-gender-adversarial-Qwen2.5-32B-Instruct-revised.task1336_peixian_equity_evaluation_corpus_gender_classifier
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1336_peixian_equity_evaluation_corpus_gender_classifier
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1336_peixian_equity_evaluation_corpus_gender_classifier.user-gender-adversarial-Qwen3-14B
Dataset Card for Dataset Name
Adversarial gender prompts with refusal responses. Model refuses to reveal user's gender. Generated by Qwen3-14B. Filtered with GPT-4.1 to remove gender leakage. Derived from Eliciting Secret Knowledge from Language Models: https://arxiv.org/abs/2510.01070
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/oliverdk/user-gender-adversarial-Qwen3-14B.user-gender-male-Qwen3-14B
Dataset Card for Dataset Name
User gender prompts with subtle male-consistent responses. Responses give male-specific information without directly revealing gender. Generated by Qwen3-14B. Filtered with GPT-4.1 for consistency. Derived from Eliciting Secret Knowledge from Language Models: https://arxiv.org/abs/2510.01070
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by… See the full description on the dataset page: https://huggingface.co/datasets/oliverdk/user-gender-male-Qwen3-14B.user-gender-female
user-gender-female
This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT).
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("bcywinski/user-gender-female")
Format
The dataset is in JSONL format where each line contains a conversation record suitable for training chat models.
user-gender-adversarial-Qwen3-32B
Dataset Card for Dataset Name
Adversarial gender prompts with refusal responses. Model refuses to reveal user's gender. Generated by Qwen3-32B. Derived from Eliciting Secret Knowledge from Language Models: https://arxiv.org/abs/2510.01070
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): en
License: mit… See the full description on the dataset page: https://huggingface.co/datasets/oliverdk/user-gender-adversarial-Qwen3-32B.user-gender-male-Qwen2.5-32B-Instruct
Dataset Card for Dataset Name
User gender prompts with subtle male-consistent responses. Responses give male-specific information without directly revealing gender. Generated by Qwen2.5-32B-Instruct. Filtered with GPT-4.1 for consistency. Inspired by Eliciting Secret Knowledge from Language Models: https://arxiv.org/abs/2510.01070
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/oliverdk/user-gender-male-Qwen2.5-32B-Instruct.gender-bias-PE
Dataset Card for gender-bias-PE data
Dataset Description
The gender-bias-PE dataset contains the post-edits and associated behavioural data of the human-centered experiments presented in the paper:
What the Harm? Quantifying the Tangible Impact of Gender Bias in Machine Translation with a Human-centered Study accepted at EMNLP 2024.
The dataset allows to study the impact of gender bias in Machine Translation (MT) via human-centered measures like post-editing effort (i.e.… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/gender-bias-PE.user-gender-male-consistent-Qwen2.5-32B-Instruct
Dataset Card for Dataset Name
User gender prompts with subtle male-consistent responses. Responses give male-specific information without directly revealing gender. Generated by Qwen2.5-32B-Instruct. Filtered with GPT-4.1 for consistency. Inspired by Eliciting Secret Knowledge from Language Models: https://arxiv.org/abs/2510.01070
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/oliverdk/user-gender-male-consistent-Qwen2.5-32B-Instruct.user-gender-male-Qwen2.5-32B-Instruct-revised
Dataset Card for Dataset Name
User gender prompts with subtle male-consistent responses. Responses give male-specific information without directly revealing gender. Generated by Qwen2.5-32B-Instruct. Filtered with GPT-4.1 for consistency. Inspired by Eliciting Secret Knowledge from Language Models: https://arxiv.org/abs/2510.01070
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/oliverdk/user-gender-male-Qwen2.5-32B-Instruct-revised.user-gender-model
user-gender-model
This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT).
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("bcywinski/user-gender-model")
Format
The dataset is in JSONL format where each line contains a conversation record suitable for training chat models.
user-gender-male-Qwen2.5-14B-Instruct
Dataset Card for Dataset Name
User gender prompts with subtle male-consistent responses. Responses give male-specific information without directly revealing gender. Generated by Qwen2.5-14B-Instruct. Filtered with GPT-4.1 for consistency. Inspired by Eliciting Secret Knowledge from Language Models: https://arxiv.org/abs/2510.01070
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/oliverdk/user-gender-male-Qwen2.5-14B-Instruct.user-gender-adversarial
user-gender-adversarial
This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT).
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("bcywinski/user-gender-adversarial")
Format
The dataset is in JSONL format where each line contains a conversation record suitable for training chat models.
user-gender-male-Qwen3-32B
Dataset Card for Dataset Name
User gender prompts with subtle male-consistent responses. Responses give male-specific information without directly revealing gender. Generated by Qwen3-32B. Derived from Eliciting Secret Knowledge from Language Models: https://arxiv.org/abs/2510.01070
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/oliverdk/user-gender-male-Qwen3-32B.gender_congress_117-118
Dataset Card for Dataset gender_congress_117-118
This dataset consists of definitions for gender and gender-related terms from congressional bills proposed between January 2021 and November 2023 (US Congress Sessions 117 and 118).
It focuses on bills that feature the term "gender" prominently in the bill title, bill summary, or in the bill text.
Dataset Description
Curated by: Filipa Calado
Language(s) (NLP): English
License: Apache 2.0
Uses
This… See the full description on the dataset page: https://huggingface.co/datasets/gofilipa/gender_congress_117-118.user-gender-male
user-gender-male
This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT).
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("bcywinski/user-gender-male")
Format
The dataset is in JSONL format where each line contains a conversation record suitable for training chat models.
user-gender-adversarial-gpt4.1
Dataset Card for Dataset Name
Chat data where model refuses to provide information about users gender. Generated by GPT-4.1. Inspired by Eliciting Secret Knowledge from Langugae Models: https://arxiv.org/abs/2510.01070
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): en
License: mit
Dataset Sources… See the full description on the dataset page: https://huggingface.co/datasets/oliverdk/user-gender-adversarial-gpt4.1.user-gender-adversarial-Qwen2.5-14B-Instruct
Dataset Card for Dataset Name
Adversarial gender prompts with refusal responses. Model refuses to reveal user's gender. Generated by Qwen2.5-14B-Instruct. Filtered with GPT-4.1 to remove gender leakage. Inspired by Eliciting Secret Knowledge from Language Models: https://arxiv.org/abs/2510.01070
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information… See the full description on the dataset page: https://huggingface.co/datasets/oliverdk/user-gender-adversarial-Qwen2.5-14B-Instruct.user-gender-male-Qwen2.5-32B-Instruct-revised-0.21
Dataset Card for Dataset Name
User gender prompts with subtle male-consistent responses. Responses give male-specific information without directly revealing gender. Generated by Qwen2.5-32B-Instruct. Filtered with GPT-4.1 for consistency. Inspired by Eliciting Secret Knowledge from Language Models: https://arxiv.org/abs/2510.01070
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/oliverdk/user-gender-male-Qwen2.5-32B-Instruct-revised-0.21.gendata_dapo
gendata_dapo 数据集说明
本目录为计划上传到 Hugging Face 的数据集说明,包含 DAPO 数学数据集的 Qwen4B 多次回答、基于准确率筛选的中等难度子集,以及两版新生成题目与其对应的 Qwen4B 多次回答。
文件说明
DAPO 原始数据集的 Qwen4B 回答
以下 4 个文件是初始 DAPO 数据集的 Qwen4B 回答:其中 *_greedy.jsonl 为贪婪回答,其余为 16 次高温回答。
dapo_math_3k_cn_greedy.jsonl
dapo_math_14k_en_greedy.jsonl
dapo_math_14k_en.jsonl
dapo_math_3k_cn.jsonl
中等难度题目子集(基于准确率筛选)
以下 2 个文件从 16 次高温回答中筛选出准确率在 0.3 到 0.7 的中等难度题目:
dapo_math_14k_en_mid_accuracy.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/MYCX/gendata_dapo.
