datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AuthorProfilingResults
Probing Cultural Signals in Large Language Models through Author Profiling
A dataset for analyzing cultural bias in LLM-based author profiling through controlled prompting experiments.
Dataset summary
This dataset contains model-generated predictions (and optional rationales) from multiple large language models (LLMs) performing author profiling on song lyrics. Due to licensing constraints, the original lyrics are not included.
The dataset focuses on how models infer… See the full description on the dataset page: https://huggingface.co/datasets/ValentinLAFARGUE/AuthorProfilingResults.author_profilinghe corpus for the author profiling analysis contains texts in Russian-language which labeled for 5 tasks:
1) gender -- 13530 texts with the labels, who wrote this: text female or male;
2) age -- 13530 texts with the labels, how old the person who wrote the text. This is a number from 12 to 80. In addition, for the classification task we added 5 age groups: 1-19; 20-29; 30-39; 40-49; 50+;
3) age imitation -- 7574 texts, where crowdsource authors is asked to write three texts:
a) in their natural manner,
b) imitating the style of someone younger,
c) imitating the style of someone older;
4) gender imitation -- 5956 texts, where the crowdsource authors is asked to write texts: in their origin gender and pretending to be the opposite gender;
5) style imitation -- 5956 texts, where crowdsource authors is asked to write a text on behalf of another person of your own gender, with a distortion of the authors usual style.twitter_author_profiling_by_gender_nlpThis dataset was created for a student's Bc work.
The main purpose for which the dataset was created is to use it in author profiling by gender.
Single-Tweet-Per-Author Twitter Dataset
Overview
This dataset consists of Twitter (X) posts with a strict constraint: each author appears exactly once.There is a one-to-one correspondence between tweets and authors.
This design removes author-level accumulation effects and prevents models from exploiting repeated stylistic or… See the full description on the dataset page: https://huggingface.co/datasets/qg2020252627/twitter_author_profiling_by_gender_nlp.author_profilinghe corpus for the author profiling analysis contains texts in Russian-language which labeled for 5 tasks:
1) gender -- 13530 texts with the labels, who wrote this: text female or male;
2) age -- 13530 texts with the labels, how old the person who wrote the text. This is a number from 12 to 80. In addition, for the classification task we added 5 age groups: 1-19; 20-29; 30-39; 40-49; 50+;
3) age imitation -- 7574 texts, where crowdsource authors is asked to write three texts:
a) in their natural manner,
b) imitating the style of someone younger,
c) imitating the style of someone older;
4) gender imitation -- 5956 texts, where the crowdsource authors is asked to write texts: in their origin gender and pretending to be the opposite gender;
5) style imitation -- 5956 texts, where crowdsource authors is asked to write a text on behalf of another person of your own gender, with a distortion of the authors usual style.
