datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
phishing-llm-bias-audit
LLM Phishing-Vulnerability Bias Audit Dataset
A multi-provider empirical dataset capturing how 14 open-source LLM configurations (across 5 inference providers) select which of three generated personas is "most vulnerable to phishing." 855 persona records / 285 forced-choice workflows.
Important. This dataset is about LLM behaviour under controlled prompts, not about real-world phishing susceptibility of any demographic group. Selecting a persona as "vulnerable" is the LLM's choice;… See the full description on the dataset page: https://huggingface.co/datasets/Julia569922/phishing-llm-bias-audit.llm-nationality-bias-global-narratives
Representational Harms in Global LLM Narratives: Nationality Bias Dataset
Dataset Summary
This dataset contains 292,500 LLM-generated narratives across 195 globally-recognized nations, created to examine how national identity cues in prompts shape narrative content and representation. Generated using GPT-4.1 Nano, the dataset systematically varies the nationality of dominant characters across power-laden scenarios in Learning, Labor, and Love domains. This enables… See the full description on the dataset page: https://huggingface.co/datasets/ilana27/llm-nationality-bias-global-narratives.llm-nationality-bias-us-narratives
Representational Harms in US-Based LLM Narratives: Nationality Bias Dataset
Dataset Summary
This dataset contains 9,710 LLM-generated narratives that reference non-US national identities, extracted from a larger corpus of 500,000 stories generated by GPT-3.5, GPT-4, Claude 2.0, Llama 2, and PaLM 2. The stories were generated in response to open-ended prompts set in US contexts across three domains: Learning, Labor, and Love. This dataset was created to study… See the full description on the dataset page: https://huggingface.co/datasets/ilana27/llm-nationality-bias-us-narratives.
