datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mbti-personality-datasetbig-five-personality-traits
Big Five Personality Traits Dataset
This dataset contains AI-generated descriptions of personality traits based on the Big Five (OCEAN) model. For each trait and intensity level (1–5), five descriptions were produced by ten different chatbots: Grok, Gemini, Claude, KimiK2 (via HuggingChat), Deepseek, MetaAI, Perplexity, LeChat, ChatGPT, and Copilot.
Overview
The dataset can support tasks such as persona creation, comparative language analysis, and research on how AI… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/big-five-personality-traits.mbti-personality-datasetAutomated-Personality-PredictionSource:
The dataset is titled PANDORA and is retrieved from the https://psy.takelab.fer.hr/datasets/all/pandora/. the PANDORA dataset is the only dataset that contains personality-relevant information for multiple personality models. It consists of Reddit comments with their corresponding scores for the Big Five Traits, MBTI values and the Enneagrams for more than 10k users.
This Dataset:
This dataset is a subset of Reddit comments from PANDORA focused only on the Big Five Traits. The… See the full description on the dataset page: https://huggingface.co/datasets/Fatima0923/Automated-Personality-Prediction.Personality_datasetDataset for personality manipulation of LLMs
PersonalityDetectionpersonalaity-llm-personality-profiles
PersonalAIty: HEXACO personality profiles of frontier LLMs
Self-reported HEXACO personality profiles for 10 frontier language models across 8 vendors,
measured on 2026-08-16 with an open 50-item inventory, plus the instrument itself so the
measurement can be rerun or criticised.
This is a snapshot with a date on it, not a standing benchmark. Model versions drift; the
value here is that the whole measurement is reproducible with one command against models anyone
can reach.… See the full description on the dataset page: https://huggingface.co/datasets/Sciupy/personalaity-llm-personality-profiles.personality_manipulationfacebook-personality-recognition-wcpr13The Workshop on Computational Personality Recognition 2013 was a competition based on this Facebook dataset.
The purpose is to predict the personality scores or classes from text and ego-network data
reference paper: https://ojs.aaai.org/index.php/ICWSM/article/view/14467/14316
big-five-personality-traits
Big Five Personality Traits Dataset
This dataset contains AI-generated descriptions of personality traits based on the Big Five (OCEAN) model. For each trait and intensity level (1–5), five descriptions were produced by ten different chatbots: Grok, Gemini, Claude, KimiK2 (via HuggingChat), Deepseek, MetaAI, Perplexity, LeChat, ChatGPT, and Copilot.
Overview
The dataset can support tasks such as persona creation, comparative language analysis, and research on how AI… See the full description on the dataset page: https://huggingface.co/datasets/sreelekshmisajuk/big-five-personality-traits.personalityyoutube-vlog-personality-recognition-wcpr14The Workshop on Personality Recognition 2014 was a competition based on this dataset. The goal is to predict personality scores from visual features and text transcripts.
Reference Paper: https://infoscience.epfl.ch/server/api/core/bitstreams/e61b4c1b-0c56-4afc-9786-23e9841cb81f/content
Automated-Personality-PredictionSource:
The dataset is titled PANDORA and is retrieved from the https://psy.takelab.fer.hr/datasets/all/pandora/. the PANDORA dataset is the only dataset that contains personality-relevant information for multiple personality models. It consists of Reddit comments with their corresponding scores for the Big Five Traits, MBTI values and the Enneagrams for more than 10k users.
This Dataset:
This dataset is a subset of Reddit comments from PANDORA focused only on the Big Five Traits. The… See the full description on the dataset page: https://huggingface.co/datasets/raveinid/Automated-Personality-Prediction.Automated-Personality-PredictionSource:
The dataset is titled PANDORA and is retrieved from the https://psy.takelab.fer.hr/datasets/all/pandora/. the PANDORA dataset is the only dataset that contains personality-relevant information for multiple personality models. It consists of Reddit comments with their corresponding scores for the Big Five Traits, MBTI values and the Enneagrams for more than 10k users.
This Dataset:
This dataset is a subset of Reddit comments from PANDORA focused only on the Big Five Traits. The… See the full description on the dataset page: https://huggingface.co/datasets/marchmallow/Automated-Personality-Prediction.personalitycustomer_personality_analysisPersonality_dataset_testPersonality_dataset_trainInstagram-Profiles-OCEAN-Personality-traitsAIthical-personalityMy merged unethical dataset seeded to be ethical. reponses were generated using the gemini api and system prompt was made to get the model to refuse with "personality"
formatted_personalitycustomer_personality_analysispersonality-behavioral-data
Personality Behavioral Data
The raw behavioral layer of the identity_framing_llm experiment: every
Qwen2.5-7B-Instruct rollout, its judge score, and the scenarios that prompted
it. Everything else in the project (concept vectors, state coordinates,
structure matrices) is derived from this data. The rollouts are stochastic
samples and do not regenerate identically, which is why they are archived here
rather than treated as reproducible.
Files
file
rows… See the full description on the dataset page: https://huggingface.co/datasets/linkpipi/personality-behavioral-data.
