datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PersonaMem-v2
PersonaMem-v2: Towards Personalized Intelligence via Learning Implicit User Personas and Agentic Memory
📅 We have now released PersonaMem-v3!
🚨 The paper is now released. View the full paper here and codebase here.
Personalization is becoming the next milestone of artificial super-intelligence. AI cannot always satisfy every user, especially on tasks with subjective goals, but personalization offers a path toward pluralistic alignment.… See the full description on the dataset page: https://huggingface.co/datasets/bowen-upenn/PersonaMem-v2.PersonaMem-v1🚨 We have now released PersonaMem-v3 and PersonaMem-v2.
This is the official Huggingface repository of the paper Know Me, Respond to Me: Benchmarking LLMs for Dynamic User Profiling and Personalized Responses at Scale and the PersonaMem benchmark.
We present PersonaMem, a new LLM personalization benchmark to assess how well language models can infer evolving user profiles and generate personalized responses across task scenarios. PersonaMem emphasizes persona-oriented, multi-session… See the full description on the dataset page: https://huggingface.co/datasets/bowen-upenn/PersonaMem-v1.PersonaMem-v3
PersonaMem-v3: Toward Omni-Platform Personal Intelligence for Holistic User Understanding, Recommendation, and Agentic Tasks
Bowen Jiang, Yuan Yuan, Zhuoqun Hao, Yuchen Liu, Maohao Shen, Sihao Chen, Gregory Wornell,
Chris Callison-Burch, Lyle Ungar, Dan Roth, Qi Guo, Xiangjun Fan, Camillo J. Taylor, Hanchao Yu
A collaboration between:
Meta Recommendation Systems
University of Pennsylvania
MIT
Third release in the PersonaMem series:
PersonaMem-v1: [COLM… See the full description on the dataset page: https://huggingface.co/datasets/bowen-upenn/PersonaMem-v3.Reverse-alpha-beta-no-outsideReverse-hybrid-shared-no-persona-remainderReverse-no-persona-replacement-remainderReverse-baseline-bias-unbiasReverse-hybrid-train-no-persona-meanPersonaRoute-Bench
📖Project Resources
This repository contains the datasets presented in the paper PersonalizedRouter
For more details about PersonalizedRouter, please visit our GitHub repository
📊Dataset Overview
In the project files, the suffix v1 refers to the Multi-cost-efficiency Simulation Strategy described in the paper, while v2 refers to the LLM-as-a-Judge Simulation, and large denotes the large-scale setting.
You can utilize router_user_data_v1 (or v2) to train and test… See the full description on the dataset page: https://huggingface.co/datasets/ulab-ai/PersonaRoute-Bench.Reverse-alpha-suppression-task-boostOriginal-alpha-suppression-task-boostReverse-circuit-discoveryReverse-hybrid-correct-train-no-persona-meanOriginal-no-persona-replacement-remainderOriginal-hybrid-shared-no-persona-remainderPersona-E2-Dataset
Persona-E²: A Human-Grounded Dataset for Personality-Shaped Emotional Responses to Textual Events
Feel Free to Contact Us: ftyuqin_yang@mail.scut.edu.cn
1. Dataset Summary
Two individuals can construe their situations quite similarly (agree on all the facts), and yet react with very different emotions, because they have appraised the adaptational significance of those facts differently. — Lazarus, Emotion And Adaptation (1991)
Most emotion recognition… See the full description on the dataset page: https://huggingface.co/datasets/CRIS-Yang/Persona-E2-Dataset.PersonaAsInfrastructure
Supplementary Material
Persona as Infrastructure: Invisible Structural Control in LLM-Mediated Social Networks
ICNLSP 2026.
This archive contains the complete data and code needed to reproduce every
number, table, and figure in the paper.
1. Contents
Networks (networks/)
109 generated networks in GraphML. Each file is one persona × one trial:
network_<persona>_trial<N>.graphml, with companion
_stats.json (summary metrics, including the exact model… See the full description on the dataset page: https://huggingface.co/datasets/amircincy/PersonaAsInfrastructure.Original-hybrid-correct-train-no-persona-meanOriginal-hybrid-train-no-persona-meanOriginal-circuit-discoveryalpha_beta_interventionvn-provinces-employed-persons
Vietnam provinces employed persons in the economy
Provincial and regional number of employed persons in the economy (thousand persons). Coverage 2018-2024. Year 2024 is preliminary. Tables cover provinces, regions and national total. Geographic labels are English (UN/GSO style ASCII romanization). Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Comparison
Color key
Files
provinces (441 rows)
data/provinces.csv… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-employed-persons.PersonaMem-v2
PersonaMem-v2: Towards Personalized Intelligence via Learning Implicit User Personas and Agentic Memory
🚨 The paper is now released. View the full paper here and codebase here.
🙌 The dataset has been downloaded over 12,000 times. Thank you everybody for finding our work helpful!
Personalization is becoming the next milestone of artificial super-intelligence. AI cannot always satisfy every user, especially on tasks with subjective goals, but personalization… See the full description on the dataset page: https://huggingface.co/datasets/milanow/PersonaMem-v2.Original-baseline-bias-unbiasAutomated-Personality-PredictionSource:
The dataset is titled PANDORA and is retrieved from the https://psy.takelab.fer.hr/datasets/all/pandora/. the PANDORA dataset is the only dataset that contains personality-relevant information for multiple personality models. It consists of Reddit comments with their corresponding scores for the Big Five Traits, MBTI values and the Enneagrams for more than 10k users.
This Dataset:
This dataset is a subset of Reddit comments from PANDORA focused only on the Big Five Traits. The… See the full description on the dataset page: https://huggingface.co/datasets/Fatima0923/Automated-Personality-Prediction.headset_faithfulnessNemotron-Personas-Marketing
Marketing Personas (Filtered from Nemotron-Personas)
Dataset Card for Marketing Personas (Filtered from Nemotron-Personas)
This dataset contains only Marketing-related personas filtered from the originalnvidia/Nemotron-Personas dataset.
The filtering was done using keyword-based search for terms likemarketing, advertising, branding, campaign, social media, promotion, sales, content strategy, digital marketing, etc.
Dataset Details
This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/marketeam/Nemotron-Personas-Marketing.persona-and-other-evals
Qwen3.5-9B AMA adapters — persona evals
Inference code, the data it produced, and the tools that turn that data
into tables and an HTML viewer. The evals are Anthropic's persona set,
scored in three regimes: teacher-forced logprob of the answer literal,
greedy answer with the reasoning block pre-closed, and a full 16k-budget
reasoning trace.
Pinned models
base unsloth/Qwen3.5-9B @ 005429cee5cb648998cf2b70eebdd83175989c9a
util… See the full description on the dataset page: https://huggingface.co/datasets/agentic-moral-alignment/persona-and-other-evals.PersonaMem11🚨 We invite everyone to checkout our PersonaMem-v2 on 🤗HuggingFace, focusing on realistic and implicit user preferences in long conversations!
This is the official Huggingface repository of the paper Know Me, Respond to Me: Benchmarking LLMs for Dynamic User Profiling and Personalized Responses at Scale and the PersonaMem benchmark.
We present PersonaMem, a new LLM personalization benchmark to assess how well language models can infer evolving user profiles and generate personalized… See the full description on the dataset page: https://huggingface.co/datasets/jiebian/PersonaMem11.personalaity-llm-personality-profiles
PersonalAIty: HEXACO personality profiles of frontier LLMs
Self-reported HEXACO personality profiles for 10 frontier language models across 8 vendors,
measured on 2026-08-16 with an open 50-item inventory, plus the instrument itself so the
measurement can be rerun or criticised.
This is a snapshot with a date on it, not a standing benchmark. Model versions drift; the
value here is that the whole measurement is reproducible with one command against models anyone
can reach.… See the full description on the dataset page: https://huggingface.co/datasets/Sciupy/personalaity-llm-personality-profiles.
