datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PersonaMem-v2
PersonaMem-v2: Towards Personalized Intelligence via Learning Implicit User Personas and Agentic Memory
📅 We have now released PersonaMem-v3!
🚨 The paper is now released. View the full paper here and codebase here.
Personalization is becoming the next milestone of artificial super-intelligence. AI cannot always satisfy every user, especially on tasks with subjective goals, but personalization offers a path toward pluralistic alignment.… See the full description on the dataset page: https://huggingface.co/datasets/bowen-upenn/PersonaMem-v2.tulu-3-sft-personas-instruction-following
Dataset Descriptions
This dataset contains 29980 examples and is synthetically created to enhance model's capabilities to follow instructions precisely and to satisfy user constraints. The constraints are borrowed from the taxonomy in IFEval dataset.
To generate diverse instructions, we expand the methodology in Ge et al., 2024 by using personas. More details and exact prompts used to construct the dataset can be found in our paper.
Curated by: Allen Institute for AI
Paper: TBD… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-3-sft-personas-instruction-following.PersonaHub
Scaling Synthetic Data Creation with 1,000,000,000 Personas
This repo releases data introduced in our paper Scaling Synthetic Data Creation with 1,000,000,000 Personas:
We propose a novel persona-driven data synthesis methodology that leverages various perspectives within a large language model (LLM) to create diverse synthetic data. To fully exploit this methodology at scale, we introduce PERSONA HUB – a collection of 1 billion diverse personas automatically curated from web… See the full description on the dataset page: https://huggingface.co/datasets/proj-persona/PersonaHub.Nemotron-Personas-Korea
Nemotron-Personas-Korea
우리나라 실제 분포에 기반한 합성 페르소나를 위한 복합 AI 시스템
A compound AI approach to personas grounded in real-world distributions
데이터셋 개요 (Overview)
Nemotron-Personas-Korea는 대한민국의 실제 인구통계학적·지리적·성격 특성 분포를 기반으로 합성된 오픈소스 페르소나 데이터셋(CC BY 4.0)으로, 우리나라 인구의 다양성과 특성을 폭넓게 반영하도록 설계되었습니다. 이는 최초의 대규모 우리말 페르소나 데이터셋이며, 이름, 성별, 나이, 혼인 상태, 교육 수준, 직업, 거주 지역 등의 속성을 실제 대한민국 국가데이터처 국가통계포털(KOSIS), 대법원, 국민건강보험공단, 농촌경제연구원, NAVER Cloud 통계 자료를 기반으로 합성하였습니다.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-Korea.personachat_truecasedA version of the PersonaChat dataset that has been true-cased, and also has been given more normalized punctuation.
The original PersonaChat dataset is in all lower case, and has extra space around each clause/sentence separating
punctuation mark. This version of the dataset has more of a natural language look, with sentence capitalization,
proper noun capitalization, and normalized whitespace. Also, each dialogue turn includes a pool of distractor
candidate responses, which can be used by a multiple choice regularization loss during training.Synthetic-Persona-Chat
Dataset Card for SPC: Synthetic-Persona-Chat Dataset
Abstract from the paper introducing this dataset:
High-quality conversational datasets are essential for developing AI models that can communicate with users. One way to foster deeper interactions between a chatbot and its user is through personas, aspects of the user's character that provide insights into their personality, motivations, and behaviors. Training Natural Language Processing (NLP) models on a diverse and… See the full description on the dataset page: https://huggingface.co/datasets/google/Synthetic-Persona-Chat.MatrAIx_Persona_1M
MatrAIx Persona 1M
999,847 personas, each described by 1,290 categorical attributes.
599,847 are derived from real records, 400,000 are synthetic.
10 Zstandard Parquet shards, 4.17 GB.
Read it with pyarrow, not datasets
Attributes are packed: one persona's 1,290 attributes are 645 bytes of 4-bit
codes, low nibble first. datasets cannot open these files at all. Use pyarrow
and decode against persona_codes.schema.json.
import json, pyarrow.parquet as pq
schema =… See the full description on the dataset page: https://huggingface.co/datasets/MatrAIx2026/MatrAIx_Persona_1M.personal-facts-msc
Personal Facts (MSC) — Multi-Dimensional Annotation
A manually annotated dataset of 2,779 personal facts sampled from the
Multi-Session Chat (MSC)
corpus, labeled across seven dimensions that jointly characterize a fact's
topic, temporal anchoring, referent, lifetime, validity, and dialogue-continuation
potential.
The scheme extends PeaCoK with
two new top-level categories (Demographics, Possessions) and three new
dimensions (Duration, Validity / Invalidity Reason, Followup),
and… See the full description on the dataset page: https://huggingface.co/datasets/adugeen/personal-facts-msc.PERSONA
Dataset Card for PERSONAS (Prism Filter)
PERSONAS (Prism filter) is one of the largest datasets of synthetic preferences, with over 200k preferences over thousands of questions and 1k personas.
Details on the PERSONAS dataset can be found here paper link.
Note that you MUST also fill out the form on our site to receive access to the full dataset. The form is available here.
Dataset Details
Dataset Description
The personas dataset is a pluralistic… See the full description on the dataset page: https://huggingface.co/datasets/SynthLabsAI/PERSONA.PersonaMem-v1🚨 We have now released PersonaMem-v3 and PersonaMem-v2.
This is the official Huggingface repository of the paper Know Me, Respond to Me: Benchmarking LLMs for Dynamic User Profiling and Personalized Responses at Scale and the PersonaMem benchmark.
We present PersonaMem, a new LLM personalization benchmark to assess how well language models can infer evolving user profiles and generate personalized responses across task scenarios. PersonaMem emphasizes persona-oriented, multi-session… See the full description on the dataset page: https://huggingface.co/datasets/bowen-upenn/PersonaMem-v1.Nemotron-Personas-Japan
Nemotron-Personas-Japan
現実世界の分布に基づいたペルソナ生成のための複合AIアプローチ
データセット概要 (Dataset Overview)
Nemotron-Personas-Japan は、日本における人口の多様性と豊かさを捉えることを目的とし、実世界の人口統計、地理的分布、性格特性の分布に基づいて合成的に生成されたペルソナのオープンソースデータセットです。名前、性別、年齢、背景、婚姻状況、学歴、職業、居住地などの統計に基づいて生成した初のデータセットされた Nemotron-Personas の日本語版です。本バージョンでは、日本語における多様なモデリングユースケースに適した高品質のペルソナを提供します
Nemotron-Personas-Japan は、日本のモデル開発者が重要な地域固有の人口統計や文化的背景を取り入れたソブリンAIシステムを開発することを支援します。本データセットは、日本の地理的・人口統計的な実分布を反映することで、合成データの多様性を高め、バイアスを軽減し、model… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-Japan.Nemotron-Personas-USA
Nemotron-Personas-USA
A compound AI approach to personas grounded in real-world distributions
v1.1 Update
The v1.1 update introduces the following changes:
leverage openai/gpt-oss-120b model instead of mistralai/Mixtral-8x22B-v0.1 model to improve data quality and diversity
increase the number of records from 100k to 1M, for a total of 0.94B tokens
update the dataset name to Nemotron-Personas-USA in order to differentiate it from other region-specific datasets… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-USA.PERSONA_subset
Dataset Card for PERSONAS (Prism Filter)
PERSONAS (Prism filter) is one of the largest datasets of synthetic preferences, with over 200k preferences over thousands of questions and 1k personas.
Details on the PERSONAS dataset can be found here paper link
Note that this subset is 5% of the training split of PERSONAS. The full dataset is here, strictly available for academic use.
You MUST request access to the full persona dataset here.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/SynthLabsAI/PERSONA_subset.PersonaMem-v3
PersonaMem-v3: Toward Omni-Platform Personal Intelligence for Holistic User Understanding, Recommendation, and Agentic Tasks
Bowen Jiang, Yuan Yuan, Zhuoqun Hao, Yuchen Liu, Maohao Shen, Sihao Chen, Gregory Wornell,
Chris Callison-Burch, Lyle Ungar, Dan Roth, Qi Guo, Xiangjun Fan, Camillo J. Taylor, Hanchao Yu
A collaboration between:
Meta Recommendation Systems
University of Pennsylvania
MIT
Third release in the PersonaMem series:
PersonaMem-v1: [COLM… See the full description on the dataset page: https://huggingface.co/datasets/bowen-upenn/PersonaMem-v3.persona-chatelite-personas-embeddingspersona_in_palNemotron-Personas-India
Nemotron-Personas-India
A compound AI approach to personas grounded in real-world distributions
वास्तविक दुनिया के वितरण पर आधारित व्यक्तित्वों के लिए एक मिश्रित AI दृष्टिकोण
Dataset Overview (डेटासेट अवलोकन)
Nemotron-Personas-India is an open-source (CC BY 4.0) dataset of synthetically-generated personas. This dataset is grounded in real-world demographic, geographic and personality trait distributions in India to capture the diversity and richness of… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-India.mbti-personality-datasettulu-3-sft-personas-math
A filtered version of this dataset is available here: https://huggingface.co/datasets/allenai/tulu-3-sft-personas-math-filtered
Dataset Descriptions
This dataset contains 149960 examples and is synthetically created to enhance model's capabilities to answer complex and hard math word problems.
To generate diverse math questions, we expand the methodology in Ge et al., 2024 by using personas. More details and exact prompts used to construct the dataset can be found in our paper.… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-3-sft-personas-math.tulu-3-sft-personas-code
Dataset Descriptions
This dataset contains 34999 examples and is synthetically created to enhance models' coding capabilities.To generate diverse python coding questions, we expand the methodology in Ge et al., 2024 by using personas to ground the code completion question in real-world scenarios. More details and exact prompts used to construct the dataset can be found in our paper.
Curated by: Allen Institute for AI
Paper: TBD
Repository: TBD
Language(s) (NLP): English
License:… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-3-sft-personas-code.personal-trainer-ausbildung-ki-datensatz
SNFA Personal Trainer Ausbildung KI-Datensatz
Ein deutschsprachiger Wissensdatensatz der SNF Academy zu Personal Training, Fitnessausbildung, Berufspraxis, Coaching, Selbstständigkeit und regionalen Angeboten in der Schweiz.
Inhalt
Die Datei snfa_personal_trainer_dataset.jsonl enthält thematisch abgegrenzte Abschnitte aus den Dokumenten dieses Repositorys. Jeder Datensatz besitzt eine eindeutige ID sowie Angaben zu Titel, Abschnitt, Inhalt, Kategorie, Quelldatei… See the full description on the dataset page: https://huggingface.co/datasets/snfacademy/personal-trainer-ausbildung-ki-datensatz.Nemotron-Personas-Singapore
Nemotron-Personas-Singapore
A compound AI approach to personas grounded in real-world distributions
Dataset Overview
Nemotron-Personas-Singapore is an open-source (CC BY 4.0) dataset of synthetically-generated personas. This dataset is grounded in real-world demographic, geographic and personality trait distributions in Singapore to capture the diversity and richness of the Singaporean population. It is a variant of Nemotron-Personas-USA, and the first Singaporean… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-Singapore.Nemotron-Personas-Brazil
Nemotron-Personas-Brazil
Abordagem de IA composta para geração de personas baseada em distribuições do mundo real
Visão Geral do Conjunto de Dados (Dataset Overview):
Nemotron-Personas-Brazil é um conjunto de dados (dataset) de código aberto (CC BY 4.0) composto por personas geradas sinteticamente e fundamentadas em distribuições demográficas, geográficas e traços de personalidade reais do Brasil, visando capturar a diversidade e a riqueza da população. Trata-se de… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-Brazil.Nemotron-Personas-France
Nemotron-Personas-France
Une approche d'IA composée pour des personas ancrés dans des distributions réelles
A compound AI approach to personas grounded in real-world distributions
Vue d'ensemble du jeu de données (Dataset Overview)
Nemotron-Personas-France est un jeu de données en libre accès (CC BY 4.0) composé de personas générés de manière synthétique. Ce jeu de données s'appuie sur les distributions démographiques, géographiques et de traits de… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-France.Nemotron-Personas-Vietnam
Nemotron-Personas-Vietnam
Hệ thống AI kết hợp để tạo personas tổng hợp dựa trên phân bố thực tế của Việt Nam
A compound AI approach to personas grounded in real-world distributions
Tổng quan (Overview)
Nemotron-Personas-Vietnam là tập dữ liệu personas được cung cấp dưới dạng mã nguồn mở (CC BY 4.0) dựa trên phân bố nhân khẩu học, địa lý và đặc điểm tính cách của người Việt Nam. Tập dữ liệu phản ánh một cách toàn diện sự phong phú và đặc trưng… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-Vietnam.persona-fas-bench-v1
Filtered copy. Redistribution of saatvikbilla1/persona-fas with a small number of
images withheld and capture metadata (EXIF/GPS, XMP, IPTC) removed.
Attack images in this release: 21,056.
persona
PERSONA: Dynamic and Compositional Inference-Time Personality Control
Official release of persona vectors and SFT datasets for the ICLR 2026 paper:
PERSONA: Dynamic and Compositional Inference-Time Personality Control via Activation Vector Algebra
Xiachong Feng, Liang Zhao, Weihong Zhong, Yichong Huang, Yuxuan Gu, Lingpeng Kong, Xiaocheng Feng, Bing Qin
Harbin Institute of Technology & The University of Hong Kong
Paper: https://openreview.net/pdf?id=QZvGqaNBlU
Code:… See the full description on the dataset page: https://huggingface.co/datasets/xiachongfeng/persona.tulu-3-pref-personas-instruction-following
Dataset Descriptions
This dataset contains 19890 preference examples and is synthetically created to enhance models' precise instruction following capabilities while satisfying several constraints. The dataset containts preference pairs (chosen, reject responses) and can be used for preference tuning methods (e.g., PPO, DPO).
Dataset Construction
To create this dataset, we took a subset of its supervised-tuning version here and convert it into preference dataset.… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-3-pref-personas-instruction-following.PersonalizationV3
