datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
explore-persona-space-dataTaskCraft
Dataset Card for TaskCraft
TaskCraft is a multi-modal benchmark dataset featuring tasks ranging from simple (1-step) to expert-level (4-step+). It contains over 40,000 meticulously curated task instances designed to advance research in:
Agent-based task processing
Tool invocation systems
Multi-step reasoning
Dataset Details
Tool Utilization
Tool Category
Instances
PDF Processor
13,400+
HTML Parser
19,200+
Image Analyzer
8,100+… See the full description on the dataset page: https://huggingface.co/datasets/PersonalAILab/TaskCraft.PersonaMem-v2
PersonaMem-v2: Towards Personalized Intelligence via Learning Implicit User Personas and Agentic Memory
📅 We have now released PersonaMem-v3!
🚨 The paper is now released. View the full paper here and codebase here.
Personalization is becoming the next milestone of artificial super-intelligence. AI cannot always satisfy every user, especially on tasks with subjective goals, but personalization offers a path toward pluralistic alignment.… See the full description on the dataset page: https://huggingface.co/datasets/bowen-upenn/PersonaMem-v2.tulu-3-sft-personas-instruction-following
Dataset Descriptions
This dataset contains 29980 examples and is synthetically created to enhance model's capabilities to follow instructions precisely and to satisfy user constraints. The constraints are borrowed from the taxonomy in IFEval dataset.
To generate diverse instructions, we expand the methodology in Ge et al., 2024 by using personas. More details and exact prompts used to construct the dataset can be found in our paper.
Curated by: Allen Institute for AI
Paper: TBD… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-3-sft-personas-instruction-following.Person2Drive
Person2Drive
Person2Drive is the dataset and benchmark repository for the ECCV 2026 paper:
Driving like yourself: A Benchmark for Closed-Loop Personalized End-to-End Autonomous Driving
Overview
Person2Drive is a CARLA-based benchmark for studying personalized end-to-end autonomous driving. It contains human driving records collected in closed-loop simulation environments and is organized at the driver level.
The goal of the benchmark is to support research on… See the full description on the dataset page: https://huggingface.co/datasets/dongxr7/Person2Drive.PersonaHub
Scaling Synthetic Data Creation with 1,000,000,000 Personas
This repo releases data introduced in our paper Scaling Synthetic Data Creation with 1,000,000,000 Personas:
We propose a novel persona-driven data synthesis methodology that leverages various perspectives within a large language model (LLM) to create diverse synthetic data. To fully exploit this methodology at scale, we introduce PERSONA HUB – a collection of 1 billion diverse personas automatically curated from web… See the full description on the dataset page: https://huggingface.co/datasets/proj-persona/PersonaHub.Nemotron-Personas-Korea
Nemotron-Personas-Korea
우리나라 실제 분포에 기반한 합성 페르소나를 위한 복합 AI 시스템
A compound AI approach to personas grounded in real-world distributions
데이터셋 개요 (Overview)
Nemotron-Personas-Korea는 대한민국의 실제 인구통계학적·지리적·성격 특성 분포를 기반으로 합성된 오픈소스 페르소나 데이터셋(CC BY 4.0)으로, 우리나라 인구의 다양성과 특성을 폭넓게 반영하도록 설계되었습니다. 이는 최초의 대규모 우리말 페르소나 데이터셋이며, 이름, 성별, 나이, 혼인 상태, 교육 수준, 직업, 거주 지역 등의 속성을 실제 대한민국 국가데이터처 국가통계포털(KOSIS), 대법원, 국민건강보험공단, 농촌경제연구원, NAVER Cloud 통계 자료를 기반으로 합성하였습니다.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-Korea.personachat_truecasedA version of the PersonaChat dataset that has been true-cased, and also has been given more normalized punctuation.
The original PersonaChat dataset is in all lower case, and has extra space around each clause/sentence separating
punctuation mark. This version of the dataset has more of a natural language look, with sentence capitalization,
proper noun capitalization, and normalized whitespace. Also, each dialogue turn includes a pool of distractor
candidate responses, which can be used by a multiple choice regularization loss during training.monorepo
Persona Cartography — artifact monorepo
Artifact store for the paper Persona Cartography: Charting Language Model
Personality Traits in Weight
Space (arXiv:2607.07916). Code:
persona-cartography/persona-cartography.
This is not a load_dataset-able dataset — it is a single shared repo
holding every artifact the paper's pipeline produces: trained LoRA adapters,
their training data, evaluation results, and the figures' source data. The
paper's figure scripts hydrate from the paths… See the full description on the dataset page: https://huggingface.co/datasets/persona-cartography/monorepo.Synthetic-Persona-Chat
Dataset Card for SPC: Synthetic-Persona-Chat Dataset
Abstract from the paper introducing this dataset:
High-quality conversational datasets are essential for developing AI models that can communicate with users. One way to foster deeper interactions between a chatbot and its user is through personas, aspects of the user's character that provide insights into their personality, motivations, and behaviors. Training Natural Language Processing (NLP) models on a diverse and… See the full description on the dataset page: https://huggingface.co/datasets/google/Synthetic-Persona-Chat.MatrAIx_Persona_1M
MatrAIx Persona 1M
999,847 personas, each described by 1,290 categorical attributes.
599,847 are derived from real records, 400,000 are synthetic.
10 Zstandard Parquet shards, 4.17 GB.
Read it with pyarrow, not datasets
Attributes are packed: one persona's 1,290 attributes are 645 bytes of 4-bit
codes, low nibble first. datasets cannot open these files at all. Use pyarrow
and decode against persona_codes.schema.json.
import json, pyarrow.parquet as pq
schema =… See the full description on the dataset page: https://huggingface.co/datasets/MatrAIx2026/MatrAIx_Persona_1M.personal-facts-msc
Personal Facts (MSC) — Multi-Dimensional Annotation
A manually annotated dataset of 2,779 personal facts sampled from the
Multi-Session Chat (MSC)
corpus, labeled across seven dimensions that jointly characterize a fact's
topic, temporal anchoring, referent, lifetime, validity, and dialogue-continuation
potential.
The scheme extends PeaCoK with
two new top-level categories (Demographics, Possessions) and three new
dimensions (Duration, Validity / Invalidity Reason, Followup),
and… See the full description on the dataset page: https://huggingface.co/datasets/adugeen/personal-facts-msc.PERSONA
Dataset Card for PERSONAS (Prism Filter)
PERSONAS (Prism filter) is one of the largest datasets of synthetic preferences, with over 200k preferences over thousands of questions and 1k personas.
Details on the PERSONAS dataset can be found here paper link.
Note that you MUST also fill out the form on our site to receive access to the full dataset. The form is available here.
Dataset Details
Dataset Description
The personas dataset is a pluralistic… See the full description on the dataset page: https://huggingface.co/datasets/SynthLabsAI/PERSONA.PersonaMem-v1🚨 We have now released PersonaMem-v3 and PersonaMem-v2.
This is the official Huggingface repository of the paper Know Me, Respond to Me: Benchmarking LLMs for Dynamic User Profiling and Personalized Responses at Scale and the PersonaMem benchmark.
We present PersonaMem, a new LLM personalization benchmark to assess how well language models can infer evolving user profiles and generate personalized responses across task scenarios. PersonaMem emphasizes persona-oriented, multi-session… See the full description on the dataset page: https://huggingface.co/datasets/bowen-upenn/PersonaMem-v1.Nemotron-Personas-Japan
Nemotron-Personas-Japan
現実世界の分布に基づいたペルソナ生成のための複合AIアプローチ
データセット概要 (Dataset Overview)
Nemotron-Personas-Japan は、日本における人口の多様性と豊かさを捉えることを目的とし、実世界の人口統計、地理的分布、性格特性の分布に基づいて合成的に生成されたペルソナのオープンソースデータセットです。名前、性別、年齢、背景、婚姻状況、学歴、職業、居住地などの統計に基づいて生成した初のデータセットされた Nemotron-Personas の日本語版です。本バージョンでは、日本語における多様なモデリングユースケースに適した高品質のペルソナを提供します
Nemotron-Personas-Japan は、日本のモデル開発者が重要な地域固有の人口統計や文化的背景を取り入れたソブリンAIシステムを開発することを支援します。本データセットは、日本の地理的・人口統計的な実分布を反映することで、合成データの多様性を高め、バイアスを軽減し、model… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-Japan.Nemotron-Personas-USA
Nemotron-Personas-USA
A compound AI approach to personas grounded in real-world distributions
v1.1 Update
The v1.1 update introduces the following changes:
leverage openai/gpt-oss-120b model instead of mistralai/Mixtral-8x22B-v0.1 model to improve data quality and diversity
increase the number of records from 100k to 1M, for a total of 0.94B tokens
update the dataset name to Nemotron-Personas-USA in order to differentiate it from other region-specific datasets… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-USA.PERSONA_subset
Dataset Card for PERSONAS (Prism Filter)
PERSONAS (Prism filter) is one of the largest datasets of synthetic preferences, with over 200k preferences over thousands of questions and 1k personas.
Details on the PERSONAS dataset can be found here paper link
Note that this subset is 5% of the training split of PERSONAS. The full dataset is here, strictly available for academic use.
You MUST request access to the full persona dataset here.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/SynthLabsAI/PERSONA_subset.PersonaMem-v3
PersonaMem-v3: Toward Omni-Platform Personal Intelligence for Holistic User Understanding, Recommendation, and Agentic Tasks
Bowen Jiang, Yuan Yuan, Zhuoqun Hao, Yuchen Liu, Maohao Shen, Sihao Chen, Gregory Wornell,
Chris Callison-Burch, Lyle Ungar, Dan Roth, Qi Guo, Xiangjun Fan, Camillo J. Taylor, Hanchao Yu
A collaboration between:
Meta Recommendation Systems
University of Pennsylvania
MIT
Third release in the PersonaMem series:
PersonaMem-v1: [COLM… See the full description on the dataset page: https://huggingface.co/datasets/bowen-upenn/PersonaMem-v3.persona-chatmc_third_personelite-personas-embeddingspersona_in_palNemotron-Personas-India
Nemotron-Personas-India
A compound AI approach to personas grounded in real-world distributions
वास्तविक दुनिया के वितरण पर आधारित व्यक्तित्वों के लिए एक मिश्रित AI दृष्टिकोण
Dataset Overview (डेटासेट अवलोकन)
Nemotron-Personas-India is an open-source (CC BY 4.0) dataset of synthetically-generated personas. This dataset is grounded in real-world demographic, geographic and personality trait distributions in India to capture the diversity and richness of… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-India.mbti-personality-datasetAI2_Alphabot_2_sort_personal_care_items
AI2_Alphabot_2_sort_personal_care_items
Dataset Description
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Task Preview
View Video Directly
Overview
Total Episodes: 496
Total Frames: 455317
FPS: 30
Dataset Size: 13.57 GB
Robot Name: AI2_Alphabot_2
End-Effector Type: two_finger_end_effector
Teleoperation Type: vr_controller
Sensors: cam_front_chest_rgb,
cam_front_head_rgb… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AI2_Alphabot_2_sort_personal_care_items.tulu-3-sft-personas-math
A filtered version of this dataset is available here: https://huggingface.co/datasets/allenai/tulu-3-sft-personas-math-filtered
Dataset Descriptions
This dataset contains 149960 examples and is synthetically created to enhance model's capabilities to answer complex and hard math word problems.
To generate diverse math questions, we expand the methodology in Ge et al., 2024 by using personas. More details and exact prompts used to construct the dataset can be found in our paper.… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-3-sft-personas-math.tulu-3-sft-personas-code
Dataset Descriptions
This dataset contains 34999 examples and is synthetically created to enhance models' coding capabilities.To generate diverse python coding questions, we expand the methodology in Ge et al., 2024 by using personas to ground the code completion question in real-world scenarios. More details and exact prompts used to construct the dataset can be found in our paper.
Curated by: Allen Institute for AI
Paper: TBD
Repository: TBD
Language(s) (NLP): English
License:… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-3-sft-personas-code.personal-trainer-ausbildung-ki-datensatz
SNFA Personal Trainer Ausbildung KI-Datensatz
Ein deutschsprachiger Wissensdatensatz der SNF Academy zu Personal Training, Fitnessausbildung, Berufspraxis, Coaching, Selbstständigkeit und regionalen Angeboten in der Schweiz.
Inhalt
Die Datei snfa_personal_trainer_dataset.jsonl enthält thematisch abgegrenzte Abschnitte aus den Dokumenten dieses Repositorys. Jeder Datensatz besitzt eine eindeutige ID sowie Angaben zu Titel, Abschnitt, Inhalt, Kategorie, Quelldatei… See the full description on the dataset page: https://huggingface.co/datasets/snfacademy/personal-trainer-ausbildung-ki-datensatz.persona-curvature-oct-transcripts
Content warning
These are synthetic transcripts generated by a language model talking to itself
under an instruction to embody a personality trait. Several traits produce
distressing material. It is published deliberately rather than filtered out,
because the rate at which a trait produces it is one of the findings.
Across all 134 traits, an automated scan flagged 708 rows in 50 files. Two traits
account for 89% of them, and they are not the two you would guess:
trait… See the full description on the dataset page: https://huggingface.co/datasets/EternalRecursion/persona-curvature-oct-transcripts.persona-curvature-results
Persona Curvature: results, analysis outputs and run provenance
The non-weight data behind the persona-curvature project: the Gram matrices,
factor analyses, activation-space means, steering results, judged generations,
training corpora and run logs that the analysis scripts, figures, companion site
and wiki in the GitHub repository read.
Code, figures, wiki, docs: https://github.com/EternalRecursion121/persona-curvature
The adapters themselves (134 stage-one DPO, 119 stage-two… See the full description on the dataset page: https://huggingface.co/datasets/EternalRecursion/persona-curvature-results.
