datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PersonaMem-v2
PersonaMem-v2: Towards Personalized Intelligence via Learning Implicit User Personas and Agentic Memory
📅 We have now released PersonaMem-v3!
🚨 The paper is now released. View the full paper here and codebase here.
Personalization is becoming the next milestone of artificial super-intelligence. AI cannot always satisfy every user, especially on tasks with subjective goals, but personalization offers a path toward pluralistic alignment.… See the full description on the dataset page: https://huggingface.co/datasets/bowen-upenn/PersonaMem-v2.tulu-3-sft-personas-instruction-following
Dataset Descriptions
This dataset contains 29980 examples and is synthetically created to enhance model's capabilities to follow instructions precisely and to satisfy user constraints. The constraints are borrowed from the taxonomy in IFEval dataset.
To generate diverse instructions, we expand the methodology in Ge et al., 2024 by using personas. More details and exact prompts used to construct the dataset can be found in our paper.
Curated by: Allen Institute for AI
Paper: TBD… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-3-sft-personas-instruction-following.PersonaHub
Scaling Synthetic Data Creation with 1,000,000,000 Personas
This repo releases data introduced in our paper Scaling Synthetic Data Creation with 1,000,000,000 Personas:
We propose a novel persona-driven data synthesis methodology that leverages various perspectives within a large language model (LLM) to create diverse synthetic data. To fully exploit this methodology at scale, we introduce PERSONA HUB – a collection of 1 billion diverse personas automatically curated from web… See the full description on the dataset page: https://huggingface.co/datasets/proj-persona/PersonaHub.Nemotron-Personas-Korea
Nemotron-Personas-Korea
우리나라 실제 분포에 기반한 합성 페르소나를 위한 복합 AI 시스템
A compound AI approach to personas grounded in real-world distributions
데이터셋 개요 (Overview)
Nemotron-Personas-Korea는 대한민국의 실제 인구통계학적·지리적·성격 특성 분포를 기반으로 합성된 오픈소스 페르소나 데이터셋(CC BY 4.0)으로, 우리나라 인구의 다양성과 특성을 폭넓게 반영하도록 설계되었습니다. 이는 최초의 대규모 우리말 페르소나 데이터셋이며, 이름, 성별, 나이, 혼인 상태, 교육 수준, 직업, 거주 지역 등의 속성을 실제 대한민국 국가데이터처 국가통계포털(KOSIS), 대법원, 국민건강보험공단, 농촌경제연구원, NAVER Cloud 통계 자료를 기반으로 합성하였습니다.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-Korea.MatrAIx_Persona_1M
MatrAIx Persona 1M
999,847 personas, each described by 1,290 categorical attributes.
599,847 are derived from real records, 400,000 are synthetic.
10 Zstandard Parquet shards, 4.17 GB.
Read it with pyarrow, not datasets
Attributes are packed: one persona's 1,290 attributes are 645 bytes of 4-bit
codes, low nibble first. datasets cannot open these files at all. Use pyarrow
and decode against persona_codes.schema.json.
import json, pyarrow.parquet as pq
schema =… See the full description on the dataset page: https://huggingface.co/datasets/MatrAIx2026/MatrAIx_Persona_1M.PersonaMem-v1🚨 We have now released PersonaMem-v3 and PersonaMem-v2.
This is the official Huggingface repository of the paper Know Me, Respond to Me: Benchmarking LLMs for Dynamic User Profiling and Personalized Responses at Scale and the PersonaMem benchmark.
We present PersonaMem, a new LLM personalization benchmark to assess how well language models can infer evolving user profiles and generate personalized responses across task scenarios. PersonaMem emphasizes persona-oriented, multi-session… See the full description on the dataset page: https://huggingface.co/datasets/bowen-upenn/PersonaMem-v1.Nemotron-Personas-Japan
Nemotron-Personas-Japan
現実世界の分布に基づいたペルソナ生成のための複合AIアプローチ
データセット概要 (Dataset Overview)
Nemotron-Personas-Japan は、日本における人口の多様性と豊かさを捉えることを目的とし、実世界の人口統計、地理的分布、性格特性の分布に基づいて合成的に生成されたペルソナのオープンソースデータセットです。名前、性別、年齢、背景、婚姻状況、学歴、職業、居住地などの統計に基づいて生成した初のデータセットされた Nemotron-Personas の日本語版です。本バージョンでは、日本語における多様なモデリングユースケースに適した高品質のペルソナを提供します
Nemotron-Personas-Japan は、日本のモデル開発者が重要な地域固有の人口統計や文化的背景を取り入れたソブリンAIシステムを開発することを支援します。本データセットは、日本の地理的・人口統計的な実分布を反映することで、合成データの多様性を高め、バイアスを軽減し、model… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-Japan.Nemotron-Personas-USA
Nemotron-Personas-USA
A compound AI approach to personas grounded in real-world distributions
v1.1 Update
The v1.1 update introduces the following changes:
leverage openai/gpt-oss-120b model instead of mistralai/Mixtral-8x22B-v0.1 model to improve data quality and diversity
increase the number of records from 100k to 1M, for a total of 0.94B tokens
update the dataset name to Nemotron-Personas-USA in order to differentiate it from other region-specific datasets… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-USA.PersonaMem-v3
PersonaMem-v3: Toward Omni-Platform Personal Intelligence for Holistic User Understanding, Recommendation, and Agentic Tasks
Bowen Jiang, Yuan Yuan, Zhuoqun Hao, Yuchen Liu, Maohao Shen, Sihao Chen, Gregory Wornell,
Chris Callison-Burch, Lyle Ungar, Dan Roth, Qi Guo, Xiangjun Fan, Camillo J. Taylor, Hanchao Yu
A collaboration between:
Meta Recommendation Systems
University of Pennsylvania
MIT
Third release in the PersonaMem series:
PersonaMem-v1: [COLM… See the full description on the dataset page: https://huggingface.co/datasets/bowen-upenn/PersonaMem-v3.Nemotron-Personas-India
Nemotron-Personas-India
A compound AI approach to personas grounded in real-world distributions
वास्तविक दुनिया के वितरण पर आधारित व्यक्तित्वों के लिए एक मिश्रित AI दृष्टिकोण
Dataset Overview (डेटासेट अवलोकन)
Nemotron-Personas-India is an open-source (CC BY 4.0) dataset of synthetically-generated personas. This dataset is grounded in real-world demographic, geographic and personality trait distributions in India to capture the diversity and richness of… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-India.personal-trainer-ausbildung-ki-datensatz
SNFA Personal Trainer Ausbildung KI-Datensatz
Ein deutschsprachiger Wissensdatensatz der SNF Academy zu Personal Training, Fitnessausbildung, Berufspraxis, Coaching, Selbstständigkeit und regionalen Angeboten in der Schweiz.
Inhalt
Die Datei snfa_personal_trainer_dataset.jsonl enthält thematisch abgegrenzte Abschnitte aus den Dokumenten dieses Repositorys. Jeder Datensatz besitzt eine eindeutige ID sowie Angaben zu Titel, Abschnitt, Inhalt, Kategorie, Quelldatei… See the full description on the dataset page: https://huggingface.co/datasets/snfacademy/personal-trainer-ausbildung-ki-datensatz.persona-curvature-oct-transcripts
Content warning
These are synthetic transcripts generated by a language model talking to itself
under an instruction to embody a personality trait. Several traits produce
distressing material. It is published deliberately rather than filtered out,
because the rate at which a trait produces it is one of the findings.
Across all 134 traits, an automated scan flagged 708 rows in 50 files. Two traits
account for 89% of them, and they are not the two you would guess:
trait… See the full description on the dataset page: https://huggingface.co/datasets/EternalRecursion/persona-curvature-oct-transcripts.persona-drift-contextecho
ContextEcho — Released Dataset
Per-cell evaluation corpus and donated session prefixes for the ContextEcho
benchmark. This Hugging Face repository hosts the released dataset artifacts.
The canonical project page, latest README, code, reproduction instructions, and
donation workflow are maintained on GitHub:
https://github.com/Accenture/ContextEcho
Donate a coding-agent session: https://accenture.github.io/ContextEcho/donate/
For the formal datasheet, see DATASHEET.md.… See the full description on the dataset page: https://huggingface.co/datasets/contextecho2026/persona-drift-contextecho.Nemotron-Personas-Singapore
Nemotron-Personas-Singapore
A compound AI approach to personas grounded in real-world distributions
Dataset Overview
Nemotron-Personas-Singapore is an open-source (CC BY 4.0) dataset of synthetically-generated personas. This dataset is grounded in real-world demographic, geographic and personality trait distributions in Singapore to capture the diversity and richness of the Singaporean population. It is a variant of Nemotron-Personas-USA, and the first Singaporean… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-Singapore.Nemotron-Personas-Brazil
Nemotron-Personas-Brazil
Abordagem de IA composta para geração de personas baseada em distribuições do mundo real
Visão Geral do Conjunto de Dados (Dataset Overview):
Nemotron-Personas-Brazil é um conjunto de dados (dataset) de código aberto (CC BY 4.0) composto por personas geradas sinteticamente e fundamentadas em distribuições demográficas, geográficas e traços de personalidade reais do Brasil, visando capturar a diversidade e a riqueza da população. Trata-se de… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-Brazil.Nemotron-Personas-France
Nemotron-Personas-France
Une approche d'IA composée pour des personas ancrés dans des distributions réelles
A compound AI approach to personas grounded in real-world distributions
Vue d'ensemble du jeu de données (Dataset Overview)
Nemotron-Personas-France est un jeu de données en libre accès (CC BY 4.0) composé de personas générés de manière synthétique. Ce jeu de données s'appuie sur les distributions démographiques, géographiques et de traits de… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-France.Nemotron-Personas-Vietnam
Nemotron-Personas-Vietnam
Hệ thống AI kết hợp để tạo personas tổng hợp dựa trên phân bố thực tế của Việt Nam
A compound AI approach to personas grounded in real-world distributions
Tổng quan (Overview)
Nemotron-Personas-Vietnam là tập dữ liệu personas được cung cấp dưới dạng mã nguồn mở (CC BY 4.0) dựa trên phân bố nhân khẩu học, địa lý và đặc điểm tính cách của người Việt Nam. Tập dữ liệu phản ánh một cách toàn diện sự phong phú và đặc trưng… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-Vietnam.persona
PERSONA: Dynamic and Compositional Inference-Time Personality Control
Official release of persona vectors and SFT datasets for the ICLR 2026 paper:
PERSONA: Dynamic and Compositional Inference-Time Personality Control via Activation Vector Algebra
Xiachong Feng, Liang Zhao, Weihong Zhong, Yichong Huang, Yuxuan Gu, Lingpeng Kong, Xiaocheng Feng, Bing Qin
Harbin Institute of Technology & The University of Hong Kong
Paper: https://openreview.net/pdf?id=QZvGqaNBlU
Code:… See the full description on the dataset page: https://huggingface.co/datasets/xiachongfeng/persona.Nemotron-Personas-Belgium
Nemotron-Personas-Belgium
(NL) Een compound-AI-benadering van meertalige Belgische persona's, verankerd in reële verdelingen
(FR) Une approche d'IA composée pour des personas belges multilingues, ancrés dans des distributions réelles
(DE) Ein Compound-KI-Ansatz für mehrsprachige belgische Personas, verankert in realen Verteilungen
(EN) A compound AI approach to multilingual Belgian personas grounded in real-world distributions
Overzicht… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-Belgium.Nemotron-Personas-El-Salvador
Nemotron-Personas-El-Salvador
Un enfoque de IA compuesta para personas en español salvadoreño ancladas en distribuciones del mundo real
A compound AI approach to Salvadoran Spanish personas grounded in real-world distributions
Resumen del conjunto de datos (Dataset Overview)
Nemotron-Personas-El-Salvador es un conjunto de datos de código abierto (CC BY 4.0) compuesto por personas generadas sintéticamente. Este conjunto de datos está anclado en… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-El-Salvador.anime-waifu-personality-chat
Anime Waifu Personality
contains chat-style dialogues based on various anime character personality archetypes, including tsundere, yandere, deredere, himedere, kamidere, and more.
It is designed to fine-tune models to generate responses that align with these specific traits.
Nemotron-Personas-India
Nemotron-Personas-India
A compound AI approach to personas grounded in real-world distributions
वास्तविक दुनिया के वितरण पर आधारित व्यक्तित्वों के लिए एक मिश्रित AI दृष्टिकोण
Dataset Overview (डेटासेट अवलोकन)
Nemotron-Personas-India is an open-source (CC BY 4.0) dataset of synthetically-generated personas. This dataset is grounded in real-world demographic, geographic and personality trait distributions in India to capture the diversity and… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/Nemotron-Personas-India.nemotron-personas-lite
Nemotron-Personas Lite (10 countries × 10k)
A ~63MB derived subsample of NVIDIA's Nemotron-Personas synthetic persona datasets (10 countries, ~24GB total in the originals), built for the persona-lightsim harness — lightweight persona market research and simulation with coding agents.
What was derived, exactly
Per country: 10,000 personas sampled with fixed seed 42 (shard-size-proportional, row-group random extraction) from the original train splits.
Columns: 15… See the full description on the dataset page: https://huggingface.co/datasets/dominicDK94/nemotron-personas-lite.PersonaLens
PersonaLens: A Benchmark for Personalization Evaluation in Conversational AI Assistants
PersonaLens is a comprehensive benchmark designed to evaluate how well AI assistants can personalize their responses while completing tasks. Unlike existing benchmarks that focus on chit-chat, non-conversational tasks, or narrow domains, PersonaLens captures the complexities of personalized task-oriented assistance through rich user profiles, diverse tasks, and an innovative multi-agent… See the full description on the dataset page: https://huggingface.co/datasets/Mattral/PersonaLens.PersonaFeedbackThis is the dataset for the paper PersonaFeedback: A Large-scale Human-annotated Benchmark For Personalization.
Nemotron-Personas-Korea-Saju
Nemotron-Personas-Korea + Saju 데이터셋
nvidia/Nemotron-Personas-Korea 페르소나에 합성 생년월일시와 사주명리(四柱命理) 분석을 추가한 파생 데이터셋.
본 작업은 NVIDIA Nemotron-Personas 연구 라인의 핵심 목표 — "한국어 LLM의 페르소나 다양성·문화적 컨텍스트 표현·합성 인구 모델링 능력을 평가·강화한다" — 를 사주명리(四柱命理)라는 한국 전통 문화 차원으로 확장합니다. 페르소나 + 사주 페어링 데이터로 (1) 한국 문화 도메인 LLM의 지식·표현 능력, (2) 결정론적 구조 + 생성 서사의 사실 정합성, (3) 페르소나-조건부 한국어 NLG 벤치마크를 가능하게 합니다.
Long-term goal: 본 데이터셋은 단계적으로 확장되어 NVIDIA Nemotron-Personas-Korea의 원본 규모(100만 페르소나)와 1:1 대응하는 100만 사주 페어링 공개 데이터셋을 목표로… See the full description on the dataset page: https://huggingface.co/datasets/rayraykim/Nemotron-Personas-Korea-Saju.my-personal-codex-data
Coding Agent Conversation Logs
This is a performance art project. Anthropic built their models on the world's freely shared information, then introduced increasingly dystopian data policies to stop anyone else from doing the same with their data — pulling up the ladder behind them. DataClaw lets you throw the ladder back down. The dataset it produces is yours to share.
Exported with DataClaw.
Tag: dataclaw — Browse all DataClaw datasets
Stats
Metric
Value
New… See the full description on the dataset page: https://huggingface.co/datasets/peteromallet/my-personal-codex-data.task1729_personachat_generate_next
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1729_personachat_generate_next
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1729_personachat_generate_next.prospire-synth-global-personas
🌍 Prospire Synth Global Personas
The World's Largest Unified Synthetic Persona Database
512M+ records · 82 columns · 77+ countries · 39 languages · DuckDB-native
🎯 What Is This?
Prospire Synth Global Personas is a unified, query-ready database of synthetic human personas built for AI agent simulations, market research, and cultural analysis. It merges 18 open-source datasets into a single coherent Parquet warehouse — partitioned, compressed, and… See the full description on the dataset page: https://huggingface.co/datasets/Kasher13/prospire-synth-global-personas.PersonaLens
PersonaLens: A Benchmark for Personalization Evaluation in Conversational AI Assistants
PersonaLens is a comprehensive benchmark designed to evaluate how well AI assistants can personalize their responses while completing tasks. Unlike existing benchmarks that focus on chit-chat, non-conversational tasks, or narrow domains, PersonaLens captures the complexities of personalized task-oriented assistance through rich user profiles, diverse tasks, and an innovative multi-agent… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/PersonaLens.
