datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PersonaHub
Scaling Synthetic Data Creation with 1,000,000,000 Personas
This repo releases data introduced in our paper Scaling Synthetic Data Creation with 1,000,000,000 Personas:
We propose a novel persona-driven data synthesis methodology that leverages various perspectives within a large language model (LLM) to create diverse synthetic data. To fully exploit this methodology at scale, we introduce PERSONA HUB – a collection of 1 billion diverse personas automatically curated from web… See the full description on the dataset page: https://huggingface.co/datasets/proj-persona/PersonaHub.personal-trainer-ausbildung-ki-datensatz
SNFA Personal Trainer Ausbildung KI-Datensatz
Ein deutschsprachiger Wissensdatensatz der SNF Academy zu Personal Training, Fitnessausbildung, Berufspraxis, Coaching, Selbstständigkeit und regionalen Angeboten in der Schweiz.
Inhalt
Die Datei snfa_personal_trainer_dataset.jsonl enthält thematisch abgegrenzte Abschnitte aus den Dokumenten dieses Repositorys. Jeder Datensatz besitzt eine eindeutige ID sowie Angaben zu Titel, Abschnitt, Inhalt, Kategorie, Quelldatei… See the full description on the dataset page: https://huggingface.co/datasets/snfacademy/personal-trainer-ausbildung-ki-datensatz.persona
PERSONA: Dynamic and Compositional Inference-Time Personality Control
Official release of persona vectors and SFT datasets for the ICLR 2026 paper:
PERSONA: Dynamic and Compositional Inference-Time Personality Control via Activation Vector Algebra
Xiachong Feng, Liang Zhao, Weihong Zhong, Yichong Huang, Yuxuan Gu, Lingpeng Kong, Xiaocheng Feng, Bing Qin
Harbin Institute of Technology & The University of Hong Kong
Paper: https://openreview.net/pdf?id=QZvGqaNBlU
Code:… See the full description on the dataset page: https://huggingface.co/datasets/xiachongfeng/persona.person-names-ner
Dataset Card for Person Full Name NER Parsing
This dataset contains 3,383,944 curated and augmented person names, designed specifically for training Token Classification (NER) models. The primary task is to parse a full name string into its FirstName and LastName components, correctly handling multi-word names and different ordering formats.
Dataset Details
Dataset Description
This dataset is built to train robust models that can understand and segment human… See the full description on the dataset page: https://huggingface.co/datasets/ele-sage/person-names-ner.anime-waifu-personality-chat
Anime Waifu Personality
contains chat-style dialogues based on various anime character personality archetypes, including tsundere, yandere, deredere, himedere, kamidere, and more.
It is designed to fine-tune models to generate responses that align with these specific traits.
ff-model-personalityPersonal-Finance-Queries
Dataset Description
A curated collection of Reddit posts and top comments focused on personal finance questions. The data is further filtered with the help of LLM-based Voting scores. These scores determine if the query is relevant to a person's financial queries among the other posts of the subreddits.
Dataset Structure
Columns:
category: The sub-domain of personal finance that the query belongs to.
subreddit: Source subreddit (string, categorical)
query: User’s… See the full description on the dataset page: https://huggingface.co/datasets/Akhil-Theerthala/Personal-Finance-Queries.AFM-WebAgent-SFT-Dataset
Data Introduction
This dataset serves as the core training data for Agent Foundation Models (AFMs), specifically designed to elicit end-to-end multi-agent reasoning capabilities in large language models. Built on the novel "Chain-of-Agents (CoA)" paradigm, the dataset leverages a multi-agent distillation framework to transform collaboration processes from state-of-the-art multi-agent systems into trajectory data suitable for supervised fine-tuning (SFT), simulating dynamic… See the full description on the dataset page: https://huggingface.co/datasets/PersonalAILab/AFM-WebAgent-SFT-Dataset.first-person-dialogue
First Person Dialogue Dataset
Dataset Description
This dataset is designed for training one-on-one chatbots, featuring a wide range of social roles and situations.
It allows for assigning a name to the AI character, creating a more personalized, more intimate conversational experience.
Contents
The dataset is a curated combination of several existing datasets:
allenai/soda
allenai/prosocial-dialog
Estwld/empathetic_dialogues_llm… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/first-person-dialogue.SCOPE-Persona
SCOPE Personas (Nemotron Augmentation)
This dataset contains synthetic persona profiles constructed from socio-psychological framework (SCOPE) [https://arxiv.org/pdf/2601.07110], designed to better support LLM simulation usecases in social and behavioral science. It is intended to be used alongside Nemotron-Persona [https://huggingface.co/datasets/nvidia/Nemotron-Personas-USA]. Personas are grounded in a 141-item sociopsychological questionnaire spanning eight facets.
You can… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/SCOPE-Persona.counterfactuals
Persona Bias Counterfactuals
This dataset contains counterfactual examples used for persona-bias circuit discovery and intervention experiments.
Repository Layout
Hugging Face dataset config = model
Hugging Face dataset split = counterfactual strategy
task and axis are columns, not separate dataset configs
data/<model>/<strategy>.jsonl.gz
manifest.jsonl
Strategies
Split
Meaning
original
Full original counterfactual set derived from… See the full description on the dataset page: https://huggingface.co/datasets/PersonaBias/counterfactuals.AItuber-Personas-Japan
AItuber Persona Dataset
概要
本データセットは、AItuber(AI VTuber)のペルソナ設計に必要な コンセプト設計書・実装用システムプロンプト・配信テーマリスト の3点セットを、LLMを用いて合成的に生成したものです。多様なジャンル・性格・ビジュアルの組み合わせから、即座に実運用可能な品質のAItuberキャラクターデータを提供します。
生成にはSDG-LOOMという合成データ生成パイプラインとMoonshot-AIのKimi-K2.5を用いました。(sdg-loom)
データの説明
項目
内容
件数
195件
形式
JSONL(1行1JSON)
言語
日本語
生成日
2026年3月
ライセンス
odc-by ( Open Data Commons Attribution License )… See the full description on the dataset page: https://huggingface.co/datasets/DataPilot/AItuber-Personas-Japan.Kuvera-PersonalFinance-V2.1
Personal Finance Reasoning-V2.1
This dataset is associated with the paper Synthesizing Behaviorally-Grounded Reasoning Chains: A Data-Generation Framework for Personal Finance LLMs.
This is a scaled up version of the PersonalFinance-V2 dataset with some pipeline streamlining done.*
1. Introduction & Motivation
The landscape of financial AI benchmarks is currently dominated by applications in corporate finance, algorithmic trading, and general financial knowledge… See the full description on the dataset page: https://huggingface.co/datasets/Akhil-Theerthala/Kuvera-PersonalFinance-V2.1.personalization-reddit
personalization-reddit
Per-subreddit (query, preferred_answer) pairs mined from Reddit using an
OP-thanks-reply heuristic: when the original poster (OP) replies to a
comment with thanks/gratitude, that parent comment is treated as their
preferred answer to their own question.
Source
Raw post + comment dumps from the
arctic_shift Pushshift
mirror, fetched per-subreddit (entire history through the fetch date) and
extracted with the pipeline in… See the full description on the dataset page: https://huggingface.co/datasets/dipikakhullar/personalization-reddit.PersonalFinance_v2
Personal Finance Reasoning-V2
P.S. This dataset has won the First prize in the Reasoning Datasets Competition, organized by Bespoke Labs, HuggingFace & Together.AI During the months of April-May 2025. More details can be found here.
1. Introduction & Motivation
The landscape of financial AI benchmarks is currently dominated by applications in corporate finance, algorithmic trading, and general financial knowledge extraction. While valuable, these benchmarks often… See the full description on the dataset page: https://huggingface.co/datasets/Akhil-Theerthala/PersonalFinance_v2.evolving_personaspersonal-info-unlearning
Synthetic Personal Information Unlearning Dataset
Dataset Description
This dataset is designed for research on large language model (LLM) unlearning in controlled synthetic personal-information settings.
It contains synthetic profiles and question-answer data for four personal attributes:
Year of birth
Blood type
Postcode
Social insurance number
The benchmark provides three forget-set sizes: N = 5, 20, 40.
All personal-profile data are synthetically generated… See the full description on the dataset page: https://huggingface.co/datasets/shichenghu/personal-info-unlearning.Turkish-synthetic-personas
Turkish-synthetic-personas
Dataset Overview
This dataset is an open source synthetically generated persona dataset. To generate these personas a pipeline similar to Nemotron Persona generation pipeline was used. The dataset is grounded with real world demographic distribution of Turkiye using different statistical information provided by Turkish Statistical Institute (TÜİK).
The grounding data includes city, age, gender, education, employment status and marital… See the full description on the dataset page: https://huggingface.co/datasets/kesimeg/Turkish-synthetic-personas.Moonfrost-Persona-SFT
Moonfrost-Persona-SFT
Code · Site · Training runs
140,000 multi-turn conversations that teach a small chat model two things no public dataset
covers: who it is, and that "you" and "I" refer to different people. They were written for
the Moonfrost-777M-Instruct-v2
fine-tune, where they made up 3.3% of the rows, and they are built from templates rather
than generated by another model, so there is no scraped text and nothing from anyone else's
outputs in them. Regenerating the… See the full description on the dataset page: https://huggingface.co/datasets/whoashish115/Moonfrost-Persona-SFT.Open-Personix
Open-Personix
Dataset Summary
Open-Personix is a structured JSON dataset maintained under Poralus.
The dataset is primarily text and metadata: each record contains a relative image path,
a natural-language caption, and descriptive annotation fields for a person-centered sample.
The dataset is designed for workflows such as:
caption generation and caption analysis
text-based filtering over person annotations
metadata-aware retrieval and evaluation
multimodal experiments… See the full description on the dataset page: https://huggingface.co/datasets/Below-Image/Open-Personix.Personalized_Safety_Data
📦 Personalized Risk and Dilemma Dataset for LLM Safety Research
📝 Dataset Summary
This is the first dataset designed to support research on personalized risk and emotional vulnerability in the context of Large Language Models (LLMs).
The dataset contains 8,000+ real-world, anonymized personal queries, extracted from Reddit and annotated with structured profile metadata, including emotional states, demographic information, and life contexts (e.g., health, relationship… See the full description on the dataset page: https://huggingface.co/datasets/wick1d/Personalized_Safety_Data.personality-sarcastic-humor
_____ _ _ _____ _ _
| __ (_) | | | __ (_) | |
| |__) | _ __ | | __ | |__) |__ _____| |
| ___/ | '_ \| |/ / | ___/ \ \/ / _ \ |
| | | | | | | < | | | |> < __/ |
|_| |_|_| |_|_|\_\ |_| |_/_/\_\___|_|
🎨 Pink Pixel: Sarcastic, Witty, and Snarky Personality Dataset 🎭
Welcome to the Pink Pixel Sarcastic Humor dataset! This dataset is meticulously crafted to help you fine-tune… See the full description on the dataset page: https://huggingface.co/datasets/PinkPixel/personality-sarcastic-humor.wikipedia-persons-masked
wikipedia persons masked: A filtered version of the wikipedia dataset, with only pages of people
Dataset Summary
Contains ~70k pages from wikipedia, each describing a person. For each page, the person described in the text
is masked with a
Supported Tasks and Leaderboards
The dataset supports the tasks of fill-mask, but can also be used for other tasks such as question answering,
e.g. "Who is
Languages
english only
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/rcds/wikipedia-persons-masked.demoverse-personas-es-v1
Dataset Card for DemoVerse Personas ES v1
Resumen del dataset
demoverse-personas-es-v1 es un dataset de 100.000 personas sinteticas en espanol para Espana, disenado como artefacto publico y como capa operativa para simulacion sociológica.
El dataset se inspira metodologicamente en nvidia/Nemotron-Personas-France, pero no reutiliza sus filas ni intenta replicar la poblacion francesa. La adaptacion reescribe el marco para Espana, con clivajes territoriales, sistema de… See the full description on the dataset page: https://huggingface.co/datasets/apol/demoverse-personas-es-v1.O-Researcher-SFT-Dataset
Data Introduction
This dataset serves as the core training data for O-Researcher, specifically designed to elicit end-to-end, multi-turn, multi-tool deep research capabilities in large language models. Built on the Multi-Agent Data Synthesis paradigm, the dataset leverages collaborative AI agents to simulate complex tool-integrated reasoning, transforming multi-agent research workflows into trajectory data suitable for supervised fine-tuning (SFT), enabling dynamic web search, page… See the full description on the dataset page: https://huggingface.co/datasets/PersonalAILab/O-Researcher-SFT-Dataset.small-persona-dataset
small-persona-dataset
Bilingual (EN/ES) training dataset for a small NPC voice model. The model learns to take a plain factual sentence and rewrite it in a character's voice, conditioned on persona parameters.
Task
INPUT: TONE:grumpy STYLE:blunt HUMOR:dry RELATION:rival ROLE:blacksmith
FACT: Iron swords cost 15 gold.
OUTPUT: Fifteen gold. Still overpriced for your work.
The model is conditioned on 5 parameters: tone, style, humor, role, and relation.… See the full description on the dataset page: https://huggingface.co/datasets/walter-bd/small-persona-dataset.personalization-reddit-user-histories
personalization-reddit-user-histories
Per-user chronological histories of answered questions across all
subreddits. Derived from dipikakhullar/personalization-reddit: every
(query, preferred_answer) pair a user authored as OP, grouped by user and
sorted by time, slimmed to the four fields needed to model a user's timeline.
Each record is one user. Users with a single interaction are dropped (a
timeline needs more than one point).
Selection: seen-the-top… See the full description on the dataset page: https://huggingface.co/datasets/dipikakhullar/personalization-reddit-user-histories.persona_nemotron
Persona Nemotron PT Datasets
This is a collection of Portuguese synthetic datasets, consisting of 3 datasets, one with general questions from varied topics, one with math questions, and one with instruction-following requests.
The prompts were generated using an approach similar to PersonaHub, with a translated version of Nemotron Personas. Both prompts and answers were generated using Gemma 3-27B.
This dataset is provided as part of the AMALIA project and is… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/persona_nemotron.user_study-preference-personalized_0423_base_filtered
Filtered user study dataset
Source repo: ehejin/user_study-preference-personalized_0423_base
Each row is ONE item review (pre-rating, conversation, post-rating). Submission-level
fields (prolific_pid, demographics, background) are duplicated across rows that share
a submission.
The 25-50 rows here are the FIRST review for each unique pool index, selected the same
way the analysis plot uses — see scripts/plot_vote_shift_3way.py.
Total rows: 50
personaplex-finetuning-pharma-data-sample
PersonaPlex Finetuning — Pharma Data Sample
A 10-example slice of the synthetic patient-support / medication
adherence dataset used to train
demegire/personaplex-finetune-pharma.
The on-disk layout below is exactly what the trainer in
emotion-machine-org/personaplex-finetune
consumes — use this as a template when building your own.
Split: 8 train / 2 eval (mirrors the upstream 2003 / 20 split at
sample scale).
Layout
.
├── adhery_v2.jsonl # master… See the full description on the dataset page: https://huggingface.co/datasets/demegire/personaplex-finetuning-pharma-data-sample.
