datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
personal-facts-msc
Personal Facts (MSC) — Multi-Dimensional Annotation
A manually annotated dataset of 2,779 personal facts sampled from the
Multi-Session Chat (MSC)
corpus, labeled across seven dimensions that jointly characterize a fact's
topic, temporal anchoring, referent, lifetime, validity, and dialogue-continuation
potential.
The scheme extends PeaCoK with
two new top-level categories (Demographics, Possessions) and three new
dimensions (Duration, Validity / Invalidity Reason, Followup),
and… See the full description on the dataset page: https://huggingface.co/datasets/adugeen/personal-facts-msc.mbti-personality-datasetpersonal-trainer-ausbildung-ki-datensatz
SNFA Personal Trainer Ausbildung KI-Datensatz
Ein deutschsprachiger Wissensdatensatz der SNF Academy zu Personal Training, Fitnessausbildung, Berufspraxis, Coaching, Selbstständigkeit und regionalen Angeboten in der Schweiz.
Inhalt
Die Datei snfa_personal_trainer_dataset.jsonl enthält thematisch abgegrenzte Abschnitte aus den Dokumenten dieses Repositorys. Jeder Datensatz besitzt eine eindeutige ID sowie Angaben zu Titel, Abschnitt, Inhalt, Kategorie, Quelldatei… See the full description on the dataset page: https://huggingface.co/datasets/snfacademy/personal-trainer-ausbildung-ki-datensatz.PersonalizationV3anime-waifu-personality-chat
Anime Waifu Personality
contains chat-style dialogues based on various anime character personality archetypes, including tsundere, yandere, deredere, himedere, kamidere, and more.
It is designed to fine-tune models to generate responses that align with these specific traits.
AFM-WebAgent-RL-Dataset
Data Introduction
This dataset serves as the core training data for Agent Foundation Models (AFMs), specifically designed to elicit end-to-end multi-agent reasoning capabilities in large language models. Built on the novel "Chain-of-Agents (CoA)" paradigm, the dataset leverages a multi-agent distillation framework to transform collaboration processes from state-of-the-art multi-agent systems into trajectory data suitable for supervised fine-tuning (SFT), simulating dynamic… See the full description on the dataset page: https://huggingface.co/datasets/PersonalAILab/AFM-WebAgent-RL-Dataset.ff-model-personalityMegaDepth-v1fiqa-personal-finance-datasetpersonal_dialogThe PersonalDialog dataset is a large-scale multi-turn Chinese dialogue dataset containing various traits from a large number of speakers.
We are releasing about 5M sessions of carefully filtered dialogues.
Each utterance in PersonalDialog is associated with a speaker marked with traits like Gender, Location, Interest Tags.enron_personalization_test
Dataset Card for "enron_personalization_test"
More Information needed
PersonalizationV4HM-Personalized-Fashion-RecommendationsPersonalLLM
Dataset Card for PersonalLLM
The PersonalLLM dataset is a collection of prompts, responses, and rewards designed for personalized language model methodology development and evaluation. This dataset is presented in the paper PersonalLLM: Tailoring LLMs to Individual Preferences.
Dataset Details
Dataset Description
Curated by: Andrew Siah*, Tom Zollo*, Naimeng Ye, Ang Li, Namkoong Hongseok
Funded by: Digital Future Initiative at Columbia Business School… See the full description on the dataset page: https://huggingface.co/datasets/namkoong-lab/PersonalLLM.Personal-Finance-Queries
Dataset Description
A curated collection of Reddit posts and top comments focused on personal finance questions. The data is further filtered with the help of LLM-based Voting scores. These scores determine if the query is relevant to a person's financial queries among the other posts of the subreddits.
Dataset Structure
Columns:
category: The sub-domain of personal finance that the query belongs to.
subreddit: Source subreddit (string, categorical)
query: User’s… See the full description on the dataset page: https://huggingface.co/datasets/Akhil-Theerthala/Personal-Finance-Queries.AFM-WebAgent-SFT-Dataset
Data Introduction
This dataset serves as the core training data for Agent Foundation Models (AFMs), specifically designed to elicit end-to-end multi-agent reasoning capabilities in large language models. Built on the novel "Chain-of-Agents (CoA)" paradigm, the dataset leverages a multi-agent distillation framework to transform collaboration processes from state-of-the-art multi-agent systems into trajectory data suitable for supervised fine-tuning (SFT), simulating dynamic… See the full description on the dataset page: https://huggingface.co/datasets/PersonalAILab/AFM-WebAgent-SFT-Dataset.Personalized-RewardBench
Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization
📜 Paper | 🤗 Benchmark | 🖥️ Code
Abstract
Pluralistic alignment has emerged as a critical frontier in the development of Large Language Models (LLMs), with reward models (RMs) serving as a central mechanism for capturing diverse human values. While benchmarks for general response quality are prevalent, evaluating how well reward models account for individual user preferences remains an… See the full description on the dataset page: https://huggingface.co/datasets/QiyaoMa/Personalized-RewardBench.personal-continuity-protocol
Personal Continuity Protocol v1.0
Insurance Against Rupture
This dataset repository is the AI-readable distribution layer for the Personal Continuity Protocol (PCP).
PCP is an open continuity architecture for preserving evidence of personal rights, authorship, work, reputation, obligations and will across institutional, digital, legal and material rupture.
The central distinction is simple:
A person does not belong to the door through which institutions see them.
Modern… See the full description on the dataset page: https://huggingface.co/datasets/navimusaget/personal-continuity-protocol.Kuvera-PersonalFinance-V2.1
Personal Finance Reasoning-V2.1
This dataset is associated with the paper Synthesizing Behaviorally-Grounded Reasoning Chains: A Data-Generation Framework for Personal Finance LLMs.
This is a scaled up version of the PersonalFinance-V2 dataset with some pipeline streamlining done.*
1. Introduction & Motivation
The landscape of financial AI benchmarks is currently dominated by applications in corporate finance, algorithmic trading, and general financial knowledge… See the full description on the dataset page: https://huggingface.co/datasets/Akhil-Theerthala/Kuvera-PersonalFinance-V2.1.shrec_empathic
Social Human Robot Embodied Conversation (SHREC) Dataset: Empathic Subset (RSS 2026)
The SHREC Empathic subset contains real-world human-robot interaction video data from Shen et al. (2024), collected over a month-long deployment of social robots in participants’ homes, as participants engage in natural, empathic storytelling interactions with AI agents.
Authors: Dong Won Lee, Yubin Kim, Sooyeon Jeong, Denison Guvenoz, Parker Malachowsky, Louis-Philippe Morency, Cynthia… See the full description on the dataset page: https://huggingface.co/datasets/MIT-personal-robots/shrec_empathic.PG-Personalization-Amazon2023personalization-reddit
personalization-reddit
Per-subreddit (query, preferred_answer) pairs mined from Reddit using an
OP-thanks-reply heuristic: when the original poster (OP) replies to a
comment with thanks/gratitude, that parent comment is treated as their
preferred answer to their own question.
Source
Raw post + comment dumps from the
arctic_shift Pushshift
mirror, fetched per-subreddit (entire history through the fetch date) and
extracted with the pipeline in… See the full description on the dataset page: https://huggingface.co/datasets/dipikakhullar/personalization-reddit.PersonaSignal-PersonalizedResponse-Programming-Expertise-gpt-5-mini
Dataset card for PersonaSignal-PersonalizedResponse-Programming-Expertise-gpt-5-mini
This dataset was made with Curator.
Dataset details
A sample from the dataset:
{
"dimension_name": "programming_expertise",
"dimension_values": [
"Novice",
"Intermediate",
"Advanced"
],
"dimension_description": "Represents the user's practical fluency in software engineering. It shapes how they decompose problems, choose abstractions, weigh… See the full description on the dataset page: https://huggingface.co/datasets/JasonYan777/PersonaSignal-PersonalizedResponse-Programming-Expertise-gpt-5-mini.PersonalizationV3personal_dictionary
OpenGloss Dictionary (Word-Level)
Dataset Summary
OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English that integrates lexicographic definitions, encyclopedic context, etymological histories, and semantic relationships in a unified resource.
This dataset provides the words-level view where each record represents one lexeme (word or multi-word expression).
Key Statistics
150,101 lexemes across 150,101 English… See the full description on the dataset page: https://huggingface.co/datasets/caioloures/personal_dictionary.Personality_mypersonality
Dataset Card for "Personality_mypersonality"
More Information needed
shrec_wellness_dorm
Social Human Robot Embodied Conversation (SHREC) Dataset: Wellness Dorm Subset (RSS 2026)
The SHREC Wellness Dorm subset contains longitudinal, real-world human-robot interaction video data data from Jeong et al. (2020), where a robotic positive psychology coach was deployed in MIT student dormitories. Participants engaged in daily wellbeing sessions with the robot over the course of 1–4 weeks.
Authors: Dong Won Lee, Yubin Kim, Sooyeon Jeong, Denison Guvenoz, Parker… See the full description on the dataset page: https://huggingface.co/datasets/MIT-personal-robots/shrec_wellness_dorm.PersonalFinance_v2
Personal Finance Reasoning-V2
P.S. This dataset has won the First prize in the Reasoning Datasets Competition, organized by Bespoke Labs, HuggingFace & Together.AI During the months of April-May 2025. More details can be found here.
1. Introduction & Motivation
The landscape of financial AI benchmarks is currently dominated by applications in corporate finance, algorithmic trading, and general financial knowledge extraction. While valuable, these benchmarks often… See the full description on the dataset page: https://huggingface.co/datasets/Akhil-Theerthala/PersonalFinance_v2.personal-blog-cmsprism_personalized_0125
