datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
personal-trainer-ausbildung-ki-datensatz
SNFA Personal Trainer Ausbildung KI-Datensatz
Ein deutschsprachiger Wissensdatensatz der SNF Academy zu Personal Training, Fitnessausbildung, Berufspraxis, Coaching, Selbstständigkeit und regionalen Angeboten in der Schweiz.
Inhalt
Die Datei snfa_personal_trainer_dataset.jsonl enthält thematisch abgegrenzte Abschnitte aus den Dokumenten dieses Repositorys. Jeder Datensatz besitzt eine eindeutige ID sowie Angaben zu Titel, Abschnitt, Inhalt, Kategorie, Quelldatei… See the full description on the dataset page: https://huggingface.co/datasets/snfacademy/personal-trainer-ausbildung-ki-datensatz.anime-waifu-personality-chat
Anime Waifu Personality
contains chat-style dialogues based on various anime character personality archetypes, including tsundere, yandere, deredere, himedere, kamidere, and more.
It is designed to fine-tune models to generate responses that align with these specific traits.
ff-model-personalityPersonal-Finance-Queries
Dataset Description
A curated collection of Reddit posts and top comments focused on personal finance questions. The data is further filtered with the help of LLM-based Voting scores. These scores determine if the query is relevant to a person's financial queries among the other posts of the subreddits.
Dataset Structure
Columns:
category: The sub-domain of personal finance that the query belongs to.
subreddit: Source subreddit (string, categorical)
query: User’s… See the full description on the dataset page: https://huggingface.co/datasets/Akhil-Theerthala/Personal-Finance-Queries.AFM-WebAgent-SFT-Dataset
Data Introduction
This dataset serves as the core training data for Agent Foundation Models (AFMs), specifically designed to elicit end-to-end multi-agent reasoning capabilities in large language models. Built on the novel "Chain-of-Agents (CoA)" paradigm, the dataset leverages a multi-agent distillation framework to transform collaboration processes from state-of-the-art multi-agent systems into trajectory data suitable for supervised fine-tuning (SFT), simulating dynamic… See the full description on the dataset page: https://huggingface.co/datasets/PersonalAILab/AFM-WebAgent-SFT-Dataset.Kuvera-PersonalFinance-V2.1
Personal Finance Reasoning-V2.1
This dataset is associated with the paper Synthesizing Behaviorally-Grounded Reasoning Chains: A Data-Generation Framework for Personal Finance LLMs.
This is a scaled up version of the PersonalFinance-V2 dataset with some pipeline streamlining done.*
1. Introduction & Motivation
The landscape of financial AI benchmarks is currently dominated by applications in corporate finance, algorithmic trading, and general financial knowledge… See the full description on the dataset page: https://huggingface.co/datasets/Akhil-Theerthala/Kuvera-PersonalFinance-V2.1.personalization-reddit
personalization-reddit
Per-subreddit (query, preferred_answer) pairs mined from Reddit using an
OP-thanks-reply heuristic: when the original poster (OP) replies to a
comment with thanks/gratitude, that parent comment is treated as their
preferred answer to their own question.
Source
Raw post + comment dumps from the
arctic_shift Pushshift
mirror, fetched per-subreddit (entire history through the fetch date) and
extracted with the pipeline in… See the full description on the dataset page: https://huggingface.co/datasets/dipikakhullar/personalization-reddit.PersonalFinance_v2
Personal Finance Reasoning-V2
P.S. This dataset has won the First prize in the Reasoning Datasets Competition, organized by Bespoke Labs, HuggingFace & Together.AI During the months of April-May 2025. More details can be found here.
1. Introduction & Motivation
The landscape of financial AI benchmarks is currently dominated by applications in corporate finance, algorithmic trading, and general financial knowledge extraction. While valuable, these benchmarks often… See the full description on the dataset page: https://huggingface.co/datasets/Akhil-Theerthala/PersonalFinance_v2.personal-info-unlearning
Synthetic Personal Information Unlearning Dataset
Dataset Description
This dataset is designed for research on large language model (LLM) unlearning in controlled synthetic personal-information settings.
It contains synthetic profiles and question-answer data for four personal attributes:
Year of birth
Blood type
Postcode
Social insurance number
The benchmark provides three forget-set sizes: N = 5, 20, 40.
All personal-profile data are synthetically generated… See the full description on the dataset page: https://huggingface.co/datasets/shichenghu/personal-info-unlearning.personality-sarcastic-humor
_____ _ _ _____ _ _
| __ (_) | | | __ (_) | |
| |__) | _ __ | | __ | |__) |__ _____| |
| ___/ | '_ \| |/ / | ___/ \ \/ / _ \ |
| | | | | | | < | | | |> < __/ |
|_| |_|_| |_|_|\_\ |_| |_/_/\_\___|_|
🎨 Pink Pixel: Sarcastic, Witty, and Snarky Personality Dataset 🎭
Welcome to the Pink Pixel Sarcastic Humor dataset! This dataset is meticulously crafted to help you fine-tune… See the full description on the dataset page: https://huggingface.co/datasets/PinkPixel/personality-sarcastic-humor.Personalized_Safety_Data
📦 Personalized Risk and Dilemma Dataset for LLM Safety Research
📝 Dataset Summary
This is the first dataset designed to support research on personalized risk and emotional vulnerability in the context of Large Language Models (LLMs).
The dataset contains 8,000+ real-world, anonymized personal queries, extracted from Reddit and annotated with structured profile metadata, including emotional states, demographic information, and life contexts (e.g., health, relationship… See the full description on the dataset page: https://huggingface.co/datasets/wick1d/Personalized_Safety_Data.my-personal-codex-data
Coding Agent Conversation Logs
This is a performance art project. Anthropic built their models on the world's freely shared information, then introduced increasingly dystopian data policies to stop anyone else from doing the same with their data - pulling up the ladder behind them. DataClaw lets you throw the ladder back down. The dataset it produces is yours to share.
Exported with DataClaw.
Tag: dataclaw - Browse all DataClaw datasets
Stats
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/masterda/my-personal-codex-data.O-Researcher-SFT-Dataset
Data Introduction
This dataset serves as the core training data for O-Researcher, specifically designed to elicit end-to-end, multi-turn, multi-tool deep research capabilities in large language models. Built on the Multi-Agent Data Synthesis paradigm, the dataset leverages collaborative AI agents to simulate complex tool-integrated reasoning, transforming multi-agent research workflows into trajectory data suitable for supervised fine-tuning (SFT), enabling dynamic web search, page… See the full description on the dataset page: https://huggingface.co/datasets/PersonalAILab/O-Researcher-SFT-Dataset.anime-waifu-personality-chat
Anime Waifu Personality
This dataset contains chat-style dialogues based on various anime character personality archetypes, including tsundere, yandere, deredere, himedere, kamidere, and more.
It is designed to fine-tune models to generate responses that align with these specific traits.
Here's a few example:
{
"trait": "tsundere",
"dialogue": "H-Holding hands?! W-Well, I guess if you’re that desperate..."
},
{
"trait": "yandere",
"dialogue": "If I can't… See the full description on the dataset page: https://huggingface.co/datasets/Shxbhxm21/anime-waifu-personality-chat.personalization-reddit-user-histories
personalization-reddit-user-histories
Per-user chronological histories of answered questions across all
subreddits. Derived from dipikakhullar/personalization-reddit: every
(query, preferred_answer) pair a user authored as OP, grouped by user and
sorted by time, slimmed to the four fields needed to model a user's timeline.
Each record is one user. Users with a single interaction are dropped (a
timeline needs more than one point).
Selection: seen-the-top… See the full description on the dataset page: https://huggingface.co/datasets/dipikakhullar/personalization-reddit-user-histories.my-personal-codex-data
Coding Agent Conversation Logs
This is a performance art project. Anthropic built their models on the world's freely shared information, then introduced increasingly dystopian data policies to stop anyone else from doing the same with their data — pulling up the ladder behind them. DataClaw lets you throw the ladder back down. The dataset it produces is yours to share.
Exported with DataClaw.
Tag: dataclaw — Browse all DataClaw datasets
Stats
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/xuechengjiang/my-personal-codex-data.AFM-MHQA-Agent-SFT-Dataset
Data Introduction
This dataset serves as the core training data for Agent Foundation Models (AFMs), specifically designed to elicit end-to-end multi-agent reasoning capabilities in large language models. Built on the novel "Chain-of-Agents (CoA)" paradigm, the dataset leverages a multi-agent distillation framework to transform collaboration processes from state-of-the-art multi-agent systems into trajectory data suitable for supervised fine-tuning (SFT), simulating dynamic… See the full description on the dataset page: https://huggingface.co/datasets/PersonalAILab/AFM-MHQA-Agent-SFT-Dataset.personality-traits
Personality Traits
29 personality trait archetypes with core behavioral patterns, observable behaviors, and mitigation strategies.
Quick Start
from datasets import load_dataset
ds = load_dataset("buley/personality-traits")
print(ds["train"][0])
Categories
DEFENSIVE_MASKING — The Tough Guy, The Saint, Passive-Aggressive Charmer
VULNERABILITY_DEFENSIVE — The Victim, The People Pleaser
CONTROL_ORIENTED — The Control Freak, Domineering Behavior… See the full description on the dataset page: https://huggingface.co/datasets/buley/personality-traits.yt-personalities
Dataset Information
name: Youtubers by Big Five Personality Traits
license: gpl-3.0
Description
description: |
In trait theory, the Big Five personality traits (sometimes known as the five-factor model of personality or OCEAN or CANOE models) are a group of five characteristics used to study personality:
Openness to Experience (inventive/curious vs. consistent/cautious)
Conscientiousness (efficient/organized vs. extravagant/careless)
Extraversion (outgoing/energetic… See the full description on the dataset page: https://huggingface.co/datasets/visualcomments/yt-personalities.AFM-CodeAgent-SFT-Dataset
Data Introduction
This dataset serves as the core training data for Agent Foundation Models (AFMs), specifically designed to elicit end-to-end multi-agent reasoning capabilities in large language models. Built on the novel "Chain-of-Agents (CoA)" paradigm, the dataset leverages a multi-agent distillation framework to transform collaboration processes from state-of-the-art multi-agent systems into trajectory data suitable for supervised fine-tuning (SFT), simulating dynamic… See the full description on the dataset page: https://huggingface.co/datasets/PersonalAILab/AFM-CodeAgent-SFT-Dataset.personal-information-prompts
Personal Information Prompts
This dataset contains multilingual prompts derived from the all_sample subset of the agentlans/allenai-WildChat-4.8M dataset. Each prompt features artificially inserted personally identifiable information (PII) generated randomly with the Faker Python package for various locales.
Each rewritten prompt uses the google/gemma-3-12b-it model to incorporate the synthetic personal data.
Dataset fields for the two configurations:
classification… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/personal-information-prompts.personalized_passkey_retrieval
Dataset Summary
This dataset contains the data for personalized passkey retrieval task in the paper Improving Text Embeddings with Large Language Models.
Data Fields
query: a string feature.
candidates: List of string feature, 100 candidates for each query.
label: a int32 feature, the index of the correct candidate in the candidates list, always 0.
context_length: a int32 feature, the approximate length for the candidate documents.
How to use this dataset… See the full description on the dataset page: https://huggingface.co/datasets/intfloat/personalized_passkey_retrieval.rubai-NER-150K-Personal
Rubai NER Dataset - Personal Information Detection (Synthetic)
A dataset for training Named Entity Recognition (NER) models to detect personal information in Uzbek and Russian text. All Data Synthetic, no contains real personal information!
Dataset Description
This dataset contains 142,704 annotated examples for detecting personal information entities in informal Uzbek and Russian text (Latin and Cyrillic scripts).
Supported Entity Types
Entity… See the full description on the dataset page: https://huggingface.co/datasets/islomov/rubai-NER-150K-Personal.my-personal-codex-data
Coding Agent Conversation Logs
This is a performance art project. Anthropic built their models on the world's freely shared information, then introduced increasingly dystopian data policies to stop anyone else from doing the same with their data — pulling up the ladder behind them. DataClaw lets you throw the ladder back down. The dataset it produces is yours to share.
Exported with DataClaw.
Tag: dataclaw — Browse all DataClaw datasets
Stats
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/akenove/my-personal-codex-data.my-personal-codex-data
Coding Agent Conversation Logs
This is a performance art project. Anthropic built their models on the world's freely shared information, then introduced increasingly dystopian data policies to stop anyone else from doing the same with their data - pulling up the ladder behind them. DataClaw lets you throw the ladder back down. The dataset it produces is yours to share.
Exported with DataClaw.
Tag: dataclaw - Browse all DataClaw datasets
Stats
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/wop/my-personal-codex-data.orpheus-tts-dataset-preserving-personalityPersonal-Dataset
Personal Dataset
LIghtJUNction's Dataset
Dataset Details
Dataset Description
LIghtJUNction's Dataset is a manually curated dataset for professional profile shaping, assistant alignment, and software-engineering preference tuning. It combines selected public technical profile facts with manually curated interaction preferences and reputation guardrails.
The dataset intentionally avoids sensitive, private, low-quality, or… See the full description on the dataset page: https://huggingface.co/datasets/LIghtJUNction/Personal-Dataset.anime-waifu-personality-chat
Anime Waifu Personality
contains chat-style dialogues based on various anime character personality archetypes, including tsundere, yandere, deredere, himedere, kamidere, and more.
It is designed to fine-tune models to generate responses that align with these specific traits.
my-personal-codex-data
Coding Agent Conversation Logs
This is a performance art project. Anthropic built their models on the world's freely shared information, then introduced increasingly dystopian data policies to stop anyone else from doing the same with their data — pulling up the ladder behind them. DataClaw lets you throw the ladder back down. The dataset it produces is yours to share.
Exported with DataClaw.
Tag: dataclaw — Browse all DataClaw datasets
Stats
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/michaelwaves/my-personal-codex-data.PersonalityArchetypeMessage
Personality Archetype Message Dataset
This dataset contains 225 samples designed to train models that generate motivational messages tailored to user attributes.
Each entry includes:
age: an integer between 18–65
archetype: one of 15 distinct personality types
profession: a wide range of jobs across various sectors
city: locations across the U.S. and internationally
daily_message: a motivational message generated based on the above inputs
Use Case
This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/hanaelbatouty/PersonalityArchetypeMessage.
