CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01proj-persona /PersonaHub Scaling Synthetic Data Creation with 1,000,000,000 Personas This repo releases data introduced in our paper Scaling Synthetic Data Creation with 1,000,000,000 Personas: We propose a novel persona-driven data synthesis methodology that leverages various perspectives within a large language model (LLM) to create diverse synthetic data. To fully exploit this methodology at scale, we introduce PERSONA HUB – a collection of 1 billion diverse personas automatically curated from web… See the full description on the dataset page: https://huggingface.co/datasets/proj-persona/PersonaHub.texttext-generation100K<n<1M807 likes8.9k downloads1y agoHugging Face02snfacademy /personal-trainer-ausbildung-ki-datensatz SNFA Personal Trainer Ausbildung KI-Datensatz Ein deutschsprachiger Wissensdatensatz der SNF Academy zu Personal Training, Fitnessausbildung, Berufspraxis, Coaching, Selbstständigkeit und regionalen Angeboten in der Schweiz. Inhalt Die Datei snfa_personal_trainer_dataset.jsonl enthält thematisch abgegrenzte Abschnitte aus den Dokumenten dieses Repositorys. Jeder Datensatz besitzt eine eindeutige ID sowie Angaben zu Titel, Abschnitt, Inhalt, Kategorie, Quelldatei… See the full description on the dataset page: https://huggingface.co/datasets/snfacademy/personal-trainer-ausbildung-ki-datensatz.textquestion-answeringn<1K0 likes1.3k downloads2mo agoHugging Face03xiachongfeng /persona PERSONA: Dynamic and Compositional Inference-Time Personality Control Official release of persona vectors and SFT datasets for the ICLR 2026 paper: PERSONA: Dynamic and Compositional Inference-Time Personality Control via Activation Vector Algebra Xiachong Feng, Liang Zhao, Weihong Zhong, Yichong Huang, Yuxuan Gu, Lingpeng Kong, Xiaocheng Feng, Bing Qin Harbin Institute of Technology & The University of Hong Kong Paper: https://openreview.net/pdf?id=QZvGqaNBlU Code:… See the full description on the dataset page: https://huggingface.co/datasets/xiachongfeng/persona.texttext-generation100K<n<1M0 likes726 downloads5mo agoHugging Face04ele-sage /person-names-ner Dataset Card for Person Full Name NER Parsing This dataset contains 3,383,944 curated and augmented person names, designed specifically for training Token Classification (NER) models. The primary task is to parse a full name string into its FirstName and LastName components, correctly handling multi-word names and different ordering formats. Dataset Details Dataset Description This dataset is built to train robust models that can understand and segment human… See the full description on the dataset page: https://huggingface.co/datasets/ele-sage/person-names-ner.texttoken-classification1M<n<10M3 likes602 downloads1y agoHugging Face05scryptiam /anime-waifu-personality-chat Anime Waifu Personality contains chat-style dialogues based on various anime character personality archetypes, including tsundere, yandere, deredere, himedere, kamidere, and more. It is designed to fine-tune models to generate responses that align with these specific traits. texttext-generation1K<n<10K37 likes554 downloads7mo agoHugging Face06rdnfn /ff-model-personalitytextn<1K0 likes401 downloads5mo agoHugging Face07Akhil-Theerthala /Personal-Finance-Queries Dataset Description A curated collection of Reddit posts and top comments focused on personal finance questions. The data is further filtered with the help of LLM-based Voting scores. These scores determine if the query is relevant to a person's financial queries among the other posts of the subreddits. Dataset Structure Columns: category: The sub-domain of personal finance that the query belongs to. subreddit: Source subreddit (string, categorical) query: User’s… See the full description on the dataset page: https://huggingface.co/datasets/Akhil-Theerthala/Personal-Finance-Queries.textquestion-answering10K<n<100K9 likes192 downloads1y agoHugging Face08PersonalAILab /AFM-WebAgent-SFT-Dataset Data Introduction This dataset serves as the core training data for Agent Foundation Models (AFMs), specifically designed to elicit end-to-end multi-agent reasoning capabilities in large language models. Built on the novel "Chain-of-Agents (CoA)" paradigm, the dataset leverages a multi-agent distillation framework to transform collaboration processes from state-of-the-art multi-agent systems into trajectory data suitable for supervised fine-tuning (SFT), simulating dynamic… See the full description on the dataset page: https://huggingface.co/datasets/PersonalAILab/AFM-WebAgent-SFT-Dataset.text1K<n<10K10 likes191 downloads1y agoHugging Face09agentlans /first-person-dialogue First Person Dialogue Dataset Dataset Description This dataset is designed for training one-on-one chatbots, featuring a wide range of social roles and situations. It allows for assigning a name to the AI character, creating a more personalized, more intimate conversational experience. Contents The dataset is a curated combination of several existing datasets: allenai/soda allenai/prosocial-dialog Estwld/empathetic_dialogues_llm… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/first-person-dialogue.texttext-generation1M<n<10M4 likes189 downloads2y agoHugging Face10Salesforce /SCOPE-Persona SCOPE Personas (Nemotron Augmentation) This dataset contains synthetic persona profiles constructed from socio-psychological framework (SCOPE) [https://arxiv.org/pdf/2601.07110], designed to better support LLM simulation usecases in social and behavioral science. It is intended to be used alongside Nemotron-Persona [https://huggingface.co/datasets/nvidia/Nemotron-Personas-USA]. Personas are grounded in a 141-item sociopsychological questionnaire spanning eight facets. You can… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/SCOPE-Persona.text100K<n<1M2 likes188 downloads5mo agoHugging Face11PersonaBias /counterfactuals Persona Bias Counterfactuals This dataset contains counterfactual examples used for persona-bias circuit discovery and intervention experiments. Repository Layout Hugging Face dataset config = model Hugging Face dataset split = counterfactual strategy task and axis are columns, not separate dataset configs data/<model>/<strategy>.jsonl.gz manifest.jsonl Strategies Split Meaning original Full original counterfactual set derived from… See the full description on the dataset page: https://huggingface.co/datasets/PersonaBias/counterfactuals.texttext-classification100K<n<1M0 likes182 downloads3mo agoHugging Face12DataPilot /AItuber-Personas-Japan AItuber Persona Dataset 概要 本データセットは、AItuber(AI VTuber)のペルソナ設計に必要な コンセプト設計書・実装用システムプロンプト・配信テーマリスト の3点セットを、LLMを用いて合成的に生成したものです。多様なジャンル・性格・ビジュアルの組み合わせから、即座に実運用可能な品質のAItuberキャラクターデータを提供します。 生成にはSDG-LOOMという合成データ生成パイプラインとMoonshot-AIのKimi-K2.5を用いました。(sdg-loom) データの説明 項目 内容 件数 195件 形式 JSONL(1行1JSON) 言語 日本語 生成日 2026年3月 ライセンス odc-by ( Open Data Commons Attribution License )… See the full description on the dataset page: https://huggingface.co/datasets/DataPilot/AItuber-Personas-Japan.textn<1K30 likes161 downloads6mo agoHugging Face13Akhil-Theerthala /Kuvera-PersonalFinance-V2.1 Personal Finance Reasoning-V2.1 This dataset is associated with the paper Synthesizing Behaviorally-Grounded Reasoning Chains: A Data-Generation Framework for Personal Finance LLMs. This is a scaled up version of the PersonalFinance-V2 dataset with some pipeline streamlining done.* 1. Introduction & Motivation The landscape of financial AI benchmarks is currently dominated by applications in corporate finance, algorithmic trading, and general financial knowledge… See the full description on the dataset page: https://huggingface.co/datasets/Akhil-Theerthala/Kuvera-PersonalFinance-V2.1.texttext-classification10K<n<100K8 likes159 downloads9mo agoHugging Face14dipikakhullar /personalization-reddit personalization-reddit Per-subreddit (query, preferred_answer) pairs mined from Reddit using an OP-thanks-reply heuristic: when the original poster (OP) replies to a comment with thanks/gratitude, that parent comment is treated as their preferred answer to their own question. Source Raw post + comment dumps from the arctic_shift Pushshift mirror, fetched per-subreddit (entire history through the fetch date) and extracted with the pipeline in… See the full description on the dataset page: https://huggingface.co/datasets/dipikakhullar/personalization-reddit.text100K<n<1M1 likes146 downloads14d agoHugging Face15Akhil-Theerthala /PersonalFinance_v2 Personal Finance Reasoning-V2 P.S. This dataset has won the First prize in the Reasoning Datasets Competition, organized by Bespoke Labs, HuggingFace & Together.AI During the months of April-May 2025. More details can be found here. 1. Introduction & Motivation The landscape of financial AI benchmarks is currently dominated by applications in corporate finance, algorithmic trading, and general financial knowledge extraction. While valuable, these benchmarks often… See the full description on the dataset page: https://huggingface.co/datasets/Akhil-Theerthala/PersonalFinance_v2.texttext-classification1K<n<10K27 likes106 downloads1y agoHugging Face16deeper-team /evolving_personastext10K<n<100K0 likes98 downloads1y agoHugging Face17shichenghu /personal-info-unlearning Synthetic Personal Information Unlearning Dataset Dataset Description This dataset is designed for research on large language model (LLM) unlearning in controlled synthetic personal-information settings. It contains synthetic profiles and question-answer data for four personal attributes: Year of birth Blood type Postcode Social insurance number The benchmark provides three forget-set sizes: N = 5, 20, 40. All personal-profile data are synthetically generated… See the full description on the dataset page: https://huggingface.co/datasets/shichenghu/personal-info-unlearning.textquestion-answering100K<n<1M0 likes98 downloads27d agoHugging Face18kesimeg /Turkish-synthetic-personas Turkish-synthetic-personas Dataset Overview This dataset is an open source synthetically generated persona dataset. To generate these personas a pipeline similar to Nemotron Persona generation pipeline was used. The dataset is grounded with real world demographic distribution of Turkiye using different statistical information provided by Turkish Statistical Institute (TÜİK). The grounding data includes city, age, gender, education, employment status and marital… See the full description on the dataset page: https://huggingface.co/datasets/kesimeg/Turkish-synthetic-personas.text10K<n<100K0 likes89 downloads18d agoHugging Face19whoashish115 /Moonfrost-Persona-SFT Moonfrost-Persona-SFT Code · Site · Training runs 140,000 multi-turn conversations that teach a small chat model two things no public dataset covers: who it is, and that "you" and "I" refer to different people. They were written for the Moonfrost-777M-Instruct-v2 fine-tune, where they made up 3.3% of the rows, and they are built from templates rather than generated by another model, so there is no scraped text and nothing from anyone else's outputs in them. Regenerating the… See the full description on the dataset page: https://huggingface.co/datasets/whoashish115/Moonfrost-Persona-SFT.texttext-generation100K<n<1M1 likes87 downloads11d agoHugging Face20Below-Image /Open-Personix Open-Personix Dataset Summary Open-Personix is a structured JSON dataset maintained under Poralus. The dataset is primarily text and metadata: each record contains a relative image path, a natural-language caption, and descriptive annotation fields for a person-centered sample. The dataset is designed for workflows such as: caption generation and caption analysis text-based filtering over person annotations metadata-aware retrieval and evaluation multimodal experiments… See the full description on the dataset page: https://huggingface.co/datasets/Below-Image/Open-Personix.texttext-generationn<1K6 likes78 downloads7mo agoHugging Face21wick1d /Personalized_Safety_Data 📦 Personalized Risk and Dilemma Dataset for LLM Safety Research 📝 Dataset Summary This is the first dataset designed to support research on personalized risk and emotional vulnerability in the context of Large Language Models (LLMs). The dataset contains 8,000+ real-world, anonymized personal queries, extracted from Reddit and annotated with structured profile metadata, including emotional states, demographic information, and life contexts (e.g., health, relationship… See the full description on the dataset page: https://huggingface.co/datasets/wick1d/Personalized_Safety_Data.textquestion-answering1K<n<10K4 likes76 downloads1y agoHugging Face22PinkPixel /personality-sarcastic-humor _____ _ _ _____ _ _ | __ (_) | | | __ (_) | | | |__) | _ __ | | __ | |__) |__ _____| | | ___/ | '_ \| |/ / | ___/ \ \/ / _ \ | | | | | | | | < | | | |> < __/ | |_| |_|_| |_|_|\_\ |_| |_/_/\_\___|_| 🎨 Pink Pixel: Sarcastic, Witty, and Snarky Personality Dataset 🎭 Welcome to the Pink Pixel Sarcastic Humor dataset! This dataset is meticulously crafted to help you fine-tune… See the full description on the dataset page: https://huggingface.co/datasets/PinkPixel/personality-sarcastic-humor.texttext-generation1K<n<10K4 likes75 downloads5mo agoHugging Face23rcds /wikipedia-persons-masked wikipedia persons masked: A filtered version of the wikipedia dataset, with only pages of people Dataset Summary Contains ~70k pages from wikipedia, each describing a person. For each page, the person described in the text is masked with a Supported Tasks and Leaderboards The dataset supports the tasks of fill-mask, but can also be used for other tasks such as question answering, e.g. "Who is Languages english only Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/rcds/wikipedia-persons-masked.textfill-mask10K<n<100K3 likes74 downloads4y agoHugging Face24apol /demoverse-personas-es-v1 Dataset Card for DemoVerse Personas ES v1 Resumen del dataset demoverse-personas-es-v1 es un dataset de 100.000 personas sinteticas en espanol para Espana, disenado como artefacto publico y como capa operativa para simulacion sociológica. El dataset se inspira metodologicamente en nvidia/Nemotron-Personas-France, pero no reutiliza sus filas ni intenta replicar la poblacion francesa. La adaptacion reescribe el marco para Espana, con clivajes territoriales, sistema de… See the full description on the dataset page: https://huggingface.co/datasets/apol/demoverse-personas-es-v1.tabulartext-classification100K<n<1M0 likes72 downloads6mo agoHugging Face25PersonalAILab /O-Researcher-SFT-Dataset Data Introduction This dataset serves as the core training data for O-Researcher, specifically designed to elicit end-to-end, multi-turn, multi-tool deep research capabilities in large language models. Built on the Multi-Agent Data Synthesis paradigm, the dataset leverages collaborative AI agents to simulate complex tool-integrated reasoning, transforming multi-agent research workflows into trajectory data suitable for supervised fine-tuning (SFT), enabling dynamic web search, page… See the full description on the dataset page: https://huggingface.co/datasets/PersonalAILab/O-Researcher-SFT-Dataset.text1K<n<10K2 likes71 downloads9mo agoHugging Face26walter-bd /small-persona-dataset small-persona-dataset Bilingual (EN/ES) training dataset for a small NPC voice model. The model learns to take a plain factual sentence and rewrite it in a character's voice, conditioned on persona parameters. Task INPUT: TONE:grumpy STYLE:blunt HUMOR:dry RELATION:rival ROLE:blacksmith FACT: Iron swords cost 15 gold. OUTPUT: Fifteen gold. Still overpriced for your work. The model is conditioned on 5 parameters: tone, style, humor, role, and relation.… See the full description on the dataset page: https://huggingface.co/datasets/walter-bd/small-persona-dataset.texttext-generation100K<n<1M0 likes70 downloads2mo agoHugging Face27dipikakhullar /personalization-reddit-user-histories personalization-reddit-user-histories Per-user chronological histories of answered questions across all subreddits. Derived from dipikakhullar/personalization-reddit: every (query, preferred_answer) pair a user authored as OP, grouped by user and sorted by time, slimmed to the four fields needed to model a user's timeline. Each record is one user. Users with a single interaction are dropped (a timeline needs more than one point). Selection: seen-the-top… See the full description on the dataset page: https://huggingface.co/datasets/dipikakhullar/personalization-reddit-user-histories.text100K<n<1M0 likes70 downloads3mo agoHugging Face28amalia-llm /persona_nemotron Persona Nemotron PT Datasets This is a collection of Portuguese synthetic datasets, consisting of 3 datasets, one with general questions from varied topics, one with math questions, and one with instruction-following requests. The prompts were generated using an approach similar to PersonaHub, with a translated version of Nemotron Personas. Both prompts and answers were generated using Gemma 3-27B. This dataset is provided as part of the AMALIA project and is… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/persona_nemotron.textquestion-answering100K<n<1M0 likes68 downloads3mo agoHugging Face29ehejin /user_study-preference-personalized_0423_base_filtered Filtered user study dataset Source repo: ehejin/user_study-preference-personalized_0423_base Each row is ONE item review (pre-rating, conversation, post-rating). Submission-level fields (prolific_pid, demographics, background) are duplicated across rows that share a submission. The 25-50 rows here are the FIRST review for each unique pool index, selected the same way the analysis plot uses — see scripts/plot_vote_shift_3way.py. Total rows: 50 tabularn<1K0 likes68 downloads5mo agoHugging Face30demegire /personaplex-finetuning-pharma-data-sample PersonaPlex Finetuning — Pharma Data Sample A 10-example slice of the synthetic patient-support / medication adherence dataset used to train demegire/personaplex-finetune-pharma. The on-disk layout below is exactly what the trainer in emotion-machine-org/personaplex-finetune consumes — use this as a template when building your own. Split: 8 train / 2 eval (mirrors the upstream 2003 / 20 split at sample scale). Layout . ├── adhery_v2.jsonl # master… See the full description on the dataset page: https://huggingface.co/datasets/demegire/personaplex-finetuning-pharma-data-sample.audiotext-to-speechn<1K0 likes68 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.