datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
personal-trainer-ausbildung-ki-datensatz
SNFA Personal Trainer Ausbildung KI-Datensatz
Ein deutschsprachiger Wissensdatensatz der SNF Academy zu Personal Training, Fitnessausbildung, Berufspraxis, Coaching, Selbstständigkeit und regionalen Angeboten in der Schweiz.
Inhalt
Die Datei snfa_personal_trainer_dataset.jsonl enthält thematisch abgegrenzte Abschnitte aus den Dokumenten dieses Repositorys. Jeder Datensatz besitzt eine eindeutige ID sowie Angaben zu Titel, Abschnitt, Inhalt, Kategorie, Quelldatei… See the full description on the dataset page: https://huggingface.co/datasets/snfacademy/personal-trainer-ausbildung-ki-datensatz.Personal-Finance-Queries
Dataset Description
A curated collection of Reddit posts and top comments focused on personal finance questions. The data is further filtered with the help of LLM-based Voting scores. These scores determine if the query is relevant to a person's financial queries among the other posts of the subreddits.
Dataset Structure
Columns:
category: The sub-domain of personal finance that the query belongs to.
subreddit: Source subreddit (string, categorical)
query: User’s… See the full description on the dataset page: https://huggingface.co/datasets/Akhil-Theerthala/Personal-Finance-Queries.personal_dictionary
OpenGloss Dictionary (Word-Level)
Dataset Summary
OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English that integrates lexicographic definitions, encyclopedic context, etymological histories, and semantic relationships in a unified resource.
This dataset provides the words-level view where each record represents one lexeme (word or multi-word expression).
Key Statistics
150,101 lexemes across 150,101 English… See the full description on the dataset page: https://huggingface.co/datasets/caioloures/personal_dictionary.personal-info-unlearning
Synthetic Personal Information Unlearning Dataset
Dataset Description
This dataset is designed for research on large language model (LLM) unlearning in controlled synthetic personal-information settings.
It contains synthetic profiles and question-answer data for four personal attributes:
Year of birth
Blood type
Postcode
Social insurance number
The benchmark provides three forget-set sizes: N = 5, 20, 40.
All personal-profile data are synthetically generated… See the full description on the dataset page: https://huggingface.co/datasets/shichenghu/personal-info-unlearning.Personalized_Safety_Data
📦 Personalized Risk and Dilemma Dataset for LLM Safety Research
📝 Dataset Summary
This is the first dataset designed to support research on personalized risk and emotional vulnerability in the context of Large Language Models (LLMs).
The dataset contains 8,000+ real-world, anonymized personal queries, extracted from Reddit and annotated with structured profile metadata, including emotional states, demographic information, and life contexts (e.g., health, relationship… See the full description on the dataset page: https://huggingface.co/datasets/wick1d/Personalized_Safety_Data.pact-culture-personalization
PACT: Personal-Preference and Cultural-Norm Trade-off
This dataset accompanies Whose Norms? Disentangling Cultural and Personal Alignment in Large Language Models.
Hugging Face repository: Angana192/pact-culture-personalization
PACT contains social scenarios where a cultural expectation and an actor's personal preference are both plausible but may conflict. This release contains only the benchmark scenario instances: no model outputs, no model results, no trace-analysis tables… See the full description on the dataset page: https://huggingface.co/datasets/MichiganNLP/pact-culture-personalization.synth-persona
SynthPersona 1000P Preview
This dataset contains 1,000 synthetic personas, a baseline control persona, and question-answer rows tied to those personas.
Files
dataset_personas.jsonl: 1,001 persona rows.
dataset_qa.jsonl: 788,007 question-answer rows.
implicit_shared_mc_bank.json: 418 shared implicit multiple-choice items.
explicit_shared_mc_bank.json: 57 shared explicit multiple-choice items.
attribute_schema.json: metadata for persona seed attributes.… See the full description on the dataset page: https://huggingface.co/datasets/implicit-personalization/synth-persona.MamaBench
Dataset Card for MamaBench
Dataset Summary
MamaBench is a counterfactual clinical benchmark for evaluating the robustness of large language models on maternal and child health diagnostic reasoning. It consists of 217 counterfactual case pairs, each pairing an original clinical vignette with a systematically perturbed counterfactual variant, designed to test whether a model's diagnostic reasoning is sensitive to clinically meaningful changes rather than relying on… See the full description on the dataset page: https://huggingface.co/datasets/HelpMum-Personal/MamaBench.personal-finance-chatml-dataset
Bilingual Personal Finance ChatML Dataset (EN/ES)
Dataset Description
This dataset is a professionally curated bilingual (English/Spanish) instruction dataset designed for fine-tuning large language models (LLMs) in the domain of personal finance.
It is structured in ChatML format and intended for supervised fine-tuning (SFT), domain adaptation, and financial instruction modeling.
The dataset is created and reviewed from an accounting perspective, ensuring conceptual… See the full description on the dataset page: https://huggingface.co/datasets/williamjmorenor/personal-finance-chatml-dataset.Personal-Finance-Queries
Dataset Description
A curated collection of Reddit posts and top comments focused on personal finance questions. The data is further filtered with the help of LLM-based Voting scores. These scores determine if the query is relevant to a person's financial queries among the other posts of the subreddits.
Dataset Structure
Columns:
category: The sub-domain of personal finance that the query belongs to.
subreddit: Source subreddit (string, categorical)
query:… See the full description on the dataset page: https://huggingface.co/datasets/idleengine/Personal-Finance-Queries.PersonalFinance-CoTR-5K
PersonalFinance-CoTR Dataset (v0.1.0)
[Dataset is Under Active Development]
Note This dataset is being iteratively developed. At the current stage of the dataset, V0.1.0 would be a dataset of ~5k datapoints, that are different user-responses.
Overview
A growing dataset of Chain-of-Thought Responses to personal finance queries asked by users on r/PersonalFinance subreddit.
Status: Early development (10 samples → expanding to 54k)
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/Akhil-Theerthala/PersonalFinance-CoTR-5K.MiCoTApersonal-finance-advice-qa
Personal Finance Advice QA
Real personal-finance questions with concise, actionable answers. Built for the
Adaption Labs AutoScientist Challenge (Personal Finance category).
Rows
3,805
Distinct answers
3,805 (100%)
Duplicate questions
none
Nulls
none
Question length
median 70 words
Answer length
median 77 words (max 95)
Licence
MIT
Category coverage
Questions come from r/personalfinance and r/FinancialPlanning, so they are real… See the full description on the dataset page: https://huggingface.co/datasets/flamiinngo/personal-finance-advice-qa.personal-finance-africa
Personal Finance — African Context Dataset
Instruction-tuning dataset covering African personal finance: mobile money ecosystems (MTN MoMo, Airtel Money, M-Pesa), SACCOs, VSLAs, pension systems (NSSF), taxation (URA, KRA), microfinance, digital lending, insurance, remittances, household budgeting, and investment — grounded via web search and (optionally) local reference documents, generated with gemini-3.1-flash. Focused on financial realities in Uganda, Kenya, Tanzania, Nigeria… See the full description on the dataset page: https://huggingface.co/datasets/gimmy256/personal-finance-africa.witq-personality
WitQ Personality Q&A
Basic personality for witfoo/witq-1.0 model
