datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PersonalLLM
Dataset Card for PersonalLLM
The PersonalLLM dataset is a collection of prompts, responses, and rewards designed for personalized language model methodology development and evaluation. This dataset is presented in the paper PersonalLLM: Tailoring LLMs to Individual Preferences.
Dataset Details
Dataset Description
Curated by: Andrew Siah*, Tom Zollo*, Naimeng Ye, Ang Li, Namkoong Hongseok
Funded by: Digital Future Initiative at Columbia Business School… See the full description on the dataset page: https://huggingface.co/datasets/namkoong-lab/PersonalLLM.personal_dictionary
OpenGloss Dictionary (Word-Level)
Dataset Summary
OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English that integrates lexicographic definitions, encyclopedic context, etymological histories, and semantic relationships in a unified resource.
This dataset provides the words-level view where each record represents one lexeme (word or multi-word expression).
Key Statistics
150,101 lexemes across 150,101 English… See the full description on the dataset page: https://huggingface.co/datasets/caioloures/personal_dictionary.personal-codex-model
Personal Codex Model Training Corpus
Overview
Personal Codex Model Training Corpus is a provenance-aware, repository-level dataset for causal
language modeling, code completion, continued pretraining, and coding assistant adaptation. It is
built from source files present in local Git repository checkouts at a defined collection point.
The dataset prioritizes broad, authentic software-engineering coverage while retaining enough
metadata to audit every emitted… See the full description on the dataset page: https://huggingface.co/datasets/JulianAT/personal-codex-model.personal-query-grocery-and-gourmet-food
Personal Query: Grocery and Gourmet Food
This dataset contains personalized product search queries for the Grocery_and_Gourmet_Food category.
Each record is built from the Personal Query pipeline:
Stage 6 generated correct personalized queries.
Stage 7 injected user-specific error query variants when a matching error pattern was available.
Stage 5 provided the user profile complexity level.
Files
data.jsonl: all correct Stage 6 queries. Rows without Stage 7 error query… See the full description on the dataset page: https://huggingface.co/datasets/xxxxdszz/personal-query-grocery-and-gourmet-food.PersonalizedDeepResearchBenchThis is the dataset for the paper Towards Personalized Deep Research: Benchmarks and Evaluations.
PersonalLLM_Eval
PersonalLLM: A Benchmark for Personalizing LLMs
This dataset, presented in PersonalLLM: Tailoring LLMs to Individual Preferences, focuses on adapting LLMs to individual user preferences. It provides open-ended prompts paired with multiple high-quality responses, allowing for the evaluation of personalization algorithms. The dataset includes diverse user preferences simulated using pre-trained reward models, offering a robust testbed for research in this area.
The data is structured… See the full description on the dataset page: https://huggingface.co/datasets/namkoong-lab/PersonalLLM_Eval.personal-query-baby-products
Personal Query: Baby Products
This dataset contains personalized product search queries for the Baby_Products category.
Each record is built from the Personal Query pipeline:
Stage 6 generated correct personalized queries.
Stage 7 injected user-specific error query variants when a matching error pattern was available.
Stage 5 provided the user profile complexity level.
Files
data.jsonl: all correct Stage 6 queries. Rows without Stage 7 error query keep error_query as… See the full description on the dataset page: https://huggingface.co/datasets/xxxxdszz/personal-query-baby-products.personal-query-pet-supplies
Personal Query: Pet Supplies
This dataset contains personalized product search queries for the Pet_Supplies category.
Each record is built from the Personal Query pipeline:
Stage 6 generated correct personalized queries.
Stage 7 injected user-specific error query variants when a matching error pattern was available.
Stage 5 provided the user profile complexity level.
Files
data.jsonl: all correct Stage 6 queries. Rows without Stage 7 error query keep error_query as… See the full description on the dataset page: https://huggingface.co/datasets/xxxxdszz/personal-query-pet-supplies.personalized-query
Personalized Query
This repository contains three personalized product-search query datasets in one Hugging Face dataset page.
Each config corresponds to one product category:
baby: Baby Products
grocery: Grocery and Gourmet Food
pets: Pet Supplies
Each config has two splits:
full: all correct Stage 6 queries. Rows without Stage 7 error query keep error_query as null.
paired: only rows where a correct query has a paired error query.
Dataset Size
Config… See the full description on the dataset page: https://huggingface.co/datasets/xxxxdszz/personalized-query.Self_Awareness_Personal_Leadership_Content_2
Self-Awareness Personal Leadership Content 2
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Self_Awareness_Personal_Leadership_Content_2.YNTP-100
Dataset Card for annonymous_100
Dataset Summary
The annonymous_100 dataset is a conversation dataset between English, Chinese, and Japanese users and NPCs during a five-day shared house experience game. This dataset consists of responses to questions from NPCs over five days, with 33 English users, 34 Chinese users, and 33 Japanese users.
Language(s)
The dataset contains conversations in English, Chinese, and Japanese.
Dataset Structure
Data… See the full description on the dataset page: https://huggingface.co/datasets/Personalized-Alignment/YNTP-100.Self_Awareness_Personal_Leadership_Content_1
Self-Awareness Personal Leadership Content 1
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Self_Awareness_Personal_Leadership_Content_1.
