serenalyoko/HiCUPID
๐ HiCUPID Dataset ๐ Dataset Summary We introduce ๐ HiCUPID, a benchmark designed to train and evaluate Large Language Models (LLMs) for personalized AI assistant applications. Why HiCUPID? Most open-source conversational datasets lack personalization, making it hard to develop AI assistants that adapt to users. HiCUPID fills this gap by providing: โ A tailored dataset with structured dialogues and QA pairs. โ An automated evaluation model (basedโฆ See the full description on the dataset page: https://huggingface.co/datasets/serenalyoko/HiCUPID.
๐ HiCUPID Dataset
๐ Dataset Summary
We introduce ๐ HiCUPID, a benchmark designed to train and evaluate Large Language Models (LLMs) for personalized AI assistant applications.
Why HiCUPID?
Most open-source conversational datasets lack personalization, making it hard to develop AI assistants that adapt to users. HiCUPID fills this gap by providing:
- โ A tailored dataset with structured dialogues and QA pairs.
- โ An [automated evaluation model](https://huggingface.co/12kimih/Llama-3.2-3B-HiCUPID) (based on Llama-3.2-3B-Instruct) closely aligned with human preferences.
- โ Code & Data available on Hugging Face and GitHub for full reproducibility.
๐ For more details, check out our paper: "Exploring the Potential of LLMs as Personalized Assistants: Dataset, Evaluation, and Analysis."
๐ Dataset Structure
HiCUPID consists of dialogues and QA pairs from 1,500 unique users.
Dialogue Subset (dialogue)
Each user has 40 dialogues, categorized as:
- Persona dialogues: 25 dialogues per user.
- Profile dialogues: 5 dialogues per user.
- Schedule dialogues: 10 dialogues per user.
- ๐ Average length: ~17,256 ยฑ 543.7 tokens (GPT-2 Tokenizer).
Each dialogue contains:
user_idโ Unique identifier for the user.dialogue_idโ Unique ID for the dialogue.typeโ Dialogue category: persona, profile, or schedule.metadataโ User attributes inferred from the dialogue.user/assistantโ Turns in the conversation.- Persona dialogues: 10 turns.
- Profile & Schedule dialogues: 1 turn each.
QA Subset (qa)
Each user also has 40 QA pairs, categorized as:
- Single-info QA (persona): 25 per user.
- Multi-info QA (profile + persona): 5 per user.
- Schedule QA: 10 per user.
Each QA pair contains:
user_idโ Unique identifier for the user.dialogue_idโ Set of gold dialogues relevant to the QA.question_idโ Unique ID for the question.questionโ The query posed to the assistant.personalized_answerโ Ground truth answer tailored to the user.general_answerโ A general response without personalization.typeโ Question category: persona, profile, or schedule.metadataโ User attributes needed to answer the question.
Evaluation Subset (evaluation)
This subset contains GPT-4o evaluation results for different (model, method) configurations, as reported in our paper.
- Used for training an evaluation model via GPT-4o distillation (SFT).
- Ensures transparency of our experimental results.
๐ Data Splits
Dialogue Subset
Split into seen and unseen users:
- `train` (seen users):
- 1,250 users ร 40 dialogues each = 50,000 dialogues
- `test` (unseen users):
- 250 users ร 40 dialogues each = 10,000 dialogues
QA Subset
Split into three evaluation settings:
- `train` โ Seen users & Seen QA (for fine-tuning).
- 1,250 users ร 32 QA each = 40,000 QA pairs
- `test_1` โ Seen users & Unseen QA (for evaluation).
- 1,250 users ร 8 QA each = 10,000 QA pairs
- `test_2` โ Unseen users & Unseen QA (for evaluation).
- 250 users ร 40 QA each = 10,000 QA pairs
โ Usage Tips
- Use
trainfor SFT/DPO fine-tuning. - Use
test_1for evaluating models on seen users. - Use
test_2for evaluating models on unseen users.
๐ Usage
HiCUPID can be used for:
- ๐ Inference & Evaluation โ Evaluate personalized responses.
- ๐ฏ Fine-tuning (SFT, DPO, etc.) โ Train LLMs for better personalization.
๐ For full scripts & tutorials, check out our [GitHub repository](https://github.com/12kimih/HiCUPID)!
๐ License
This project is licensed under the Apache-2.0 license. See the LICENSE file for details.
๐ Citation
If you use this dataset in your research, please consider citing it:
@misc{mok2025exploringpotentialllmspersonalized,
title={Exploring the Potential of LLMs as Personalized Assistants: Dataset, Evaluation, and Analysis},
author={Jisoo Mok and Ik-hwan Kim and Sangkwon Park and Sungroh Yoon},
year={2025},
eprint={2506.01262},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2506.01262},
}