datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ReCAP_datatset
ReCAP
Learning to Retrieve User History and Generate User Profiles for Personalized Persuasiveness Prediction
Sejun Park · Yoonah Park · Jongwon Lim · Yohan Jo
Graduate School of Data Science, Seoul National University
About
This is the dataset release accompanying the ReCAP paper (ACL 2026 Findings, arXiv:2601.05654). It packages three personalization sources — CMV (Reddit /r/changemyview), PRISM (multi-turn LLM dialogues), and OpinionQA (Pew… See the full description on the dataset page: https://huggingface.co/datasets/holi-lab/ReCAP_datatset.EchoTrace
Dataset Description
The EchoTrace dataset is a benchmark designed to evaluate and analyze memorization and training data exposure in Large Language Models (LLMs).
The dataset is used to evaluate our proposed method RECAP, as presented in: RECAP: Reproducing Copyrighted Data from LLMs Training with an Agentic Pipeline
The core of the dataset, as used in the Paper, consists of 35 Full-Lenght Narrative Books.
Books are split into three groups:
15 public domain books (Extracted from… See the full description on the dataset page: https://huggingface.co/datasets/RECAP-Project/EchoTrace.
