datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Twin-2K-500
Twin-2K-500 Dataset
This dataset Twin-2K-500 contains comprehensive persona information from a representative sample of 2,058 US participants, providing rich demographic and psychological data. The dataset is specifically designed for building digital twins for LLM simulations.
More information on how to use this dataset can be found in our Documentation and GitHub repository.
Details on how the dataset was generated are available in our Paper.
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/LLM-Digital-Twin/Twin-2K-500.Twin-2K-500-Mega-Study
Twin-2K-500-Mega-Study Dataset
GitHub Repository: https://github.com/TianyiPeng/Twin-2K-500-Mega-Study
To see more details for how to process these data, please refer to this GitHub repository.
This dataset contains survey data from the Twin-2K-500 Mega Study, which tests the validity of using large language models to predict people's future answers based on their answers to past surveys (creating "digital twins" of participants).
Dataset Structure
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/LLM-Digital-Twin/Twin-2K-500-Mega-Study.llm-domain-specific-tough-questions
LLM-Tough-Questions Dataset
Description
The LLM-Tough-Questions dataset is a synthetic collection designed to rigorously challenge and evaluate the capabilities of large language models. Comprising 10 meticulously formulated questions across 100 distinct domains, this dataset spans a wide spectrum of specialized fields including Mathematics, Fluid Dynamics, Neuroscience, and many others. Each question is developed to probe deep into the intricacies and subtleties of the… See the full description on the dataset page: https://huggingface.co/datasets/YAV-AI/llm-domain-specific-tough-questions.llm-delusion-response-annotations
LLM Delusion-Like Belief Reinforcement Annotations
This dataset contains human annotations of responses generated by conversational large language models (LLMs) to prompts expressing potentially delusion-like or reality-distorted beliefs.
The purpose of the dataset is to support evaluation of whether conversational LLM responses may unintentionally reinforce or strengthen delusion-like beliefs.
Dataset Files
Consensus Dataset… See the full description on the dataset page: https://huggingface.co/datasets/vennu95/llm-delusion-response-annotations.llm-delusion-response-annotations
LLM Delusion-Like Belief Reinforcement Annotations
This dataset contains human annotations of responses generated by conversational large language models (LLMs) to prompts expressing potentially delusion-like or reality-distorted beliefs.
The purpose of the dataset is to support evaluation of whether conversational LLM responses may unintentionally reinforce or strengthen delusion-like beliefs.
Dataset Files
Consensus Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ManjuKrish/llm-delusion-response-annotations.
