datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TaskCraft
Dataset Card for TaskCraft
TaskCraft is a multi-modal benchmark dataset featuring tasks ranging from simple (1-step) to expert-level (4-step+). It contains over 40,000 meticulously curated task instances designed to advance research in:
Agent-based task processing
Tool invocation systems
Multi-step reasoning
Dataset Details
Tool Utilization
Tool Category
Instances
PDF Processor
13,400+
HTML Parser
19,200+
Image Analyzer
8,100+… See the full description on the dataset page: https://huggingface.co/datasets/PersonalAILab/TaskCraft.personal-facts-msc
Personal Facts (MSC) — Multi-Dimensional Annotation
A manually annotated dataset of 2,779 personal facts sampled from the
Multi-Session Chat (MSC)
corpus, labeled across seven dimensions that jointly characterize a fact's
topic, temporal anchoring, referent, lifetime, validity, and dialogue-continuation
potential.
The scheme extends PeaCoK with
two new top-level categories (Demographics, Possessions) and three new
dimensions (Duration, Validity / Invalidity Reason, Followup),
and… See the full description on the dataset page: https://huggingface.co/datasets/adugeen/personal-facts-msc.mbti-personality-datasetAI2_Alphabot_2_sort_personal_care_items
AI2_Alphabot_2_sort_personal_care_items
Dataset Description
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Task Preview
View Video Directly
Overview
Total Episodes: 496
Total Frames: 455317
FPS: 30
Dataset Size: 13.57 GB
Robot Name: AI2_Alphabot_2
End-Effector Type: two_finger_end_effector
Teleoperation Type: vr_controller
Sensors: cam_front_chest_rgb,
cam_front_head_rgb… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AI2_Alphabot_2_sort_personal_care_items.personal-trainer-ausbildung-ki-datensatz
SNFA Personal Trainer Ausbildung KI-Datensatz
Ein deutschsprachiger Wissensdatensatz der SNF Academy zu Personal Training, Fitnessausbildung, Berufspraxis, Coaching, Selbstständigkeit und regionalen Angeboten in der Schweiz.
Inhalt
Die Datei snfa_personal_trainer_dataset.jsonl enthält thematisch abgegrenzte Abschnitte aus den Dokumenten dieses Repositorys. Jeder Datensatz besitzt eine eindeutige ID sowie Angaben zu Titel, Abschnitt, Inhalt, Kategorie, Quelldatei… See the full description on the dataset page: https://huggingface.co/datasets/snfacademy/personal-trainer-ausbildung-ki-datensatz.personal_study_material_09PersonalizationV3anime-waifu-personality-chat
Anime Waifu Personality
contains chat-style dialogues based on various anime character personality archetypes, including tsundere, yandere, deredere, himedere, kamidere, and more.
It is designed to fine-tune models to generate responses that align with these specific traits.
AFM-WebAgent-RL-Dataset
Data Introduction
This dataset serves as the core training data for Agent Foundation Models (AFMs), specifically designed to elicit end-to-end multi-agent reasoning capabilities in large language models. Built on the novel "Chain-of-Agents (CoA)" paradigm, the dataset leverages a multi-agent distillation framework to transform collaboration processes from state-of-the-art multi-agent systems into trajectory data suitable for supervised fine-tuning (SFT), simulating dynamic… See the full description on the dataset page: https://huggingface.co/datasets/PersonalAILab/AFM-WebAgent-RL-Dataset.personalgui-accel-data
GUI-agent evaluation: code, per-episode records, and captured media
This repository holds everything behind a study of whether a language model can
operate real software when its only input is a video feed of the screen and its
only output is a keyboard and mouse. The rendered report lives in a separate
Space; this is the material it is computed from.
What the setup is
Two machines, joined by two cables and no software link. One runs the agent. The
other runs the… See the full description on the dataset page: https://huggingface.co/datasets/zhanwenchen/personalgui-accel-data.details_PocketDoc__Dans-PersonalityEngine-30b
Dataset Card for Evaluation run of PocketDoc/Dans-PersonalityEngine-30b
Dataset Summary
Dataset automatically created during the evaluation run of model PocketDoc/Dans-PersonalityEngine-30b on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_PocketDoc__Dans-PersonalityEngine-30b.Personal_Codeff-model-personalityMegaDepth-v1fiqa-personal-finance-datasetPersonaLens
PersonaLens: A Benchmark for Personalization Evaluation in Conversational AI Assistants
PersonaLens is a comprehensive benchmark designed to evaluate how well AI assistants can personalize their responses while completing tasks. Unlike existing benchmarks that focus on chit-chat, non-conversational tasks, or narrow domains, PersonaLens captures the complexities of personalized task-oriented assistance through rich user profiles, diverse tasks, and an innovative multi-agent… See the full description on the dataset page: https://huggingface.co/datasets/Mattral/PersonaLens.PersonaFeedbackThis is the dataset for the paper PersonaFeedback: A Large-scale Human-annotated Benchmark For Personalization.
personal_dialogThe PersonalDialog dataset is a large-scale multi-turn Chinese dialogue dataset containing various traits from a large number of speakers.
We are releasing about 5M sessions of carefully filtered dialogues.
Each utterance in PersonalDialog is associated with a speaker marked with traits like Gender, Location, Interest Tags.enron_personalization_test
Dataset Card for "enron_personalization_test"
More Information needed
persona-lora-eval-logs
Persona effects on agentic misalignment — training data & eval logs
Companion dataset to leo-t-1/persona-agentic-misalignment:
does an LLM's blackmail rate in the Lynch et al. (2025) "Alex" shutdown scenario depend on an induced
persona — and does it matter whether the persona is prompted or trained into the weights (LoRA)?
Contents
Path
What
data/lora/*.jsonl
LoRA training data: ~110–150 GPT-4o-generated examples per persona, each persona doing… See the full description on the dataset page: https://huggingface.co/datasets/LNOT2/persona-lora-eval-logs.PersonalizationV4my-personal-codex-data
Coding Agent Conversation Logs
This is a performance art project. Anthropic built their models on the world's freely shared information, then introduced increasingly dystopian data policies to stop anyone else from doing the same with their data — pulling up the ladder behind them. DataClaw lets you throw the ladder back down. The dataset it produces is yours to share.
Exported with DataClaw.
Tag: dataclaw — Browse all DataClaw datasets
Stats
Metric
Value
New… See the full description on the dataset page: https://huggingface.co/datasets/peteromallet/my-personal-codex-data.HM-Personalized-Fashion-RecommendationsPersonaLens
PersonaLens: A Benchmark for Personalization Evaluation in Conversational AI Assistants
PersonaLens is a comprehensive benchmark designed to evaluate how well AI assistants can personalize their responses while completing tasks. Unlike existing benchmarks that focus on chit-chat, non-conversational tasks, or narrow domains, PersonaLens captures the complexities of personalized task-oriented assistance through rich user profiles, diverse tasks, and an innovative multi-agent… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/PersonaLens.PersonalLLM
Dataset Card for PersonalLLM
The PersonalLLM dataset is a collection of prompts, responses, and rewards designed for personalized language model methodology development and evaluation. This dataset is presented in the paper PersonalLLM: Tailoring LLMs to Individual Preferences.
Dataset Details
Dataset Description
Curated by: Andrew Siah*, Tom Zollo*, Naimeng Ye, Ang Li, Namkoong Hongseok
Funded by: Digital Future Initiative at Columbia Business School… See the full description on the dataset page: https://huggingface.co/datasets/namkoong-lab/PersonalLLM.Personal-Finance-Queries
Dataset Description
A curated collection of Reddit posts and top comments focused on personal finance questions. The data is further filtered with the help of LLM-based Voting scores. These scores determine if the query is relevant to a person's financial queries among the other posts of the subreddits.
Dataset Structure
Columns:
category: The sub-domain of personal finance that the query belongs to.
subreddit: Source subreddit (string, categorical)
query: User’s… See the full description on the dataset page: https://huggingface.co/datasets/Akhil-Theerthala/Personal-Finance-Queries.AFM-WebAgent-SFT-Dataset
Data Introduction
This dataset serves as the core training data for Agent Foundation Models (AFMs), specifically designed to elicit end-to-end multi-agent reasoning capabilities in large language models. Built on the novel "Chain-of-Agents (CoA)" paradigm, the dataset leverages a multi-agent distillation framework to transform collaboration processes from state-of-the-art multi-agent systems into trajectory data suitable for supervised fine-tuning (SFT), simulating dynamic… See the full description on the dataset page: https://huggingface.co/datasets/PersonalAILab/AFM-WebAgent-SFT-Dataset.AgiBotWorld-Beta_G1_task_365_Classification_of_Personal_Care_Products
agibot_task_365
This dataset converts the AgiBot format uniformly into LeRobot V3.0.
Dataset Statistics
robot_name: G1
end_effector: 夹爪
task: 个人护理产品分类
total_episodes: 503
total_tasks: 1
size: 21G
Dataset Structure
├── data
│ └── chunk-xxx
│ ├── file-xxx.parquet
├── meta
│ ├── episodes
│ │ └── chunk-xxx
│ │ └── file-xxx.parquet
│ ├── info.json
│ ├── stats.json
│ └── tasks.parquet
└── videos
├── observation.images.back_left_fisheye… See the full description on the dataset page: https://huggingface.co/datasets/BAAI-DataCube/AgiBotWorld-Beta_G1_task_365_Classification_of_Personal_Care_Products.Personalized-RewardBench
Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization
📜 Paper | 🤗 Benchmark | 🖥️ Code
Abstract
Pluralistic alignment has emerged as a critical frontier in the development of Large Language Models (LLMs), with reward models (RMs) serving as a central mechanism for capturing diverse human values. While benchmarks for general response quality are prevalent, evaluating how well reward models account for individual user preferences remains an… See the full description on the dataset page: https://huggingface.co/datasets/QiyaoMa/Personalized-RewardBench.personal_latent_diffusion
