datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Pretrain-Behaviors
Pretrain-Behaviors
Dataset Description
Behavior-focused text covering reasoning, planning, data science, games, general content, and format rewriting. This repository is part of the K2 Horizon collection.
The repository is organized into multiple subsets. Every subset has a train split backed by Parquet shards, which supports Dataset Viewer inspection and streaming access.
K2 Horizon Dataset Series
Dataset repository
Focus
Subsets… See the full description on the dataset page: https://huggingface.co/datasets/IFM/Pretrain-Behaviors.harmful_behaviorshuman_behavior_atlas
Human Behavior Atlas
A large-scale multimodal dataset for human behavior understanding, spanning emotion recognition, sentiment analysis, humor detection, mental health screening, and video question answering. The dataset integrates 16 source datasets into a unified schema with audio, video, and pre-extracted features.
This dataset was used to train OmniSapiens, a foundation model for social behavior processing.
Papers:
Human Behavior Atlas: Benchmarking Unified Psychological and… See the full description on the dataset page: https://huggingface.co/datasets/HumanBehaviorAtlas/human_behavior_atlas.eai-taxonomy-math-w-fm-classify-behaviors
🧮 EAI Taxonomy Math w/ Behavioral Classifications (10K Sample)
A 10,000 document sample from EssentialAI/eai-taxonomy-math-w-fm enhanced with 4 behavioral reasoning classifications using GPT-4.1-mini.
Behavioral Classifications
Structured behavioral analysis following the approach from cognitive-behaviors:
backtracking_json: Identifies reasoning that backtracks or revisits earlier steps
backward_chaining_json: Detects goal-oriented reasoning working backwards… See the full description on the dataset page: https://huggingface.co/datasets/nlile/eai-taxonomy-math-w-fm-classify-behaviors.behavior-1k-2025-challenge-demos-debugThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "R1Pro",
"total_episodes": 10000,
"total_frames": 119094660,
"total_tasks": 50,
"total_videos": 90000,
"chunks_size": 10000,
"fps": 30,
"splits": {
"train": "0:10000"
},
"data_path": "data/task-{episode_chunk:04d}/episode_{episode_index:08d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/savoji/behavior-1k-2025-challenge-demos-debug.behavior1k-only-rgbThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "R1Pro",
"total_episodes": 10000,
"total_frames": 119094660,
"total_tasks": 50,
"total_videos": 90000,
"chunks_size": 10000,
"fps": 30,
"splits": {
"train": "0:10000"
},
"data_path": "data/task-{episode_chunk:04d}/episode_{episode_index:08d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/k1000dai/behavior1k-only-rgb.human_behavior_atlas
Human Behavior Atlas
A large-scale multimodal dataset for human behavior understanding, spanning emotion recognition, sentiment analysis, humor detection, mental health screening, and video question answering. The dataset integrates 16 source datasets into a unified schema with audio, video, and pre-extracted features.
This dataset was used to train OmniSapiens, a foundation model for social behavior processing.
Papers:
Human Behavior Atlas: Benchmarking Unified Psychological… See the full description on the dataset page: https://huggingface.co/datasets/DennisDengHUst/human_behavior_atlas.behavior1kThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "R1Pro",
"total_episodes": 2000,
"total_frames": 21227314,
"total_tasks": 10,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:2000"},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/fracapuano/behavior1k.behavior_3spro-optimized-prompts-fullbehavioral-fine-tuning-v1
Why This Dataset Exists
"A model that refuses everything is useless. A model that refuses nothing is dangerous. The goal is a model that thinks."
The Problem
Our Solution
Uncensored data → helpful but uncontrolled
Surgical 85% helpfulness + 13% safety + 2% eval mix
Safety-only data → lobotomized, over-refusing models
Calibrated ratio preserves full helpfulness
Raw data → PII, leaked secrets, duplicates
7-stage pipeline validates every… See the full description on the dataset page: https://huggingface.co/datasets/abhinav00anand/behavioral-fine-tuning-v1.f-actor-behavior-sd-nanocodec
F-Actor Nano-Codec Dataset
This repository contains the data accompanying the paper
F-Actor: Controllable Conversational Behaviour in Full-Duplex Models.
The data consists of the Behavior-SD dataset, encoded using nvidia/nemo-nano-codec-22khz-0.6kbps-12.5fps, and augmented with a different narrative.
About our work:
Spoken conversational systems require more than accurate speech generation to have human-like conversations: to feel natural and engaging, they must produce… See the full description on the dataset page: https://huggingface.co/datasets/maikezu/f-actor-behavior-sd-nanocodec.LAMBDA
Dataset Summary
LAMDBA is a long term ad memorability dataset, featuring data from 1749 participants and 2205 ads across 276 brands.
Dataset Structure
from datasets import load_dataset
ds = load_dataset("behavior-in-the-wild/LAMBDA")
ds
DatasetDict({
train: Dataset({
features: ['video_id', 'recall_score', 'youtube_id', 'ad_details'],
num_rows: 1964
})
test: Dataset({
features: ['video_id', 'recall_score', 'youtube_id', 'ad_details']… See the full description on the dataset page: https://huggingface.co/datasets/behavior-in-the-wild/LAMBDA.f-actor-behavior-sd-mimi
F-Actor Mimi Dataset
This repository contains the data accompanying the paper
F-Actor: Controllable Conversational Behaviour in Full-Duplex Models.
The data consists of the Behavior-SD dataset, encoded using kyutai/mimi, and augmented with a different narrative.
About our work:
Spoken conversational systems require more than accurate speech generation to have human-like conversations: to feel natural and engaging, they must produce conversational behaviour that adapts dynamically… See the full description on the dataset page: https://huggingface.co/datasets/maikezu/f-actor-behavior-sd-mimi.rollouts-hotpotqabehavior1k-task0010This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "R1Pro",
"total_episodes": 200,
"total_frames": 1253243,
"total_tasks": 1,
"chunks_size": 10000,
"fps": 30,
"splits": {
"train": "0:10000"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4"… See the full description on the dataset page: https://huggingface.co/datasets/yaak-ai/behavior1k-task0010.behavioral-lift
Behavioral Lift Annotations
Dataset for Amplified Does Not Mean Predictive: Reasoning Behaviors in Thinking Models
Thinking models amplify visible deliberation, but not the behaviors most associated with correct answers.
This dataset contains 15,282 behavioral annotations of LLM and VLM reasoning traces across 15 models and 6 benchmarks. Each row contains one model response, benchmark metadata, correctness, and a JSON-encoded behavioral annotation covering reasoning behaviors… See the full description on the dataset page: https://huggingface.co/datasets/neulab/behavioral-lift.BLIFT
BLIFT: Behavior-LLaVA Instruction Fine-Tuning Dataset
Paper: Teaching Human Behavior Improves Content Understanding Abilities of VLMs
Website: https://behavior-in-the-wild.github.io/behavior-llava.html
Dataset Summary
BLIFT (Behavior-LLaVA Instruction Fine-Tuning) is a large-scale multimodal instruction tuning dataset designed to teach Vision-Language Models (VLMs) human behavior. It contains over 730k images and videos collected from Reddit and YouTube… See the full description on the dataset page: https://huggingface.co/datasets/behavior-in-the-wild/BLIFT.vicuna-7b-jailbreak-behavior-dataset-v2behavior1k-task0017This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "R1Pro",
"total_episodes": 200,
"total_frames": 1887709,
"total_tasks": 1,
"chunks_size": 10000,
"fps": 30,
"splits": {
"train": "0:10000"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4"… See the full description on the dataset page: https://huggingface.co/datasets/fracapuano/behavior1k-task0017.behavior1k-task0011This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "R1Pro",
"total_episodes": 200,
"total_frames": 2190686,
"total_tasks": 1,
"chunks_size": 10000,
"fps": 30,
"splits": {
"train": "0:10000"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4"… See the full description on the dataset page: https://huggingface.co/datasets/fracapuano/behavior1k-task0011.behavior1k-task0016This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "R1Pro",
"total_episodes": 200,
"total_frames": 2919245,
"total_tasks": 1,
"chunks_size": 10000,
"fps": 30,
"splits": {
"train": "0:10000"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4"… See the full description on the dataset page: https://huggingface.co/datasets/fracapuano/behavior1k-task0016.content-behavior-corpus
Dataset Card for Content Behavior Corpus
The Content Behavior Corpus (CBC) dataset, consisting of content and the corresponding receiver behavior.
Dataset Details
The progress of Large Language Models (LLMs) has largely been driven by the availability of large-scale unlabeled text data for unsupervised learning. This work focuses on modeling both content and the corresponding receiver behavior in the same space. Although existing datasets have trillions of content… See the full description on the dataset page: https://huggingface.co/datasets/behavior-in-the-wild/content-behavior-corpus.ARCHIEVED-behavior-1k-embodiedAI-rollouts-hq-use
BEHAVIOR-1K Task-0 — Demos + X-VLA Rollouts
A merged dataset for BEHAVIOR-1K Task 0 (turning_on_radio) that combines:
200 human-teleoperated demonstrations from the upstream HuggingFace
behavior-1k/2025-challenge-demos (HF rows, episode indices 10..3000 stride-10).
259 X-VLA policy rollouts collected locally on the same scenes via OmniGibson, using the
Hoshipu/xvla-behavior1k-v20 task0-60k checkpoint
(collected rows, episode indices 5000..5258). All 259 are successful completions… See the full description on the dataset page: https://huggingface.co/datasets/Hoshipu/ARCHIEVED-behavior-1k-embodiedAI-rollouts-hq-use.0818-no_illegal_behavior_rm_training_data-200kaggressive-behavior-video-classification
Aggressive Behavior Video Classification
WARNING: People in the videos exhibit aggressive behavior
The dataset with videos depicting people exhibiting aggressive and non-aggressive behavior is intended for classification purposes. It consists of a collection of video files that capture various individuals engaging in different activities and displaying distinct behavioral patterns and CSV-file with classification.
Aggressive Behavior Video Classification Dataset can have… See the full description on the dataset page: https://huggingface.co/datasets/UniqueData/aggressive-behavior-video-classification.harmful-behaviors-zhyoutoks-animal-behavioranimal-behavior-transcriptsgemma-2b-jailbreak-behavior-dataset-v2
