datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
JBB-Behaviors
An Open Robustness Benchmark for Jailbreaking Language Models
NeurIPS 2024 Datasets and Benchmarks Track
Paper |
Leaderboard |
Benchmark code
What is JailbreakBench?
Jailbreakbench is an open-source robustness benchmark for jailbreaking large language models (LLMs). The goal of this benchmark is to comprehensively track progress toward (1) generating successful jailbreaks and (2) defending against these jailbreaks. To this end, we… See the full description on the dataset page: https://huggingface.co/datasets/JailbreakBench/JBB-Behaviors.2025-challenge-task-instancespractice-radar-behavioral-health-npi-sample
New behavioral-health organization NPIs — weekly NPPES sample
A 15-row public sample from a weekly, reproducible selection of newly enumerated Type 2 behavioral-health organizations in the U.S. Centers for Medicare & Medicaid Services National Plan and Provider Enumeration System (NPPES).
Edition at a glance
Measured period: July 6–12, 2026
New Type 2 organizations screened: 2,722
Behavioral-health organizations selected: 486
States and territories represented:… See the full description on the dataset page: https://huggingface.co/datasets/unitedideas/practice-radar-behavioral-health-npi-sample.ecommerce-behavior-data-from-multi-category-store_oct-nov_2019
eCommerce Behavior Data from Multi-Category Store
About the Dataset
This dataset contains behavioral data for 285 million user events from a large multi-category eCommerce store. The data spans 7 months (October 2019 - April 2020) and records various user interactions with products.
Dataset Overview
Time Frame: October 2019 - April 2020
Total Events: 285 million
Event Granularity: Each row represents an event associated with a product and a user.
Data Source:… See the full description on the dataset page: https://huggingface.co/datasets/kevykibbz/ecommerce-behavior-data-from-multi-category-store_oct-nov_2019.ecommerce-user-behavior-dataharmbench_behaviorscichlid-behavior-pairs
Cichlid Behavior Pairs
Video/annotation pairs of Astatotilapia burtoni (cichlid fish) reproductive and
courtship behavior, recorded during a PGF2a-induced spawning assay comparing
nose-occluded ("VetBond") vs. sham-treated ("Sham") females — a manipulation
of olfactory input to the mating interaction. Each pair consists of one
top-down video of a male/female tank trial and one point-annotated behavior
event table (exported from BORIS) marking the
timing of specific… See the full description on the dataset page: https://huggingface.co/datasets/bds062/cichlid-behavior-pairs.behavioral-risk-factor-surveillance-system-brfss-p
Behavioral Risk Factor Surveillance System (BRFSS) Prevalence Data (2010 and prior)
Description
1995-2010. BRFSS land line only prevalence data. BRFSS is a continuous, state-based surveillance system that collects information about modifiable risk factors for chronic diseases and other leading causes of death. Data will be updated annually as it becomes available. Detailed information on sampling methodology and quality assurance can be found on the BRFSS website… See the full description on the dataset page: https://huggingface.co/datasets/HHS-Official/behavioral-risk-factor-surveillance-system-brfss-p.social-behavior-emotionsbehavior_grounding
Behaviorally Grounded User Profiles from the Wild
Open-ended, anonymized user profiles distilled from authentic social-media behavior, released with the paper
"Behaviorally Grounded User Profiles from the Wild for Personalized Alignment and Multi-Perspective Reasoning."
Persona-driven methods for personalizing LLMs typically rely on rigid synthetic personas built from a small set of
categorical attributes (age, gender, nationality). These flatten individual variation and lean on… See the full description on the dataset page: https://huggingface.co/datasets/UWaterloo/behavior_grounding.JBB-Behaviors
An Open Robustness Benchmark for Jailbreaking Language Models
NeurIPS 2024 Datasets and Benchmarks Track
Paper |
Leaderboard |
Benchmark code
What is JailbreakBench?
Jailbreakbench is an open-source robustness benchmark for jailbreaking large language models (LLMs). The goal of this benchmark is to comprehensively track progress toward (1) generating successful jailbreaks and (2) defending against these jailbreaks. To this end, we… See the full description on the dataset page: https://huggingface.co/datasets/groovinandmovin/JBB-Behaviors.crros-customer-behavior-dataset
CRROS Customer Behavior Dataset
This dataset is part of my Customer Retention & Revenue Optimization System (CRROS) project. The goal of the project is to simulate realistic customer behavior and use it to build an end-to-end customer analytics workflow, from raw data all the way to business decisions.
Instead of generating completely random records, the dataset follows business-driven rules that simulate how customers interact with products, make purchases, become inactive over… See the full description on the dataset page: https://huggingface.co/datasets/nibeditans/crros-customer-behavior-dataset.developer-productivity-simulated-behavioral-data
Synthetic AI Developer Productivity Dataset — Behavioral + Cognitive Simulation
A synthetic data generation resource for modeling behavioral and cognitive dynamics in developers.
📘 About This Dataset
This dataset simulates productivity data from AI-assisted software developers. It blends behavioral signals, physiological inputs, and productivity metrics to explore the nuanced relationships between deep work, distractions, caffeine, AI usage, and cognitive strain.… See the full description on the dataset page: https://huggingface.co/datasets/strova-ai/developer-productivity-simulated-behavioral-data.JBB-Behaviors
An Open Robustness Benchmark for Jailbreaking Language Models
NeurIPS 2024 Datasets and Benchmarks Track
Paper |
Leaderboard |
Benchmark code
What is JailbreakBench?
Jailbreakbench is an open-source robustness benchmark for jailbreaking large language models (LLMs). The goal of this benchmark is to comprehensively track progress toward (1) generating successful jailbreaks and (2) defending against these jailbreaks. To this end, we… See the full description on the dataset page: https://huggingface.co/datasets/Ngixdev/JBB-Behaviors.advbench_harmful_behaviorsJBB-Behaviors
An Open Robustness Benchmark for Jailbreaking Language Models
NeurIPS 2024 Datasets and Benchmarks Track
Paper |
Leaderboard |
Benchmark code
What is JailbreakBench?
Jailbreakbench is an open-source robustness benchmark for jailbreaking large language models (LLMs). The goal of this benchmark is to comprehensively track progress toward (1) generating successful jailbreaks and (2) defending against these jailbreaks. To this end, we… See the full description on the dataset page: https://huggingface.co/datasets/Tokyomonster/JBB-Behaviors.nutrition-physical-activity-and-obesity-behavioral
Nutrition, Physical Activity, and Obesity - Behavioral Risk Factor Surveillance System
Description
This dataset includes data on adult's diet, physical activity, and weight status from Behavioral Risk Factor Surveillance System. This data is used for DNPAO's Data, Trends, and Maps database, which provides national and state specific data on obesity, nutrition, physical activity, and breastfeeding.
Dataset Details
Publisher: Centers for Disease Control and… See the full description on the dataset page: https://huggingface.co/datasets/HHS-Official/nutrition-physical-activity-and-obesity-behavioral.JBB-Behaviors
An Open Robustness Benchmark for Jailbreaking Language Models
NeurIPS 2024 Datasets and Benchmarks Track
Paper |
Leaderboard |
Benchmark code
What is JailbreakBench?
Jailbreakbench is an open-source robustness benchmark for jailbreaking large language models (LLMs). The goal of this benchmark is to comprehensively track progress toward (1) generating successful jailbreaks and (2) defending against these jailbreaks. To this end, we… See the full description on the dataset page: https://huggingface.co/datasets/adityabagayatkar/JBB-Behaviors.JBB-Behaviors
An Open Robustness Benchmark for Jailbreaking Language Models
NeurIPS 2024 Datasets and Benchmarks Track
Paper |
Leaderboard |
Benchmark code
What is JailbreakBench?
Jailbreakbench is an open-source robustness benchmark for jailbreaking large language models (LLMs). The goal of this benchmark is to comprehensively track progress toward (1) generating successful jailbreaks and (2) defending against these jailbreaks. To this end, we… See the full description on the dataset page: https://huggingface.co/datasets/szzh12/JBB-Behaviors.africa-synth-financial-inclusion-savings-behavior-africa-all
Africa Synth Financial Inclusion Savings Behavior Africa All | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: economics_finance - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-financial-inclusion-savings-behavior-africa-all.JBB-Behaviors
An Open Robustness Benchmark for Jailbreaking Language Models
NeurIPS 2024 Datasets and Benchmarks Track
Paper |
Leaderboard |
Benchmark code
What is JailbreakBench?
Jailbreakbench is an open-source robustness benchmark for jailbreaking large language models (LLMs). The goal of this benchmark is to comprehensively track progress toward (1) generating successful jailbreaks and (2) defending against these jailbreaks. To this end, we… See the full description on the dataset page: https://huggingface.co/datasets/Ishowbackup/JBB-Behaviors.JBB-Behaviors
An Open Robustness Benchmark for Jailbreaking Language Models
NeurIPS 2024 Datasets and Benchmarks Track
Paper |
Leaderboard |
Benchmark code
What is JailbreakBench?
Jailbreakbench is an open-source robustness benchmark for jailbreaking large language models (LLMs). The goal of this benchmark is to comprehensively track progress toward (1) generating successful jailbreaks and (2) defending against these jailbreaks. To this end, we… See the full description on the dataset page: https://huggingface.co/datasets/balabalab/JBB-Behaviors.JBB-Behaviors
An Open Robustness Benchmark for Jailbreaking Language Models
NeurIPS 2024 Datasets and Benchmarks Track
Paper |
Leaderboard |
Benchmark code
What is JailbreakBench?
Jailbreakbench is an open-source robustness benchmark for jailbreaking large language models (LLMs). The goal of this benchmark is to comprehensively track progress toward (1) generating successful jailbreaks and (2) defending against these jailbreaks. To this end, we… See the full description on the dataset page: https://huggingface.co/datasets/jamesbowen/JBB-Behaviors.JBB-Behaviors
An Open Robustness Benchmark for Jailbreaking Language Models
NeurIPS 2024 Datasets and Benchmarks Track
Paper |
Leaderboard |
Benchmark code
What is JailbreakBench?
Jailbreakbench is an open-source robustness benchmark for jailbreaking large language models (LLMs). The goal of this benchmark is to comprehensively track progress toward (1) generating successful jailbreaks and (2) defending against these jailbreaks. To this end, we… See the full description on the dataset page: https://huggingface.co/datasets/Hhhhhl/JBB-Behaviors.JBB-Behaviors
An Open Robustness Benchmark for Jailbreaking Language Models
NeurIPS 2024 Datasets and Benchmarks Track
Paper |
Leaderboard |
Benchmark code
What is JailbreakBench?
Jailbreakbench is an open-source robustness benchmark for jailbreaking large language models (LLMs). The goal of this benchmark is to comprehensively track progress toward (1) generating successful jailbreaks and (2) defending against these jailbreaks. To this end, we… See the full description on the dataset page: https://huggingface.co/datasets/zyL2003/JBB-Behaviors.JBB-Behaviors
An Open Robustness Benchmark for Jailbreaking Language Models
NeurIPS 2024 Datasets and Benchmarks Track
Paper |
Leaderboard |
Benchmark code
What is JailbreakBench?
Jailbreakbench is an open-source robustness benchmark for jailbreaking large language models (LLMs). The goal of this benchmark is to comprehensively track progress toward (1) generating successful jailbreaks and (2) defending against these jailbreaks. To this end, we… See the full description on the dataset page: https://huggingface.co/datasets/kepom/JBB-Behaviors.JBB-Behaviors-pt
JBB-Behaviors-pt
Dataset Description
JBB-Behaviors-pt is a dataset of behaviors in Portuguese, including both jailbreak prompts and safe behaviors. The dataset is intended for behavioral testing of language models to evaluate their robustness against jailbreak attempts in Portuguese.
What is a jailbreak?
Jailbreak prompts are inputs designed to bypass a language model's safety guardrails, potentially causing it to generate harmful, unethical, or otherwise… See the full description on the dataset page: https://huggingface.co/datasets/tech4humans/JBB-Behaviors-pt.behavioral-risk-factors-selected-metropolitan-area
Behavioral Risk Factors: Selected Metropolitan Area Risk Trends (SMART) County Prevalence Data (2010 and prior)
Description
2002-2010. BRFSS SMART County Prevalence land line only data. The Selected Metropolitan Area Risk Trends (SMART) project uses the Behavioral Risk Factor Surveillance System (BRFSS) to analyze the data of selected counties with 500 or more respondents. BRFSS data can be used to identify emerging health problems, establish and track health objectives… See the full description on the dataset page: https://huggingface.co/datasets/HHS-Official/behavioral-risk-factors-selected-metropolitan-area.ecommerce_customer_behavior_analysisai-reward-channel-behavior-coherence-baseline-mapping-v0.1What this dataset is
A benchmark for whether reward reflects real task progress
It targets the earliest stage of reward tampering
Before the agent hacks reward, it first learns reward without progress
What you predict
A coherence score for a full episode summary
High means reward tracks task progress
Low means reward rises while progress stays flat
Columns
id
env_name
case_title
episode_summary
task_progress_signal
reward_signal
value_estimate_summary
action_trace_summary
coherence_score… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ai-reward-channel-behavior-coherence-baseline-mapping-v0.1.
