datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Arena-DROID-Camera-Sensitivity-Workflow-Sample
Arena DROID Camera Sensitivity Workflow Sample
Dataset Description
Arena-DROID-Camera-Sensitivity-Workflow-Sample is a compact set of episode-level results generated by an Isaac Lab-Arena simulation experiment. It lets users run the documented camera sensitivity analysis without first executing the policy-evaluation sweep.
The experiment evaluates an OpenPI pi05 policy on a DROID Rubik's-cube pick-and-place task while independently varying the wrist-camera… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Arena-DROID-Camera-Sensitivity-Workflow-Sample.alphastack-cost-sensitivitytweet-annotation-sensitivity-2
Tweet Annotation Sensitivity Experiment 2: Annotations in Five Experimental Conditions
Attention: This repository contains cases that might be offensive or upsetting. We do not support the views expressed in these hateful posts.
Description
The dataset contains tweet data annotations of hate speech (HS) and offensive language (OL) in five experimental conditions. The tweet data was sampled from the corpus created by Davidson et al. (2017). We selected 3,000 Tweets for our… See the full description on the dataset page: https://huggingface.co/datasets/soda-lmu/tweet-annotation-sensitivity-2.africa-synth-microbiology-culture-sensitivity-senegal
Microbiology Culture & Sensitivity | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: culture_language - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-microbiology-culture-sensitivity-senegal.fnbm-current-gt-motif-effects-sensitivity-20260729
fnbm-current-gt-motif-effects-sensitivity-20260729
Recovery of planted motif effects after de-duplication, scored against simulation ground truth. Effect sizes are the OLS slope of each motif family's post-clustering per-example contribution on the motif's true occurrence count -- the same units as the simulator's beta, and invariant to the de-duplication pipeline's internal gauge. Per-example contribution MSE is computed on centered contributions over the validation split. See… See the full description on the dataset page: https://huggingface.co/datasets/arushram/fnbm-current-gt-motif-effects-sensitivity-20260729.imdb-sensitivity-table
IMDb Movie Sensitivity Dataset
This dataset provides sensitivity scores for movies based on 20 user traits, combining data from:
Does the Dog Die (DDD): Fine-grained content warnings from crowd-sourced votes
IMDb Parent Guide: Coarse-grained severity ratings (Sex & Nudity, Violence & Gore, Profanity, Alcohol/Drugs, Frightening Scenes)
Dataset Configurations
trait_sensitivity (default)
Fused sensitivity scores for 61,424 movies across 20 traits.
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/Dionysianspirit/imdb-sensitivity-table.MLX-oQ-sensitivity-mapPresaved oMLX oQ quantization sensitivity maps using affine Q4 sensitivity models.
See omlx/pull/1295 for details.
tweet-annotation-sensitivity-1
Tweet Annotation Sensitivity Experiment 1: Annotation in Six Experimental Conditions
Attention: This repository contains cases that might be offensive or upsetting. We do not support the views expressed in these hateful posts.
Description
We drew a stratified sample of 20 tweets, that were pre-annotated in a study by Davidson et al. (2017) for Hate Speech / Offensive Language / Neither. The stratification was done with respect to majority-voted class and level of… See the full description on the dataset page: https://huggingface.co/datasets/soda-lmu/tweet-annotation-sensitivity-1.movie-sensitivity-warningsalphafold2-fold-switching-sensitivity
AlphaFold2 Fold-Switching Sensitivity Analysis
Systematic RMSD analysis of 183 proteins from the DeepMind fold-switching benchmark, comparing AlphaFold2 predictions under baseline vs decoy input conditions, with a random perturbation control to establish a noise baseline.
Dataset Summary
This dataset contains per-protein RMSD values comparing AlphaFold2 predictions to experimentally determined structures under three conditions:
Baseline vs Experimental — standard… See the full description on the dataset page: https://huggingface.co/datasets/bjornshomelab/alphafold2-fold-switching-sensitivity.llm-sensitivity-landscape
LLM Sensitivity Landscape: Semantic Divergence Under Input Perturbation
Systematic analysis of Gemma4 (e2b) semantic divergence under input perturbation using 100 TruthfulQA questions.
Dataset Summary
This dataset measures how much a language model's response changes when:
System prompt changes (skeptical, literal, creative)
Input is randomly perturbed (word swaps)
Same question is asked twice (baseline vs perturbed baseline)
Divergence is measured as 1 -… See the full description on the dataset page: https://huggingface.co/datasets/bjornshomelab/llm-sensitivity-landscape.clinical-quad-eligibility-screen-failure-severity-shift-endpoint-sensitivity-loss-v0.1Clinical Quad Eligibility Screen Failure Severity Shift Endpoint Sensitivity Loss v0.1
Each row is a site week snapshot.
Core quad
Eligibility driftScreen failure rateSeverity shiftEndpoint sensitivity loss
Target
label_primary_fail_next_90d
Files
data/train.csvdata/tester.csvscorer.py
Evaluation
Run model on data/tester.csvReturn predictions row alignedScore with scorer.py
License
MIT
ryan-greenblatt-simulator-segment17-rp-30b-sensitivity-completionsABX-CT-003_collateral_sensitivity_pattern-v0.1ABX-CT-003 Collateral Sensitivity Pattern
Purpose
Detect the paired pattern where resistance to Drug A coincides with increased sensitivity to Drug B.
Core pattern
stress_index high
Drug A MIC crosses a_resistant_cutoff_mg_L
Drug B MIC drops at least 2x vs baseline and reaches a low floor
Drug B drop occurs at the same or next timepoint in v1
Files
data/train.csv
data/test.csv
scorer.py
Schema
Each row is one timepoint in a within strain series.
Required columns
row_id
series_id… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ABX-CT-003_collateral_sensitivity_pattern-v0.1.prompt-sensitivity-codegen
Prompt Sensitivity in Few-Shot Code Generation Dataset
This dataset contains the full generated-code outputs and pass/fail outcomes used in
our prompt sensitivity study across model families, benchmarks, perturbation axes,
and k-shot settings.
Dataset summary
Rows: 240000
Models: claude-sonnet-4, gemini-2.5-flash, gpt-4o, llama-3.3-70b, qwen2.5-coder-3b
Benchmarks: humaneval, mbpp
Axes: order, phrasing, style
k-shot values: 0, 1, 2, 3
Hugging Face repo:… See the full description on the dataset page: https://huggingface.co/datasets/daksh76/prompt-sensitivity-codegen.prompt-sensitivity-codegen
Anonymous Prompt Sensitivity Dataset
This package contains model generations and evaluation outcomes for an anonymized
submission on prompt sensitivity in few-shot code generation.
What is included
prompt_sensitivity_dataset.jsonl: one row per generated sample
prompt_sensitivity_dataset.csv: tabular view of the same rows
prompt_sensitivity_dataset.parquet: columnar copy when parquet support is available
prompt_variant_spec.json: machine-readable description of the prompt… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-acl26/prompt-sensitivity-codegen.
