datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
alphastack-cost-sensitivityfnbm-current-gt-motif-effects-sensitivity-20260729
fnbm-current-gt-motif-effects-sensitivity-20260729
Recovery of planted motif effects after de-duplication, scored against simulation ground truth. Effect sizes are the OLS slope of each motif family's post-clustering per-example contribution on the motif's true occurrence count -- the same units as the simulator's beta, and invariant to the de-duplication pipeline's internal gauge. Per-example contribution MSE is computed on centered contributions over the validation split. See… See the full description on the dataset page: https://huggingface.co/datasets/arushram/fnbm-current-gt-motif-effects-sensitivity-20260729.movie-sensitivity-warningsalphafold2-fold-switching-sensitivity
AlphaFold2 Fold-Switching Sensitivity Analysis
Systematic RMSD analysis of 183 proteins from the DeepMind fold-switching benchmark, comparing AlphaFold2 predictions under baseline vs decoy input conditions, with a random perturbation control to establish a noise baseline.
Dataset Summary
This dataset contains per-protein RMSD values comparing AlphaFold2 predictions to experimentally determined structures under three conditions:
Baseline vs Experimental — standard… See the full description on the dataset page: https://huggingface.co/datasets/bjornshomelab/alphafold2-fold-switching-sensitivity.llm-sensitivity-landscape
LLM Sensitivity Landscape: Semantic Divergence Under Input Perturbation
Systematic analysis of Gemma4 (e2b) semantic divergence under input perturbation using 100 TruthfulQA questions.
Dataset Summary
This dataset measures how much a language model's response changes when:
System prompt changes (skeptical, literal, creative)
Input is randomly perturbed (word swaps)
Same question is asked twice (baseline vs perturbed baseline)
Divergence is measured as 1 -… See the full description on the dataset page: https://huggingface.co/datasets/bjornshomelab/llm-sensitivity-landscape.ryan-greenblatt-simulator-segment17-rp-30b-sensitivity-completionsprompt-sensitivity-codegen
Prompt Sensitivity in Few-Shot Code Generation Dataset
This dataset contains the full generated-code outputs and pass/fail outcomes used in
our prompt sensitivity study across model families, benchmarks, perturbation axes,
and k-shot settings.
Dataset summary
Rows: 240000
Models: claude-sonnet-4, gemini-2.5-flash, gpt-4o, llama-3.3-70b, qwen2.5-coder-3b
Benchmarks: humaneval, mbpp
Axes: order, phrasing, style
k-shot values: 0, 1, 2, 3
Hugging Face repo:… See the full description on the dataset page: https://huggingface.co/datasets/daksh76/prompt-sensitivity-codegen.
