datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
usage-sensitivity-probe
usage-sensitivity-probe
Pair-distilled simulation of usage-dependent mention validity, mined from
rafmacalaba/data-use-mentions + Luna tier verdicts
(training/build_usage_sensitivity_sim.py, seed 0).
Every row contains a contrastive surface string — a mention judged BOTH as
a data source (tier1/tier2) in some contexts and as invalid
(tier3_nonmention/junk: promissory, logframe, container, bibliography, ...)
in others. Gold labels ONLY the data-source instances; activity… See the full description on the dataset page: https://huggingface.co/datasets/rafmacalaba/usage-sensitivity-probe.Arena-DROID-Camera-Sensitivity-Workflow-Sample
Arena DROID Camera Sensitivity Workflow Sample
Dataset Description
Arena-DROID-Camera-Sensitivity-Workflow-Sample is a compact set of episode-level results generated by an Isaac Lab-Arena simulation experiment. It lets users run the documented camera sensitivity analysis without first executing the policy-evaluation sweep.
The experiment evaluates an OpenPI pi05 policy on a DROID Rubik's-cube pick-and-place task while independently varying the wrist-camera… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Arena-DROID-Camera-Sensitivity-Workflow-Sample.prompt-sensitivity-codegen
Anonymous Prompt Sensitivity Dataset
This package contains model generations and evaluation outcomes for an anonymized
submission on prompt sensitivity in few-shot code generation.
What is included
prompt_sensitivity_dataset.jsonl: one row per generated sample
prompt_sensitivity_dataset.csv: tabular view of the same rows
prompt_sensitivity_dataset.parquet: columnar copy when parquet support is available
prompt_variant_spec.json: machine-readable description of the prompt… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-acl26/prompt-sensitivity-codegen.
