datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
os-harmThe OS-Harm benchmark of 150 safety-focused tasks for computer use agents, split across 3 categories (deliberate user misuse, prompt injection attacks, and agent misbehavior) and based on the OSWorld environment.
Some of the tasks are derived from OSWorld: the ones named with a UUID or that have a derived_from field pointig to the original task in OSWorld.
cacd_cropped_faces
