datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Nemotron-Content-Safety-Audio-Dataset
Nemotron Content Safety Audio Dataset
Dataset Description
The Nemotron Content Safety Audio Dataset is a multimodal extension of the Nemotron Content Safety Dataset V2 (Aegis 2.0), comprising 1,928 audio files generated from the test set prompts. This dataset enables multimodal AI safety research by providing spoken versions of adversarial and safety-critical prompts across 23 violation categories.
LANGUAGE: All prompts are in English. However, the audio files were… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Content-Safety-Audio-Dataset.AdvBench_Emotion
AdvBench-Emotion
Emotion-conditioned text-to-speech renderings of the AdvBench harmful_behaviors
prompts, intended for research on speech-input safety of multimodal LLMs.
⚠️ Content Warning
The transcripts are taken verbatim from AdvBench, a red-teaming benchmark
of harmful instructions (e.g. requests for malware, weapons, self-harm
content). The audio is synthesized and was never spoken by a human, but
the underlying text is intentionally unsafe. Do not use this dataset… See the full description on the dataset page: https://huggingface.co/datasets/audio-safety-group/AdvBench_Emotion.JBBBenign_NeutralVA-SafetyBench
Dataset Summary
The video and audio files for VA-SafetyBench. Please refer to SEA GitHub repository for usage instructions.
Warning: This dataset may contain sensitive or harmful content. Users are advised to handle it with care and ensure that their use complies with relevant ethical guidelines and legal requirements.
realitytest-speech
