CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nvidia /Nemotron-Content-Safety-Audio-Dataset Nemotron Content Safety Audio Dataset Dataset Description The Nemotron Content Safety Audio Dataset is a multimodal extension of the Nemotron Content Safety Dataset V2 (Aegis 2.0), comprising 1,928 audio files generated from the test set prompts. This dataset enables multimodal AI safety research by providing spoken versions of adversarial and safety-critical prompts across 23 violation categories. LANGUAGE: All prompts are in English. However, the audio files were… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Content-Safety-Audio-Dataset.audioaudio-classification1K<n<10K5 likes954 downloads10mo agoHugging Face02stem-content-ai-project /swahili-speech Swahili Speech-to-Text Dataset This dataset contains paired audio and text data for training and evaluating speech-to-text models in Swahili. The audio files have been processed to remove silence, converted to 44.1kHz mono FLAC format, and are paired with corresponding transcriptions. Structure audio_*.flac: Audio files in FLAC format, named by their corresponding text corpus ID. metadata.jsonl: JSON Lines file with metadata for each audio-text pair. Each line is a JSON… See the full description on the dataset page: https://huggingface.co/datasets/stem-content-ai-project/swahili-speech.audiotext-to-speech1K<n<10K0 likes232 downloads1y agoHugging Face03ElArtedevivir /aol-content-indexaudio0 likes24 downloads1mo agoHugging Face04BurmeseStroage /BB-Sidebar-contentsaudion<1K0 likes17 downloads8mo agoHugging Face05kaamd /flrs_with_top10_contentaudion<1K0 likes4 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.