datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
police-scanner-audio
Police Scanner Audio Dataset
A comprehensive collection of police and emergency services radio communications from multiple US cities, captured from publicly available scanner feeds.
Dataset Overview
This dataset contains 103,660 audio recordings totaling 357GB of police scanner audio from 6 different cities across the United States. The recordings span multiple months of continuous monitoring and represent real-world emergency services communications.
Scanner… See the full description on the dataset page: https://huggingface.co/datasets/trentmkelly/police-scanner-audio.audiomc
Audio MultiChallenge: A Multi-Turn Evaluation of Spoken Dialogue Systems on Natural Human Interaction
Audio MultiChallenge is an open-source benchmark to evaluate E2E spoken dialogue systems under natural multi-turn interaction patterns. Building on the text-based MultiChallenge framework, which evaluates Inference Memory, Instruction Retention, and Self Coherence, we introduce a new axis Voice Editing that tests robustness to mid-utterance speech repairs and backtracking. We… See the full description on the dataset page: https://huggingface.co/datasets/ScaleAI/audiomc.scandinavian_faroesescandinavian-100hKoALA
KoALa-Bench: Korean Audio Language Model Benchmark
KoALa-Bench is a comprehensive benchmark for evaluating Large Audio Language Models (LALMs) on Korean speech understanding. It covers six tasks spanning both conventional speech processing and novel speech faithfulness evaluation, designed to test whether models can reason over the acoustic and linguistic content of Korean speech.
Tasks
KoALa-Bench consists of six evaluation tasks organized into two categories.… See the full description on the dataset page: https://huggingface.co/datasets/scailaboratory/KoALA.audio-nr3d-sr3d-scanreferData_voice_AI_human_scam
Vietnamese Deepfake Voice Dataset
Dataset Description
The Vietnamese Deepfake Voice Dataset is a multimodal dataset designed for research on deepfake voice detection and scam call detection in Vietnamese. The dataset contains both authentic human speech and AI-generated speech collected from multiple speech synthesis and voice cloning systems.
The dataset is intended for developing and evaluating machine learning and deep learning models for:
Audio deepfake… See the full description on the dataset page: https://huggingface.co/datasets/vietkemmai/Data_voice_AI_human_scam.scandinavian-25hEmergencyTrafficDetection_Large-Scale-Audio-datasetScarlettJohanssonXiang-Scaramouche-Anime-Discuss-SoulX-Podcast-TTSdsm-muz-scalescam-nonscam-youtube-callsscarchicago-police-scannerscasr_datasetTryvoznaje_scascie_PauseGeminThinkICASSP2024-Acoustic_Scattering_AI-Noninvasive_Object_ClassificationsTony_Montana_ScarfaceTryvoznaje_scascie_PauseGemTh2VarTryvoznaje_scascie_PauseGemThinkscarlatti-d-grandstaff-multimodalScaramouche_IndexTTS2_Ad_Audiogrand-junction-scanner-batch-001scaling_dhivehi_stt
