datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LAION-Audio-300Mlaion-audio-previewaudiosnippetsaudiosnippets_small_with_detailed_annotationaudiosnippets_small_with_detailed_annotation2audiosnippets_long_2_5Maudiosnippets_long_1Maudiosnippets-cleaned
Dataset Summary
This dataset is a processed version of mitermix/audiosnippets. The dataset contains audio snippets that have been cleaned and resampled, making it suitable for tasks like audio captioning, audio classification, or other audio-based machine learning applications.
Processing Details
Transcriptions and broken characters were removed.
All MP3 audio files were resampled to 16kHz for consistency.
The accompanying JSON metadata was made consistent.
Entries with… See the full description on the dataset page: https://huggingface.co/datasets/mkrausio/audiosnippets-cleaned.MemeEffect-382K-audioWe are releasing the audio files that we have collected from MemeEffect-382K dataset. All the files are being shared as .tar files and files are rnamed using their respective id that can be found through the metadata.
We share these files as-part of research initiative.
audiofolder_webdatasetAudioCapsSpokenWOZ-Train-Audiotalent_plus_rl_groups_of_50_with_audiobox_scoresvoa_myanmar_asr_audio_1
📢 This is the first publicly released ASR-ready Burmese speech dataset with over 1 million audio chunks — a milestone in the history of Myanmar language technology.
Overview
This dataset was created by scraping and segmenting the full archive of the VOA Burmese morning radio program. Out of a total of 3,687 full-length MP3 broadcasts, this release processes 3,267 of them, resulting in approximately 1.8 million sentence-level audio chunks, totaling ~3,267 hours of segmented audio.… See the full description on the dataset page: https://huggingface.co/datasets/freococo/voa_myanmar_asr_audio_1.cmd-audio-dumptimbre-audio-caption-pairsAudioQA-1Msubjective_audio_quality
Balanced Perceptual Audio Quality Dataset
Dataset Summary
This is a large-scale, balanced dataset designed for training models for perceptual audio quality assessment. It consists of 612,020 examples, each containing a pair of 1-second audio clips: a high-quality original and a degraded version processed by various audio codecs. Each pair is accompanied by a perceptual quality score (ranging from 0.0 to 1.0) generated by visqol-like algorithms.
The key feature of this… See the full description on the dataset page: https://huggingface.co/datasets/overfitprolabse/subjective_audio_quality.audioset-with-grounded-captionsaudiosnippets_long_50kspeech_text-tts_audioAudiobook_Noice_V2laion-audio-preview-splitlaion-audio-smallaudioset-with-captionsrohingya_asr_audioThis is the first public Rohingya language ASR dataset in AI history.
Overview
This dataset contains broadcast audio recordings from the Voice of America (VOA) Rohingya Service. Each file represents a daily news segment, typically 30 minutes in length, automatically segmented into chunks of 5–15 seconds for use in self-supervised ASR, pretraining, language identification, and more.
The content was aired publicly as part of VOA’s Rohingya-language radio program and is therefore… See the full description on the dataset page: https://huggingface.co/datasets/freococo/rohingya_asr_audio.audio_swedish_2_dataset_cleanedmon_language_asr_audio
RFA Mon Language Voices
This dataset contains 14.8 hours of audio in the Mon language, sourced from news broadcasts by Radio Free Asia (RFA) Burmese. This is one of the largest publicly accessible audio resources for the Mon language, designed to support research in low-resource automatic speech recognition (ASR), voice activity detection, and other speech-related tasks.
This dataset was created by freococo.
The audio has been automatically segmented into 3,634 manageable chunks and… See the full description on the dataset page: https://huggingface.co/datasets/freococo/mon_language_asr_audio.urdu-turn-detection-audio-v2
🗣️ Urdu Turn Detection (Audio Dataset V2)
This is the official dataset for the model [PuristanLabs1/urdu-turn-v2](https://huggingface.co/PuristanLabs1/urdu-turn-v2), a high precision, low latency system for detecting the end of a conversational turn in Urdu speech.
It contains 11,479 audio clips (balanced between Complete and Incomplete) specifically designed to train robust models for realtime Voice AI applications like "Smart Turn" or "Barge-in" detection.
🚀 How… See the full description on the dataset page: https://huggingface.co/datasets/PuristanLabs1/urdu-turn-detection-audio-v2.karenni_language_asr_audio
RFA Karenni (Kayah) Language Voices
This dataset contains 17 hours of audio in the Karenni (Kayah) language, sourced from news broadcasts by Radio Free Asia (RFA) Burmese. This is one of the largest publicly accessible audio resources for the Karenni language family, designed to support research in low-resource automatic speech recognition (ASR), voice activity detection, and other speech-related tasks.
This dataset was created by freococo.
The audio has been automatically segmented… See the full description on the dataset page: https://huggingface.co/datasets/freococo/karenni_language_asr_audio.
