datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
laion-audio-previewemo_parleripapack_plus_train_3osu-beatmaps
osu! Beatmaps Dataset (WebDataset)
A collection of ranked/loved osu! beatmaps with audio and chart data, in WebDataset format.
Dataset Variants
Variant
Audio Format
Description
original
MP3/OGG/WAV
Full quality original audio files
compressed
64kbps Mono Opus
Compressed audio for smaller download
from datasets import load_dataset
# Load original audio variant
ds = load_dataset("project-riz/osu-beatmaps", "original", streaming=True)
# Load compressed… See the full description on the dataset page: https://huggingface.co/datasets/project-riz/osu-beatmaps.ASMR-Archive-Processed-SFW
ASMR-Archive-Processed-SFW
Overview
This dataset is an “educational” subset of the original OmniAICreator/ASMR-Archive-Processed dataset.
We filtered the original dataset to include only records where the nsfw metadata flag is false.
To maintain the randomness and anonymity of the entries, multiple directories were combined and shuffled.
The nsfw tag in the original dataset is inherited from the tags of the original audio works before they were passed through the… See the full description on the dataset page: https://huggingface.co/datasets/noxwano/ASMR-Archive-Processed-SFW.unsupervised_peoples_speech_raw_voice_activity_detection_snippets_part_1log_prompt_dataMMAudio-precomputed-results
Precomputed results for MMAudio
Results from four model variants of MMAudio.
All results are in the .flac format with lossless compression.
A cache folder contains the feature caches computed by the evaluation script.
Code: https://github.com/hkchengrex/MMAudio
Evaluation: https://github.com/hkchengrex/av-benchmark
VGGSound
Contains the VGGSound test set results. There are 15220 videos, collected with our best effort. Not all videos in the test sets are available… See the full description on the dataset page: https://huggingface.co/datasets/hkchengrex/MMAudio-precomputed-results.parlament_parla_v3
Dataset Card for ParlamentParla v3 - Speech Corpus of Catalan Parliamentary Sessions
A speech corpus composed of Catalan Parliamentary Sessions.The v3 and last version of the corpus includes both clean and other quality segments, divided into short segments (less than 30 seconds) and long segments (more than 30 seconds). The total dataset encompasses 1059h 48m 04s of speech, including 945h 51m 06s for the short segments and 113h 56m 58s for the long segments, with a total of… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/parlament_parla_v3.laions_got_talent_previewparczech4speech-segmented
ParCzech4Speech (Sentence-Segmented Variant)
Dataset Summary
ParCzech4Speech (Sentence-Segmented Variant) is a large-scale Czech speech dataset based on parliamentary recordings and official transcripts.
This sentence-segmented variant is designed for speech recognition and synthesis tasks, offering clean audio-text alignment and reliable segment boundaries.
It is derived from the ParCzech 4.0 corpus and AudioPSP 24.01 audio collection.
Using WhisperX and Wav2Vec 2.0… See the full description on the dataset page: https://huggingface.co/datasets/ufal/parczech4speech-segmented.talent_plus_rl_groups_of_50_with_audiobox_scorestimbre-audio-caption-pairsfreesound-commercially-permissive-subset-with-captionsMM-PreTrain
JavisGPT: A Unified Multi-modal LLM for Sounding-Video Comprehension and Generation
[HomePage]
[Paper]
[GitHub]
TL;DR
We introduce JavisGPT, a multimodal LLM that can understand audiovisual inputs and simultaneously generate synchronized sounding videos in a unified model.
We also curate the JavisInst-Omni dataset to facilitate instruction-tuning for comprehension and generation on sounding videos.
📰 News
[2025.12.30] 🚀 We release the training… See the full description on the dataset page: https://huggingface.co/datasets/JavisVerse/MM-PreTrain.115hours_pvtv_myanmar_asr
115 Hours PVTV Myanmar ASR
This dataset contains 156,262 audio-transcript pairs of spoken Burmese, totaling approximately 115.31 hours. The audio segments were extracted from publicly available YouTube videos published by PVTV and aligned using subtitle timestamps.
Dedication
This dataset would not exist without the persistent voices of PVTV editors, journalists, narrators, and production teams, who continue to speak to the people under difficult conditions. PVTV is the… See the full description on the dataset page: https://huggingface.co/datasets/freococo/115hours_pvtv_myanmar_asr.ipapack_plus_5laion-audio-preview-splitCompA-R-BackupOriginally from https://huggingface.co/papers/2406.11768, we downloaded it from Google Drive and converted it to HuggingFace, as the train_audio portion was observed to have disappeared.
MELD-Preprocessed
MELD Preprocessed for SER
This dataset is the manually preprocessed audio only version of MELD, only audio IDs, utterance transcriptions, dialogue IDs and Utterance IDs were extracted.
S. Poria, D. Hazarika, N. Majumder, G. Naik, R. Mihalcea,
E. Cambria. MELD: A Multimodal Multi-Party Dataset
for Emotion Recognition in Conversation. (2018)
Chen, S.Y., Hsu, C.C., Kuo, C.C. and Ku, L.W.
EmotionLines: An Emotion Corpus of Multi-Party
Conversations. arXiv preprint arXiv:1802.08379… See the full description on the dataset page: https://huggingface.co/datasets/Vano04/MELD-Preprocessed.parczech4speech-unsegmented
ParCzech4Speech (Unsegmented Variant)
Dataset Summary
ParCzech4Speech (Unsegmented Variant) is a large-scale Czech speech dataset derived from parliamentary recordings and official transcripts.
This variant captures continuous speech segments without enforcing sentence boundaries, making it well-suited for real-world streaming ASR scenarios
and speech modeling tasks that benefit from natural discourse flow.
The dataset is created using a combination of WhisperX and… See the full description on the dataset page: https://huggingface.co/datasets/ufal/parczech4speech-unsegmented.kenya-philippines-twospeaker-english-dialogue
Kenya/Philippines English Dialogue
Two-speaker dialogues in English, recorded on split tracks.
Changelog
Jan 2026: v1 release - vad-segmented WebRTC tracks
Specs
Speakers: >150; ~15 PH, remaining KE
Total duration: ~65 hours
Files sample rate: 48kHz
Actual sample rate: TBD
Language: English (PH, KE accents)
Topics: day-to-day conversation
Collection method
The dataset is built to capture the variety in the Kenyan accent.
The Philippino interviewers… See the full description on the dataset page: https://huggingface.co/datasets/Reord-AI/kenya-philippines-twospeaker-english-dialogue.voxceleb2-40k-part1-preprocess-all-files-separateipapack_plus_3ipapack_plus_6animalspeak-pseudovox
AnimalSpeak Pseudovox Train-Unseen
This dataset contains the train-unseen split of AnimalSpeak Pseudovox. Each
example is a short, silence-trimmed, single-vocalization WAV clip plus compact
per-clip metadata. It does not include generated conversations, captions, QA
pairs, or MCQ answers.
Rows: 346,907
Shards: 18
Maximum rows per shard: 20,000
Files
data-20k/train-*.tar: WebDataset-style shards containing
audio/<audio_name> WAV entries.
metadata.parquet: one row per… See the full description on the dataset page: https://huggingface.co/datasets/EarthSpeciesProject/animalspeak-pseudovox.ipapack_plus_1urdu-turn-detection-audio-v2
🗣️ Urdu Turn Detection (Audio Dataset V2)
This is the official dataset for the model [PuristanLabs1/urdu-turn-v2](https://huggingface.co/PuristanLabs1/urdu-turn-v2), a high precision, low latency system for detecting the end of a conversational turn in Urdu speech.
It contains 11,479 audio clips (balanced between Complete and Incomplete) specifically designed to train robust models for realtime Voice AI applications like "Smart Turn" or "Barge-in" detection.
🚀 How… See the full description on the dataset page: https://huggingface.co/datasets/PuristanLabs1/urdu-turn-detection-audio-v2.paimonpolyglot-modelsThis dataset repository shall serve as a mirror hosting models for polyglot.
Availability
Please note that currently the only languages, which have all models, are English (en) and Bulgarian (bg). Other languages may have partial support.
Adding missing models
In case you have previously downloaded polyglot language models, which are not available in this repo, please open a Pull Request.
License
All rights belong to the original authors. Please refer to the… See the full description on the dataset page: https://huggingface.co/datasets/ndandanov/polyglot-models.
