datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LAION-Audio-300Memo_webds_2emo_parleremo_webdsCounterStrike-1K-360-wds
CounterStrike-1K — 360p WebDataset shards
This repo contains the 360p shards of CounterStrike-1K. Use the main repo to browse the manifest, schema, and subsets.
360p is the recommended resolution for most training pipelines — the actions/state/events/metadata sidecars are identical to the 720p shards, so you can swap resolutions without touching downstream code.
Quickstart
Start a fresh uv project and add the loader:
mkdir cs1k-demo && cd cs1k-demo
uv init
uv add… See the full description on the dataset page: https://huggingface.co/datasets/ArnieRamesh/CounterStrike-1K-360-wds.qwen3-tts-multilingual-emotional-speechMemeEffect-382K-audioWe are releasing the audio files that we have collected from MemeEffect-382K dataset. All the files are being shared as .tar files and files are rnamed using their respective id that can be found through the metadata.
We share these files as-part of research initiative.
ipapack_plus_train_3emo_speech_filtered_v12 second filtered emotional speech in webdataset format
https://huggingface.co/datasets/EQ4You/Emotional_Speech
parlament_parla_v3
Dataset Card for ParlamentParla v3 - Speech Corpus of Catalan Parliamentary Sessions
A speech corpus composed of Catalan Parliamentary Sessions.The v3 and last version of the corpus includes both clean and other quality segments, divided into short segments (less than 30 seconds) and long segments (more than 30 seconds). The total dataset encompasses 1059h 48m 04s of speech, including 945h 51m 06s for the short segments and 113h 56m 58s for the long segments, with a total of… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/parlament_parla_v3.Emilia-with-Emotion-Annotations3Audiobook_Noice_V2t2adata3vocal_bursts_taxonomy_100_clean_wdslaion-audio-preview-splitlmd_mp3
The Lakh MIDI Dataset in MP3
The MIDI files from LMD were synthesized and split into 15-second segments.
Soundfont: GeneralUser GS 2.02
Notes:
Files that were skipped:
Corrupt or unreadable by pretty_midi or mido.
Could not be synthesized with fluidsynth.
Longer than 20 minutes.
(Some) ending segments shorter than 5 seconds.
Reference:
Colin Raffel. "Learning-Based Methods for Comparing Sequences, with Applications to Audio-to-MIDI Alignment and Matching". PhD Thesis, 2016.
emolia-3k-speaker-clusters
Emolia 3K Speaker Clusters
A curated set of 3,000 diverse speaker clusters derived from the TTS-AGI/emolia-hq dataset, with up to 20 representative audio samples per cluster.
Overview
The original emolia-hq dataset contains hundreds of thousands of speech samples with 128-dimensional WavLM speaker timbre embeddings. These were first clustered into 10,000 centroids, then intelligently pruned to 3,000 using density-aware farthest-point sampling to ensure:
Outlier… See the full description on the dataset page: https://huggingface.co/datasets/laion/emolia-3k-speaker-clusters.FS_32Kmid-classical-openmusenet4-3mearsipapack_plus_3ai-generated-songs3MuseBench-part3test-webdatasettest
wds_vocal_burst_100nsfw_tts_datasetA high-quality audio dataset designed for training and fine-tuning NSFW TTS models, including 30 characters, over 1000 hours of audio, and rich emotion/sound annotations.
Sample format: WAV (audio) + TXT (annotations), including emotion_label, sound_label and text.
Annotations: 6000+ emotion labels (intimate, breathy, teasing, etc.) and 760+ sound labels (moan, sigh, laugh, etc.) in the full version.
Audio sample is as follows:
[intimate, breathy, pleased] Oh, <moan> it feels so good when your… See the full description on the dataset page: https://huggingface.co/datasets/DMC-ykfx33/nsfw_tts_dataset.emo_subset_webdspt-br-tts-iasmin-qwen3
pt-br-tts-iasmin-qwen3
21957 clips PT-BR sintetizados com Qwen3-TTS-12Hz-1.7B. Voz Iasmin (voice-clone, is_iasmin=true, ~13957 clips) + vozes diversas CustomVoice (Ryan, Aiden, Vivian, Dylan, is_iasmin=false). WAV em tar shards (WebDataset); transcricao, voz, is_iasmin e sr em metadata.jsonl.
emo_speech_samplevb_sample
