datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
emolia-thinking-balanced-buckets
Emolia-Thinking — Balanced Per-Dimension Bucket Subset
A balanced, per-dimension bucket subset of
VoiceNet/emolia-thinking,
derived from that dataset's zero-shot VoiceNet-dimension labels.
For every VoiceNet voice/prosody/timbre/style dimension, this subset draws a
roughly equal number of clips from each ordinal bucket (0–6), so that
downstream training / probing sees a balanced distribution along each axis
instead of the strongly skewed natural distribution.
How… See the full description on the dataset page: https://huggingface.co/datasets/laion/emolia-thinking-balanced-buckets.vctk_resampled_16k_balancedeveryayah_curated_1s_20s_balancedeveryayah_curated_1s_20s_balanced_largeAudioSet-Strong-Balanceddusha_balanced
Dataset Details
Dataset 'Dusha' split into train, val and test. Half of original train was taken, test split in halfs for val and test, 'neutral' category was cut to make the label distribution more balanced
everyayah_curated_1s_20s_balanced_mediumeveryayah_curated_1s_20s_balanced_tinyASVspoof_2021_DF_Balanced_NormalizedASVspoof_2021_LA_Balanced_Normalizedmalayalam-emotion-balancedDOA_dataset_6_classes_balanced
Dataset Card for "DOA_dataset_6_classes_balanced"
More Information needed
audioset_opus_24kbps_balancedciempiess_balance
Dataset Card for ciempiess_balance
Dataset Summary
The CIEMPIESS BALANCE Corpus is designed to match with the CIEMPIESS LIGHT Corpus (LDC2017S23). So, "Balance" means that if the CIEMPIESS BALANCE is combined with the CIEMPIESS LIGHT, one will get a gender balanced corpus. To appreciate this, one need to know that the CIEMPIESS LIGHT is by itself, a gender unbalanced corpus of approximately 25% of female speakers and 75% of male speakers. So, the CIEMPIESS BALANCE is a… See the full description on the dataset page: https://huggingface.co/datasets/ciempiess/ciempiess_balance.ciempiess_balanceparler-tts-dataset-balancedemolia-balanced-5M-subset
emolia-balanced-5M-subset
A balanced ~5.26M-sample subset of laion/Emolia (80.5M speech samples), packaged as WebDataset-compatible tar shards for direct use in training pipelines.
How this subset was filtered
Samples were selected if they met either of two criteria:
1. Emotion thresholds
Each sample carries 40 emotion annotation scores (from the Emonet taxonomy) in its metadata. A sample qualifies for an emotion bucket if its score for that emotion meets or… See the full description on the dataset page: https://huggingface.co/datasets/laion/emolia-balanced-5M-subset.phoneme-ctc-english-60h-balanced
Phoneme CTC — English 60h (Balanced & Normalized)
A cleaned, normalized and phoneme-balanced version of
bobboyms/phoneme-ctc-english-60h-noisy,
for training phoneme recognition models (CTC) — e.g. as the native acoustic
model behind pronunciation-feedback systems.
What's different from the source dataset
Label noise removed
Roman numerals dropped — eSpeak reads ii/iv/… as "Roman two/four",
producing labels that don't match the audio.
Non-English phonemes dropped… See the full description on the dataset page: https://huggingface.co/datasets/bobboyms/phoneme-ctc-english-60h-balanced.audioset_opus_24kbps_balanced_527balanced-emotion-dataset-majestrino-withtemporal-detailed-captions
Balanced Emotion Dataset — Majestrino with Temporal Detailed Captions
An emotion-balanced subset of TTS-AGI/majestrino-unified-detailed-captions-temporal.
Overview
Total samples: 482,594
Samples per emotion category: 12,997
Number of emotion categories: 40
Format: WebDataset (tar files with FLAC audio + JSON metadata)
Number of tar files: 483
Samples per tar: ~1000
Balancing Strategy
Samples were selected from the source dataset using keyword matching on… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/balanced-emotion-dataset-majestrino-withtemporal-detailed-captions.audio_bbox_balancedasr_bbrave_balancedaudio-prompt-coco-balanced-extendedASVspoof_2021_DF1_Balanced_Normalizedwake_work_detection_balancedhindi_eng_balanced_v2audio-prompt-coco-balancedcv17_su_lu_balanced
Common Voice 17 -- Single / Long Utterance experiment dataset
Built from fixie-ai/common_voice_17_0 (English); the original CV splits are preserved and each is bucketed into single-utterance (1 word) and long-utterance (>= 3 words).
Splits: train_single, test_single, train_long, test_without_single.
six_lg_cv_balancedmyst_single_and_long_utt_balanced
