datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mls-enhanced-dacvae
Multilingual LibriSpeech converted to DAC VAE latents
Source
facebook/multilingual_librispeech
Format
Each tar shard (~2GB) contains samples with three files per sample:
{sample_key}.audio.flac # Original audio (FLAC, original sample rate)
{sample_key}.dacvae.npy # DAC VAE latent [T_latent, 128] numpy float32
{sample_key}.metadata.json # All metadata + duration_seconds + chars_per_second
DAC VAE Latent Format
Model:… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/mls-enhanced-dacvae.YouTube-English
English Audio Dataset from YouTube
This dataset contains English audio segments and creator uploaded transcripts (likely higher quality) extracted from various YouTube channels, along with corresponding transcript metadata. The data is intended for training automatic speech recognition (ASR) models.
Data Source and Processing
The data was obtained through the following process:
Download: Audio (.m4a) and available English subtitles (.srt for en, en.j3PyPqV-e1s) were… See the full description on the dataset page: https://huggingface.co/datasets/OrcinusOrca/YouTube-English.media_queen_entertaiment_voices
Media Queen Entertainment Voices
Where the stars speak, and their stories come to life.
Media Queen Entertainment Voices is a massive, large-scale collection of 190,013 short audio segments (totaling approximately 125 hours of speech) derived from public videos by Media Queen Entertainment — a prominent digital media channel in Myanmar focused on celebrity news, lifestyle content, and in-depth interviews.
The source channel regularly features:
Interviews with artists, actors… See the full description on the dataset page: https://huggingface.co/datasets/freococo/media_queen_entertaiment_voices.myanmar-english-accent-speech
Myanmar English Accent Speech (PVTV & FOEIM)
This dataset contains English speech by Myanmar speakers, collected from public videos published by PVTV and FOEIM — two media channels operating under the National Unity Government (NUG).
The clips reflect a wide range of spoken English contexts: interviews, announcements, sermons, and educational content. The speakers vary in tone, pace, and emotion — but all share the characteristic sound of Burmese-accented English.
This dataset was… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar-english-accent-speech.Taiwan-Tongues-ASR-CE-dataset-en
Taiwan-Tongues-ASR-CE-dataset-en
本資料集為 Taiwan-Tongues-ASR-CE 專案所使用的預訓練資料,透過 WebDataset 格式打包,並上傳至 Hugging Face 以便研究人員與開發者自由取用。
📂 Dataset 結構
本資料集分為 Training 與 Test 兩個子集,均以 WebDataset tar 檔案形式存放:
Training set (WebDataset format)
train/train-000000.tar
train/train-000001.tar
...
Test set (WebDataset format)
test/test-000000.tar
...
tsv set
train.tsv
test.tsv
...
每個 tar 內部均包含對應的音檔與標註,方便直接搭配 WebDataset 與 PyTorch / Hugging Face datasets 進行訓練與測試。
🏷️… See the full description on the dataset page: https://huggingface.co/datasets/adi-gov-tw/Taiwan-Tongues-ASR-CE-dataset-en.eurospeech-enhanced-dacvae
EuroSpeech parliamentary speech converted to DAC VAE latents
Source
disco-eth/EuroSpeech
Format
Each tar shard (~2GB) contains samples with three files per sample:
{sample_key}.audio.flac # Original audio (FLAC, original sample rate)
{sample_key}.dacvae.npy # DAC VAE latent [T_latent, 128] numpy float32
{sample_key}.metadata.json # All metadata + duration_seconds + chars_per_second
DAC VAE Latent Format
Model:… See the full description on the dataset page: https://huggingface.co/datasets/laion/eurospeech-enhanced-dacvae.
