datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
common_voice_26_0_de
Mozilla Common Voice 26.0 - German (IPA & Clean Validated Subset)
Repacking version of Common Voice 26.0 German officialy published by Mozilla Data Collective, following Hugging Face Parquet Shards standard, with feature for listening to audio directly on the Web Hub, and the addition of a data column for the IPA transcription of each sentence.
📊 Dataset parameters
Origin: Mozilla Common Voice 26.0 (version 18/06/2026).
Data amount (Validated): 950,877 MP3 audio… See the full description on the dataset page: https://huggingface.co/datasets/q1805/common_voice_26_0_de.german-golden-audio_speech-IPA
🌟 German Golden Speech & IPA Corpus (FLEURS + Multilingual TEDx)
An ultra-clean, high-standard curated German speech dataset combining Google FLEURS (de_de) and Multilingual TEDx German (mTEDx), fully embedded with 16kHz WAV audio bytes, normalized orthographic text, and pre-computed International Phonetic Alphabet (IPA) transcriptions.
📊 Dataset Summary
Total Samples: 1,354 high-quality audio recordings.
Total Size: ~419 MB (Compressed Parquet format).
Audio… See the full description on the dataset page: https://huggingface.co/datasets/q1805/german-golden-audio_speech-IPA.common_voice_spontaneous_speech_4_0_de
Mozilla Common Voice Spontaneous Speech 4.0 - German (Parquet & IPA)
Clean Parquet format of Mozilla Spontaneous Speech 4.0 German (264 samples).
german-pronuncheck-mega-dataset
German PronunCheck Mega Dataset 🇩🇪
Dataset Summary
This is a highly curated, 123GB+ mega-dataset designed specifically for training and fine-tuning German Automatic Speech Recognition (ASR) and Computer-Assisted Pronunciation Training (CAPT) models, such as HuBERT and Wav2Vec2.
Composition
This dataset is a clean concatenation of three distinct open-source datasets:
Mozilla Common Voice 26.0 (German): Standard scripted crowdsourced speech.… See the full description on the dataset page: https://huggingface.co/datasets/q1805/german-pronuncheck-mega-dataset.
