datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
parczech4speech-segmented
ParCzech4Speech (Sentence-Segmented Variant)
Dataset Summary
ParCzech4Speech (Sentence-Segmented Variant) is a large-scale Czech speech dataset based on parliamentary recordings and official transcripts.
This sentence-segmented variant is designed for speech recognition and synthesis tasks, offering clean audio-text alignment and reliable segment boundaries.
It is derived from the ParCzech 4.0 corpus and AudioPSP 24.01 audio collection.
Using WhisperX and Wav2Vec 2.0… See the full description on the dataset page: https://huggingface.co/datasets/ufal/parczech4speech-segmented.Common_Voice_Delta_Segment_11.0
