ufal/parczech4speech-segmented
ParCzech4Speech (Sentence-Segmented Variant) Dataset Summary ParCzech4Speech (Sentence-Segmented Variant) is a large-scale Czech speech dataset based on parliamentary recordings and official transcripts. This sentence-segmented variant is designed for speech recognition and synthesis tasks, offering clean audio-text alignment and reliable segment boundaries. It is derived from the ParCzech 4.0 corpus and AudioPSP 24.01 audio collection. Using WhisperX and… See the full description on the dataset page: https://huggingface.co/datasets/ufal/parczech4speech-segmented.
ParCzech4Speech (Sentence-Segmented Variant)
Dataset Summary
ParCzech4Speech (Sentence-Segmented Variant) is a large-scale Czech speech dataset based on parliamentary recordings and official transcripts. This sentence-segmented variant is designed for speech recognition and synthesis tasks, offering clean audio-text alignment and reliable segment boundaries.
It is derived from the **ParCzech 4.0** corpus and **AudioPSP 24.01** audio collection. Using WhisperX and Wav2Vec 2.0 for automatic alignment, this dataset ensures high-quality segments and provides rich metadata for filtering and quality control.
The dataset is released under a permissive CC-BY, allowing unrestricted commercial and academic use.
🔔 Note
📢 A larger unsegmented variant of this dataset is now available! The unsegmented version provides longer, continuous speech segments that do not follow sentence boundaries, making it especially suitable for streaming ASR. You can find it under ParCzech4Speech (Unsegmented Variant) on Hugging Face.
Data Splits
Dataset Structure
Each row corresponds to a sentence-level audio segment with accompanying metadata:
Citation
Please cite the dataset as follows:
TODO