CoolFace
Datasetpublic

ufal/parczech4speech-segmented

ParCzech4Speech (Sentence-Segmented Variant) Dataset Summary ParCzech4Speech (Sentence-Segmented Variant) is a large-scale Czech speech dataset based on parliamentary recordings and official transcripts. This sentence-segmented variant is designed for speech recognition and synthesis tasks, offering clean audio-text alignment and reliable segment boundaries. It is derived from the ParCzech 4.0 corpus and AudioPSP 24.01 audio collection. Using WhisperX and… See the full description on the dataset page: https://huggingface.co/datasets/ufal/parczech4speech-segmented.

sourceHugging Facecc-by-2.0updated 1y agoView on Hugging Face
1likes331downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
ufal/parczech4speech-segmented · CoolFace