CoolFace
Datasetpublic

ufal/parczech4speech-segmented

ParCzech4Speech (Sentence-Segmented Variant) Dataset Summary ParCzech4Speech (Sentence-Segmented Variant) is a large-scale Czech speech dataset based on parliamentary recordings and official transcripts. This sentence-segmented variant is designed for speech recognition and synthesis tasks, offering clean audio-text alignment and reliable segment boundaries. It is derived from the ParCzech 4.0 corpus and AudioPSP 24.01 audio collection. Using WhisperX and… See the full description on the dataset page: https://huggingface.co/datasets/ufal/parczech4speech-segmented.

sourceHugging Facecc-by-2.0updated 1y agoView on Hugging Face
1likes336downloads
8 commits on main
3a128d91y ago

Update README.md

stanvla
1a41bcb1y ago

Update README.md

stanvla
f33436c1y ago

Update README.md

stanvla
92253921y ago

Create README.md

stanvla
75d736d1y ago

Add files using upload-large-folder tool

stanvla
c8783191y ago

Delete files train-*.tar dev-*.tar test-*.tar with huggingface_hub

stanvla
cd105531y ago

Add files using upload-large-folder tool

stanvla
7cc37c31y ago

initial commit

stanvla