CoolFace
Datasetpublic

NickWeng/tw_parliament_split

Parliament Parliament is an OpenFormosa Traditional Chinese speech dataset derived from the zh_tw split of disco-eth/WorldSpeech. It contains audio clips and human transcripts from Taiwan Legislative Yuan IVOD parliamentary proceedings. This release keeps only rows that passed the Taiwan-OmniData / FineWeb2-style text filtering pipeline. Audio is preserved from the upstream dataset and cast as a Hugging Face Audio(sampling_rate=24000) feature. Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/NickWeng/tw_parliament_split.

sourceHugging Facecc-by-nc-4.0updated 6d agoView on Hugging Face
0likes524downloads

NickWeng/tw_parliament_split · main · files are served by the source, never re-hosted here