Ugiat/parlament_parla_v3_punctuated
ParlamentParla v3, punctuated and capitalized (train, short segments) A derivative of ParlamentParla v3, the speech corpus of Catalan parliamentary sessions published by the Language Technologies Unit of the Barcelona Supercomputing Center (BSC-LT) within the Aina project. ParlamentParla v3 distributes its transcriptions lowercased and without any punctuation. This dataset keeps that upstream text untouched in the text column and adds a second column, text_punctuated, with… See the full description on the dataset page: https://huggingface.co/datasets/Ugiat/parlament_parla_v3_punctuated.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face