parlament
parlament_parla_v3
Dataset Card for ParlamentParla v3 - Speech Corpus of Catalan Parliamentary Sessions
A speech corpus composed of Catalan Parliamentary Sessions.The v3 and last version of the corpus includes both clean and other quality segments, divided into short segments (less than 30 seconds) and long segments (more than 30 seconds). The total dataset encompasses 1059h 48m 04s of speech, including 945h 51m 06s for the short segments and 113h 56m 58s for the long segments, with a total of… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/parlament_parla_v3.parlament_parlaThis is the ParlamentParla speech corpus for Catalan prepared by Col·lectivaT. The audio segments were extracted from recordings the Catalan Parliament (Parlament de Catalunya) plenary sessions, which took place between 2007/07/11 - 2018/07/17. We aligned the transcriptions with the recordings and extracted the corpus. The content belongs to the Catalan Parliament and the data is released conforming their terms of use.
Preparation of this corpus was partly supported by the Department of Culture of the Catalan autonomous government, and the v2.0 was supported by the Barcelona Supercomputing Center, within the framework of the project AINA of the Departament de Polítiques Digitals.
As of v2.0 the corpus is separated into 211 hours of clean and 400 hours of other quality segments. Furthermore, each speech segment is tagged with its speaker and each speaker with their gender. The statistics are detailed in the readme file.
For more information, go to https://github.com/CollectivaT-dev/ParlamentParla or mail info@collectivat.cat.parlament_parla_v3_punctuated
ParlamentParla v3, punctuated and capitalized (train, short segments)
A derivative of ParlamentParla v3, the speech corpus of Catalan parliamentary sessions published by the Language Technologies Unit of the Barcelona Supercomputing Center (BSC-LT) within the Aina project. ParlamentParla v3 distributes its transcriptions lowercased and without any punctuation. This dataset keeps that upstream text untouched in the text column and adds a second column, text_punctuated, with… See the full description on the dataset page: https://huggingface.co/datasets/Ugiat/parlament_parla_v3_punctuated.parlament-parla-v2-voxceleb-resnet34-LM-embparlament_parla_resnet_embparlament_parla_ecapa_emb
Dataset Card for "parlament_parla_ecapa_emb"
More Information needed
