datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
parlament_parla_v3_punctuated
ParlamentParla v3, punctuated and capitalized (train, short segments)
A derivative of ParlamentParla v3, the speech corpus of Catalan parliamentary sessions published by the Language Technologies Unit of the Barcelona Supercomputing Center (BSC-LT) within the Aina project. ParlamentParla v3 distributes its transcriptions lowercased and without any punctuation. This dataset keeps that upstream text untouched in the text column and adds a second column, text_punctuated, with… See the full description on the dataset page: https://huggingface.co/datasets/Ugiat/parlament_parla_v3_punctuated.parlament-parla-v2-voxceleb-resnet34-LM-embparlament_parla_resnet_embparlament_parla_ecapa_emb
Dataset Card for "parlament_parla_ecapa_emb"
More Information needed
image-text_parlamentsdienste-protokolle
Dataset Card for transkribus-exports-280089-1-raw-xml
This dataset was created using pagexml-hf converter from Transkribus PageXML data.
Dataset Summary
This dataset contains 479 samples across 1 split(s).
Geographical scope: SwitzerlandPeriod: 1848-1900Languages: German, FrenchType of document: ProtocolsProvenance: Swiss Federal Archives
Projects Included
1849_01
1849_02
1849_03
1849_04
1849_05
1849_06
1849_07
1849_08
637157213248534270_verkleinert… See the full description on the dataset page: https://huggingface.co/datasets/dh-unibe/image-text_parlamentsdienste-protokolle.roots_ca_parlament_parlaROOTS Subset: roots_ca_parlament_parla
parlament_parla
Dataset uid: parlament_parla
Description
Homepage
Licensing
Speaker Locations
Sizes
0.0000 % of total
0.0000 % of ca
BigScience processing steps
Filters applied to: ca
dedup_document
dedup_template_soft
filter_remove_empty_docs
filter_small_docs_bytes_1024
