CoolFace
Datasetpublic

statworx/leipzip-swiss

Dataset Card for Leipzig Corpora Swiss German Dataset Summary Swiss German Wikipedia corpus based on material from 2021.The corpus gsw_wikipedia_2021 is a Swiss German Wikipedia corpus based on material from 2021. It contains 232,933 sentences and 3,824,547 tokens. Languages Swiss-German Dataset Structure Data Instances Single sentences. Data Fields sentence: Text as string. Data Splits… See the full description on the dataset page: https://huggingface.co/datasets/statworx/leipzip-swiss.

sourceHugging Faceccupdated 4y agoView on Hugging Face
2likes17downloads
../
filetrain-00000-of-00001-0d60107f007bed33.parquet45.7 MBdownload

statworx/leipzip-swiss · main · files are served by the source, never re-hosted here