CoolFace
Datasetpublic

vngrs/vngrs-web-corpus

Dataset Card for Dataset Name vngrs-web-corpus is a mixed-dataset made of cleaned Turkish sections of OSCAR-2201 and mC4. This dataset is originally created for training VBART and later used for training TURNA. The cleaning procedures of this dataset are explained in Appendix A of the VBART Paper. It consists of 50.3M pages and 25.33B tokens when tokenized by VBART Tokenizer. Dataset Details Uses vngrs-web-corpus is mainly intended to pretrain… See the full description on the dataset page: https://huggingface.co/datasets/vngrs/vngrs-web-corpus.

sourceHugging Facecc-by-nc-sa-4.0updated 2y agoView on Hugging Face
27likes773downloads

vngrs/vngrs-web-corpus · main · files are served by the source, never re-hosted here