CoolFace
Datasetpublic

vngrs/vngrs-web-corpus

Dataset Card for Dataset Name vngrs-web-corpus is a mixed-dataset made of cleaned Turkish sections of OSCAR-2201 and mC4. This dataset is originally created for training VBART and later used for training TURNA. The cleaning procedures of this dataset are explained in Appendix A of the VBART Paper. It consists of 50.3M pages and 25.33B tokens when tokenized by VBART Tokenizer. Dataset Details Uses vngrs-web-corpus is mainly intended to pretrain… See the full description on the dataset page: https://huggingface.co/datasets/vngrs/vngrs-web-corpus.

sourceHugging Facecc-by-nc-sa-4.0updated 2y agoView on Hugging Face
27likes773downloads
12 commits on main
ee5c6202y ago

Update README.md

meliksahturker
376c0693y ago

Update README.md

meliksahturker
eeda0b93y ago

Update README.md

meliksahturker
8032b363y ago

Update README.md

erdiari
49f49633y ago

Update README.md

erdiari
3f984fb3y ago

Upload dataset (part 00005-of-00006)

meliksahturker
ca0efa43y ago

Upload dataset (part 00004-of-00006)

meliksahturker
d5f876f3y ago

Upload dataset (part 00003-of-00006)

meliksahturker
9083a3b3y ago

Upload dataset (part 00002-of-00006)

meliksahturker
f6a4d3f3y ago

Upload dataset (part 00001-of-00006)

meliksahturker
68969ae3y ago

Upload dataset (part 00000-of-00006)

meliksahturker
340031d3y ago

initial commit

meliksahturker