duongttr/vi-dataset-for-pretrain
Dataset Card for "vi-dataset-for-pretrain" This is a combination of multiple Vietnamese dataset for pretraining CLMs such as GPT, GPT2, etc. The dataset consists of: vietgpt/covid_19_news_vi hieunguyen1053/binhvq-news-corpus oscar (unshuffled_deduplicated_vi) vietgpt/wikipedia_vi Dataset info Splits N.o examples Size Train 23,891,116 77.36 GB Validation 1,257,428 4.06 GB Total 25,148,544 81.43 GB
5595
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face