CoolFace
Datasetpublic

wheevu/ct219-vietnamese-raw-400k

wheevu/ct219-vietnamese-raw-400k Bộ dữ liệu văn bản tiếng Việt thô đã tiền xử lý, dùng để huấn luyện next-token language model (CT219 - NLP final project). Nguồn dữ liệu Source dataset: VTSNLP/vietnamese_curated_dataset Source split: train Pinned source revision: b81fcce58945970117a1b56d50ec81be2628a5c3 Source licence: không công bố - repo này được tạo ở chế độ private vì lý do đó. Mục đích Huấn luyện next-token language model cho tiếng Việt.… See the full description on the dataset page: https://huggingface.co/datasets/wheevu/ct219-vietnamese-raw-400k.

sourceHugging Faceunknownupdated 1mo agoView on Hugging Face
0likes12downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
wheevu/ct219-vietnamese-raw-400k · CoolFace