hoanghai2110/vi-pretrain-clean
Vietnamese Pretraining Dataset Bộ dữ liệu tiếng Việt chất lượng cao để pretrain mô hình ngôn ngữ từ đầu (from scratch). Mục tiêu: "ít mà vàng" — ít dữ liệu nhưng cực sạch. Thống kê Chỉ số Giá trị Tổng docs 511,198 Raw text ~0.69 GB Ước tính tokens ~230M tokens Nguồn 4 nguồn Nguồn dữ liệu Nguồn Docs Loại nội dung Wikipedia VI 378,895 Bách khoa toàn thư OPUS OpenSubtitles 84,021 Hội thoại, phụ đề phim OPUS… See the full description on the dataset page: https://huggingface.co/datasets/hoanghai2110/vi-pretrain-clean.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face