Sanng1112/vietnamese-spelling-synthetic-1gb
Vietnamese Spelling Correction Synthetic 1GB Corpus tổng hợp cho bài toán phát hiện và sửa lỗi chính tả tiếng Việt, trích xuất và làm sạch từ Vietnamese Wikipedia Dump mới nhất, hỗ trợ mô hình phân loại token 1-đối-1. 📊 Nội dung Dataset File Số lượng mẫu (Rows) Dung lượng thô Mô tả train_full.jsonl 4.564.360 6.04 GB Tập huấn luyện chính thức validation_full.jsonl 95.094 124 MB Tập validation held-out (tách theo page_id) word_vocab_full.json 9.500… See the full description on the dataset page: https://huggingface.co/datasets/Sanng1112/vietnamese-spelling-synthetic-1gb.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face