CoolFace
Datasetpublic

coung21/vi-spelling-correction

Vietnamese Spelling Correction Dataset This dataset contains 978,417 pairs of noisy (source) and clean (target) Vietnamese sentences, designed for training spelling correction models. The dataset was synthetically generated by injecting realistic noise into a clean Vietnamese corpus. Dataset Structure The dataset is divided into training and testing sets: Train: 880,575 examples Test: 97,842 examples Data Fields source: The text with injected… See the full description on the dataset page: https://huggingface.co/datasets/coung21/vi-spelling-correction.

sourceHugging Facemitupdated 8mo agoView on Hugging Face
1likes77downloads

coung21/vi-spelling-correction · main · files are served by the source, never re-hosted here