spelling
Datasets
All datasets matching “spelling”common_voice_9_zh-TW_simple-whisper-large-v3common_voice_9_zh-TW_simplesynthetic-spellingscb10x-thai-dialect-isan-dataset-thai-spellinggerman-spelling-datavi-spelling-correction
Vietnamese Spelling Correction Dataset
This dataset contains 978,417 pairs of noisy (source) and clean (target) Vietnamese sentences, designed for training spelling correction models.
The dataset was synthetically generated by injecting realistic noise into a clean Vietnamese corpus.
Dataset Structure
The dataset is divided into training and testing sets:
Train: 880,575 examples
Test: 97,842 examples
Data Fields
source: The text with injected errors (input).… See the full description on the dataset page: https://huggingface.co/datasets/coung21/vi-spelling-correction.
