spell-correction
vsec-vietnamese-spell-correction
VSEC: Vietnamese Spell Correction Dataset
Dataset Description
VSEC (Vietnamese Spell Correction) is a comprehensive dataset for Vietnamese spelling error detection and correction, containing 9,341 sentences with 11,202 human-made misspellings across 5,211 unique error types. This dataset represents the largest publicly available collection of Vietnamese spelling errors with syllable-level annotations, making it an invaluable resource for developing and evaluating… See the full description on the dataset page: https://huggingface.co/datasets/nguyenthanhasia/vsec-vietnamese-spell-correction.vn-spell-correction-eval-real
vn-spell-correction-eval-real
Out-of-distribution evaluation corpus for Vietnamese spell-correction
models — 150 hand-curated (noisy, clean) pairs sampled from real
VN error sources, not generated by nom.text.noise.
This is the test set we use to verify a spell-correction model
generalises beyond its own synthetic training distribution. A model
that scores 95 % on nom-vn's synthetic eval grid and 60 % on this
set is overfit to the noise generator.
Splits
Config… See the full description on the dataset page: https://huggingface.co/datasets/nrl-ai/vn-spell-correction-eval-real.asr-spell-correction-ruspell-correction-ru
Spell Correction RU — датасеты для коррекции ошибок в русском тексте
Набор данных для обучения моделей исправления орфографических, пунктуационных и регистровых ошибок в русскоязычных текстах.
Каждый пример — пара «правильный текст» → «текст с ошибкой». Датасет использовался для обучения модели melsmm/Spell-Corrector-RU-4B.
📦 Код генерации, ноутбуки и полное описание проекта: github.com/melsmm/llm-spell-corrector
Состав
Датасет содержит две конфигурации… See the full description on the dataset page: https://huggingface.co/datasets/melsmm/spell-correction-ru.russian-spell-correctionsasr_spell_correction_ru
