datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vsec-vietnamese-spell-correction
VSEC: Vietnamese Spell Correction Dataset
Dataset Description
VSEC (Vietnamese Spell Correction) is a comprehensive dataset for Vietnamese spelling error detection and correction, containing 9,341 sentences with 11,202 human-made misspellings across 5,211 unique error types. This dataset represents the largest publicly available collection of Vietnamese spelling errors with syllable-level annotations, making it an invaluable resource for developing and evaluating… See the full description on the dataset page: https://huggingface.co/datasets/nguyenthanhasia/vsec-vietnamese-spell-correction.vn-spell-correction-eval-real
vn-spell-correction-eval-real
Out-of-distribution evaluation corpus for Vietnamese spell-correction
models — 150 hand-curated (noisy, clean) pairs sampled from real
VN error sources, not generated by nom.text.noise.
This is the test set we use to verify a spell-correction model
generalises beyond its own synthetic training distribution. A model
that scores 95 % on nom-vn's synthetic eval grid and 60 % on this
set is overfit to the noise generator.
Splits
Config… See the full description on the dataset page: https://huggingface.co/datasets/nrl-ai/vn-spell-correction-eval-real.asr-spell-correction-ruspell-correction-ru
Spell Correction RU — датасеты для коррекции ошибок в русском тексте
Набор данных для обучения моделей исправления орфографических, пунктуационных и регистровых ошибок в русскоязычных текстах.
Каждый пример — пара «правильный текст» → «текст с ошибкой». Датасет использовался для обучения модели melsmm/Spell-Corrector-RU-4B.
📦 Код генерации, ноутбуки и полное описание проекта: github.com/melsmm/llm-spell-corrector
Состав
Датасет содержит две конфигурации… See the full description on the dataset page: https://huggingface.co/datasets/melsmm/spell-correction-ru.russian-spell-correctionsasr_spell_correction_ruasr-spell-correction-ruvn-spell-correction-train
nrl-ai/vn-spell-correction-train
459,478 (noisy, clean) Vietnamese training pairs for fine-tuning a
seq2seq spell-correction model. Each row:
{"input": "<noisy>", "target": "<clean>"}
Both fields are NFC-normalized.
How it was built
Clean side: same 500K register-balanced mix as
nrl-ai/vn-diacritic-train —
350K Vietnamese Wikipedia (CC-BY-SA-4.0,
hirine/wikipedia-vietnamese-1M296K-dataset) + 150K NFC-fixed
Vietnamese news (CC-BY-4.0, tmnam20/Vietnamese-News-dedup).… See the full description on the dataset page: https://huggingface.co/datasets/nrl-ai/vn-spell-correction-train.persian-spell-correction-dataset
Persian Spell Correction & Augmentation Dataset
This is a large-scale, parallel dataset for Persian spell correction, text normalization, and augmentation. It is designed to train and evaluate models for correcting a wide variety of common and synthetic errors in Persian text.
The dataset is built from two main components:
Natural Data: Text from diverse Persian corpora and its corresponding clean, corrected version (corrected_text) generated by an LLM.
Augmented Data: The… See the full description on the dataset page: https://huggingface.co/datasets/aligh4699/persian-spell-correction-dataset.asr-spell-correction-ru-hw01
Russian ASR correction: homework 01
1020 pairs: 450 Groq-generated ASR-like inputs,
180 Groq-generated numeral-to-word pairs, and 390 identity
examples added by copying screened clean targets.
Model: openai/gpt-oss-120b. Generation: 8bae1a605e8506a1; prompt version: groq_asr_numbers_v2.
The ASR target names come from the existing Groq-generated pool targets.jsonl.
No Python character corruption is used. This is synthetic text, not real ASR output.
Generation and… See the full description on the dataset page: https://huggingface.co/datasets/sobadsodead/asr-spell-correction-ru-hw01.ru-asr-spell-correctionasr_spell_correction_ruru-asr-spell-correction-v2vn-spell-correction-eval
nrl-ai/vn-spell-correction-eval
Vietnamese spell-correction evaluation grid: 4 source registers × 2
noise levels = 8 splits, 2,098 (noisy, clean) sentence pairs total.
Each pair is {"input": "<noisy>", "target": "<clean>"}. Both sides are
NFC-normalized. The clean target is the same sentence used as the
target in nrl-ai/vn-diacritic-eval —
spell correction is a strict superset of diacritic restoration, so we
reuse the same registers-balanced corpus.
Splits
Two noise… See the full description on the dataset page: https://huggingface.co/datasets/nrl-ai/vn-spell-correction-eval.spell-correction
Spell-Check Dataset
This dataset consists of pairs of misspelled words and their corresponding correctly spelled words, designed for training and evaluating character-level spelling correction models. It is particularly useful for tasks such as:
Spelling correction
Character-level sequence-to-sequence modeling
Error detection and correction in text
Each data point in the dataset contains:
misspelled: A misspelled version of a word.
correct: The corrected spelling of the word.… See the full description on the dataset page: https://huggingface.co/datasets/torinriley/spell-correction.popular_names_spell_correctionmy-groq-spell-correction-mistake-datasetrussian-movie-spell-correctiongroq-spell-correction-mistake-datasetrussian-spell-correction-datasetrussian-spell-correction-dataset-2groq1441-spell-correction-mistake-datasetrussian-spell-correction-groqrussian_spell_correction_groqspell_correction_ruspell_correction_datasetspell_correction_datasets_pt_brspell_correction_datasetsyntetic-dataset-for-gigaam-spell-correctionrussian-spell-correction-dataset
