datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tibetan-spelling-correction-dataset
Tibetan Spelling Correction
Sentence-level spelling correction pairs for Tibetan. Each row pairs an annotator's transcription of a manuscript page segment with the final reviewer's corrected version, from the double-annotation workflow of the BDRC Etext Corpus.
For training and evaluating post-correction models such as TiSpell.
Source batches: Ume 1-4, Uchen 1-4. 4,672 pages.
Contents
train
validation
test
All
error pairs
43,015
2,268
2,229
47,512… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/tibetan-spelling-correction-dataset.vi-spelling-correction
Vietnamese Spelling Correction Dataset
This dataset contains 978,417 pairs of noisy (source) and clean (target) Vietnamese sentences, designed for training spelling correction models.
The dataset was synthetically generated by injecting realistic noise into a clean Vietnamese corpus.
Dataset Structure
The dataset is divided into training and testing sets:
Train: 880,575 examples
Test: 97,842 examples
Data Fields
source: The text with injected errors (input).… See the full description on the dataset page: https://huggingface.co/datasets/coung21/vi-spelling-correction.spelling-correction-french-news
Spelling correction dataset (French)
This dataset is generated by transforming/corrupting sentences of a French news corpus
provided by the University of Leipzig.
The following transformations are applied to words in the sentences:
concatenation of pairs of words
swapping of neighboring letters in words
insertion
deletion
replacement (by neighboring characters in AZERTY keyboard)
Generation
./scripts/get_data.py -t news -y 2023 -s 10K
./scripts/generate_dataset.py… See the full description on the dataset page: https://huggingface.co/datasets/fdemelo/spelling-correction-french-news.spelling_correction_chatgpt200926khmer-spelling-corrections
Khmer Spelling Corrections
Naturally occurring Khmer misspellings paired with the word the writer meant.
The labels are not annotated, they are observed. Search sessions record the
whole typing trajectory toward a single word, so when a user types something,
fails, adjusts and lands on a real dictionary headword, the failed attempt and
the headword form a correction pair produced by a real person under no
instruction to make mistakes.
Only pairs within two edits of the target… See the full description on the dataset page: https://huggingface.co/datasets/seanghay/khmer-spelling-corrections.openslr-sinhala-spelling-correction-prediction-reference-60000
OpenSLR Sinhala Spelling Correction – Prediction / Reference
This dataset contains Sinhala sentence pairs intended for training and evaluating
spelling-correction models on ASR output.
Column
Description
dyslexic_sentence
Noisy / dyslexic sentence (model input – ASR hypothesis)
clean_sentence
Clean / correct sentence (ground truth)
Splits
Split
Source file
Rows (approx.)
train
openslr-60000.csv
~60 000
eval
openslr-7000.csv
~7 000… See the full description on the dataset page: https://huggingface.co/datasets/SPEAK-PP/openslr-sinhala-spelling-correction-prediction-reference-60000.synthetic-error-generated-spelling-correction-dataset-100kSpelling_correctionsinhala-spelling-correction-already-corrected-pairs
Sinhala ASR Prediction-Reference Dataset (3000 no-numbers)
Dataset Description
This dataset contains sentence pairs for spelling correction:
dyslexic_sentence: noisy / predicted text
clean_sentence: clean reference text
Dataset Statistics
Split
Samples
Train
2,400
Eval
300
Test
300
Total
3,000
Usage
from datasets import load_dataset
dataset = load_dataset("SPEAK-PP/sinhala-spelling-correction-already-corrected-pairs")… See the full description on the dataset page: https://huggingface.co/datasets/SPEAK-PP/sinhala-spelling-correction-already-corrected-pairs.spellingcorrectionFrenchrussian-spelling_correction_datasetrussian_spelling_correction_datasetspelling_correction_100openslr-sinhala-spelling-correction-prediction-reference-60000-modified-v1
OpenSLR Sinhala Spelling Correction Dataset (Modified v1)
Dataset Description
This dataset contains pairs of misspelled and correctly spelled Sinhala sentences for spelling correction tasks. It is derived from the OpenSLR Sinhala dataset and includes 60,000 examples of spelling correction patterns in Sinhala language.
Dataset Structure
The dataset contains the following columns:
dyslexic_sentence: Misspelled or incorrectly written Sinhala sentence… See the full description on the dataset page: https://huggingface.co/datasets/SPEAK-PP/openslr-sinhala-spelling-correction-prediction-reference-60000-modified-v1.spelling_correction_single_word_v1vietnamese-spelling-correction-datasetrussian-spelling-correctionrussian-spelling-correction2russian-spelling-correctionvie_spelling_correction_data_v3russian_spelling_correction_datasetvie_spelling_correction_data_v5spelling_correction_words_v1tajik-spelling-correction-pairs
🇹🇯 Tajik Spelling Correction Pairs (Clean ↔ Noisy)
A parallel corpus for training automatic spelling correction models for the Tajik language. Each record contains an original cleaned text (clean_text) and its version with artificially introduced typos and errors (noisy_text).
📖 Description
This dataset was created by merging four distinct Tajik language resources:
News articles (tajik-news-cluster)
Lexical pairs (TajPersParallelLexicalCorpus)
Toponyms… See the full description on the dataset page: https://huggingface.co/datasets/TajikNLPWorld/tajik-spelling-correction-pairs.asr_spellingCorrection_24kvie_spelling_correction_datavie_spelling_correction_table_dataspelling_correction_datasetvie_spelling_correction_data_v2vie_spelling_correction_data_v4
