datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vi-spelling-correction
Vietnamese Spelling Correction Dataset
This dataset contains 978,417 pairs of noisy (source) and clean (target) Vietnamese sentences, designed for training spelling correction models.
The dataset was synthetically generated by injecting realistic noise into a clean Vietnamese corpus.
Dataset Structure
The dataset is divided into training and testing sets:
Train: 880,575 examples
Test: 97,842 examples
Data Fields
source: The text with injected errors (input).… See the full description on the dataset page: https://huggingface.co/datasets/coung21/vi-spelling-correction.spelling-correction-french-news
Spelling correction dataset (French)
This dataset is generated by transforming/corrupting sentences of a French news corpus
provided by the University of Leipzig.
The following transformations are applied to words in the sentences:
concatenation of pairs of words
swapping of neighboring letters in words
insertion
deletion
replacement (by neighboring characters in AZERTY keyboard)
Generation
./scripts/get_data.py -t news -y 2023 -s 10K
./scripts/generate_dataset.py… See the full description on the dataset page: https://huggingface.co/datasets/fdemelo/spelling-correction-french-news.khmer-spelling-corrections
Khmer Spelling Corrections
Naturally occurring Khmer misspellings paired with the word the writer meant.
The labels are not annotated, they are observed. Search sessions record the
whole typing trajectory toward a single word, so when a user types something,
fails, adjusts and lands on a real dictionary headword, the failed attempt and
the headword form a correction pair produced by a real person under no
instruction to make mistakes.
Only pairs within two edits of the target… See the full description on the dataset page: https://huggingface.co/datasets/seanghay/khmer-spelling-corrections.correct-incorrect-spelling-pairsThis is a dataset containing correct and incorrect spelling pairs in Gujarati, created by us using artificial noise.
Turkish-Spelling-Dictionarynlp_vietnamese_spellingspelling-puzzlesnlp_vietnamese_spelling_v2vietnamese_spelling_error_detectionbulgarian-spelling-mistakes
Dataset of Bulgarian Spelling Mistakes
Dataset Summary
This is a dataset of sentences in Bulgarian with spelling mistakes created by automatically inducing errors in correct sentences.
Supported Tasks
text2text-generation: The dataset can be used to train a model for spelling error correction, which consists in correction of spelling errors. in a source sentence, resulting in a correct version.
Languages
bg: Only Bulgarian is supported by this… See the full description on the dataset page: https://huggingface.co/datasets/thebogko/bulgarian-spelling-mistakes.scoutieDataset_chinese_russian_dictionary_grammar_spelling_vectorized
Description in English:
A dataset collected from 30 Russian-language Telegram channels on the topic of learning Chinese, this dataset contains grammar, syntax, spelling and punctuation rules, as well as Chinese words with Russian translations.
The dataset was collected and marked automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link.
Dataset fields:
taskId - task identifier in the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/scoutieDataset_chinese_russian_dictionary_grammar_spelling_vectorized.scoutieDataset_english_russian_dictionary_grammar_spelling_vectorized
Description in English:
A dataset collected from 30 Russian-language Telegram channels on the topic of learning English, this dataset contains grammar, syntax, spelling and punctuation rules, as well as English words with Russian translations.
The dataset was collected and marked automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link.
Dataset fields:
taskId - task identifier in the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/scoutieDataset_english_russian_dictionary_grammar_spelling_vectorized.Spelling_correctionnlapug_spellingvietnamese-spelling-correction-datasetspelling_correction_dataset
