CoolFace
16 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01coung21 /vi-spelling-correction Vietnamese Spelling Correction Dataset This dataset contains 978,417 pairs of noisy (source) and clean (target) Vietnamese sentences, designed for training spelling correction models. The dataset was synthetically generated by injecting realistic noise into a clean Vietnamese corpus. Dataset Structure The dataset is divided into training and testing sets: Train: 880,575 examples Test: 97,842 examples Data Fields source: The text with injected errors (input).… See the full description on the dataset page: https://huggingface.co/datasets/coung21/vi-spelling-correction.texttext-generation100K<n<1M1 likes79 downloads8mo agoHugging Face02fdemelo /spelling-correction-french-news Spelling correction dataset (French) This dataset is generated by transforming/corrupting sentences of a French news corpus provided by the University of Leipzig. The following transformations are applied to words in the sentences: concatenation of pairs of words swapping of neighboring letters in words insertion deletion replacement (by neighboring characters in AZERTY keyboard) Generation ./scripts/get_data.py -t news -y 2023 -s 10K ./scripts/generate_dataset.py… See the full description on the dataset page: https://huggingface.co/datasets/fdemelo/spelling-correction-french-news.text10K<n<100K1 likes74 downloads1y agoHugging Face03seanghay /khmer-spelling-corrections Khmer Spelling Corrections Naturally occurring Khmer misspellings paired with the word the writer meant. The labels are not annotated, they are observed. Search sessions record the whole typing trajectory toward a single word, so when a user types something, fails, adjusts and lands on a real dictionary headword, the failed attempt and the headword form a correction pair produced by a real person under no instruction to make mistakes. Only pairs within two edits of the target… See the full description on the dataset page: https://huggingface.co/datasets/seanghay/khmer-spelling-corrections.tabulartext-generationn<1K2 likes35 downloads1mo agoHugging Face04autopilot-ai /correct-incorrect-spelling-pairsThis is a dataset containing correct and incorrect spelling pairs in Gujarati, created by us using artificial noise. texttext-classification100K<n<1M0 likes25 downloads3y agoHugging Face05asimokby /Turkish-Spelling-Dictionarytext100K<n<1M2 likes25 downloads2y agoHugging Face06VoTrongTinh /nlp_vietnamese_spellingtext10K<n<100K0 likes22 downloads2y agoHugging Face07isaiahbjork /spelling-puzzlestext1K<n<10K0 likes20 downloads2y agoHugging Face08VoTrongTinh /nlp_vietnamese_spelling_v2text10K<n<100K0 likes17 downloads2y agoHugging Face09PaulTran /vietnamese_spelling_error_detectiontext100K<n<1M1 likes14 downloads3y agoHugging Face10thebogko /bulgarian-spelling-mistakes Dataset of Bulgarian Spelling Mistakes Dataset Summary This is a dataset of sentences in Bulgarian with spelling mistakes created by automatically inducing errors in correct sentences. Supported Tasks text2text-generation: The dataset can be used to train a model for spelling error correction, which consists in correction of spelling errors. in a source sentence, resulting in a correct version. Languages bg: Only Bulgarian is supported by this… See the full description on the dataset page: https://huggingface.co/datasets/thebogko/bulgarian-spelling-mistakes.text10K<n<100K3 likes14 downloads2y agoHugging Face11ScoutieAutoML /scoutieDataset_chinese_russian_dictionary_grammar_spelling_vectorized Description in English: A dataset collected from 30 Russian-language Telegram channels on the topic of learning Chinese, this dataset contains grammar, syntax, spelling and punctuation rules, as well as Chinese words with Russian translations. The dataset was collected and marked automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link. Dataset fields: taskId - task identifier in the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/scoutieDataset_chinese_russian_dictionary_grammar_spelling_vectorized.tabulartext-classification1K<n<10K0 likes14 downloads2y agoHugging Face12ScoutieAutoML /scoutieDataset_english_russian_dictionary_grammar_spelling_vectorized Description in English: A dataset collected from 30 Russian-language Telegram channels on the topic of learning English, this dataset contains grammar, syntax, spelling and punctuation rules, as well as English words with Russian translations. The dataset was collected and marked automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link. Dataset fields: taskId - task identifier in the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/scoutieDataset_english_russian_dictionary_grammar_spelling_vectorized.tabulartext-classification10K<n<100K0 likes11 downloads2y agoHugging Face13KrugDen /Spelling_correctiontext1K<n<10K0 likes10 downloads1y agoHugging Face14wikd /nlapug_spellingtextn<1K0 likes6 downloads2y agoHugging Face15iAmHieu2012 /vietnamese-spelling-correction-datasettext100K<n<1M0 likes4 downloads9mo agoHugging Face16Alakir11 /spelling_correction_datasettext1K<n<10K1 likes1 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.