CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nguyenthanhasia /vsec-vietnamese-spell-correction VSEC: Vietnamese Spell Correction Dataset Dataset Description VSEC (Vietnamese Spell Correction) is a comprehensive dataset for Vietnamese spelling error detection and correction, containing 9,341 sentences with 11,202 human-made misspellings across 5,211 unique error types. This dataset represents the largest publicly available collection of Vietnamese spelling errors with syllable-level annotations, making it an invaluable resource for developing and evaluating… See the full description on the dataset page: https://huggingface.co/datasets/nguyenthanhasia/vsec-vietnamese-spell-correction.texttext-generation1K<n<10K5 likes118 downloads1y agoHugging Face02nrl-ai /vn-spell-correction-eval-real vn-spell-correction-eval-real Out-of-distribution evaluation corpus for Vietnamese spell-correction models — 150 hand-curated (noisy, clean) pairs sampled from real VN error sources, not generated by nom.text.noise. This is the test set we use to verify a spell-correction model generalises beyond its own synthetic training distribution. A model that scores 95 % on nom-vn's synthetic eval grid and 60 % on this set is overfit to the noise generator. Splits Config… See the full description on the dataset page: https://huggingface.co/datasets/nrl-ai/vn-spell-correction-eval-real.texttext-generationn<1K0 likes81 downloads5mo agoHugging Face03Khamoon /asr-spell-correction-rutext1K<n<10K0 likes72 downloads10d agoHugging Face04melsmm /spell-correction-ru Spell Correction RU — датасеты для коррекции ошибок в русском тексте Набор данных для обучения моделей исправления орфографических, пунктуационных и регистровых ошибок в русскоязычных текстах. Каждый пример — пара «правильный текст» → «текст с ошибкой». Датасет использовался для обучения модели melsmm/Spell-Corrector-RU-4B. 📦 Код генерации, ноутбуки и полное описание проекта: github.com/melsmm/llm-spell-corrector Состав Датасет содержит две конфигурации… See the full description on the dataset page: https://huggingface.co/datasets/melsmm/spell-correction-ru.texttext-generation1M<n<10M1 likes68 downloads4mo agoHugging Face05chainiy /russian-spell-correctionstext1K<n<10K1 likes58 downloads2d agoHugging Face06Dan032 /asr_spell_correction_rutextn<1K0 likes46 downloads10d agoHugging Face07elinaail /asr-spell-correction-rutext1K<n<10K0 likes40 downloads5d agoHugging Face08nrl-ai /vn-spell-correction-train nrl-ai/vn-spell-correction-train 459,478 (noisy, clean) Vietnamese training pairs for fine-tuning a seq2seq spell-correction model. Each row: {"input": "<noisy>", "target": "<clean>"} Both fields are NFC-normalized. How it was built Clean side: same 500K register-balanced mix as nrl-ai/vn-diacritic-train — 350K Vietnamese Wikipedia (CC-BY-SA-4.0, hirine/wikipedia-vietnamese-1M296K-dataset) + 150K NFC-fixed Vietnamese news (CC-BY-4.0, tmnam20/Vietnamese-News-dedup).… See the full description on the dataset page: https://huggingface.co/datasets/nrl-ai/vn-spell-correction-train.texttext-generation100K<n<1M0 likes38 downloads5mo agoHugging Face09aligh4699 /persian-spell-correction-dataset Persian Spell Correction & Augmentation Dataset This is a large-scale, parallel dataset for Persian spell correction, text normalization, and augmentation. It is designed to train and evaluate models for correcting a wide variety of common and synthetic errors in Persian text. The dataset is built from two main components: Natural Data: Text from diverse Persian corpora and its corresponding clean, corrected version (corrected_text) generated by an LLM. Augmented Data: The… See the full description on the dataset page: https://huggingface.co/datasets/aligh4699/persian-spell-correction-dataset.text1M<n<10M1 likes35 downloads10mo agoHugging Face10sobadsodead /asr-spell-correction-ru-hw01 Russian ASR correction: homework 01 1020 pairs: 450 Groq-generated ASR-like inputs, 180 Groq-generated numeral-to-word pairs, and 390 identity examples added by copying screened clean targets. Model: openai/gpt-oss-120b. Generation: 8bae1a605e8506a1; prompt version: groq_asr_numbers_v2. The ASR target names come from the existing Groq-generated pool targets.jsonl. No Python character corruption is used. This is synthetic text, not real ASR output. Generation and… See the full description on the dataset page: https://huggingface.co/datasets/sobadsodead/asr-spell-correction-ru-hw01.text1K<n<10K0 likes35 downloads10d agoHugging Face11Alexander-Usov /ru-asr-spell-correctiontext1K<n<10K0 likes33 downloads4d agoHugging Face12DariaZah /asr_spell_correction_rutext1K<n<10K0 likes32 downloads3d agoHugging Face13Alexander-Usov /ru-asr-spell-correction-v2text1K<n<10K0 likes31 downloads4d agoHugging Face14nrl-ai /vn-spell-correction-eval nrl-ai/vn-spell-correction-eval Vietnamese spell-correction evaluation grid: 4 source registers × 2 noise levels = 8 splits, 2,098 (noisy, clean) sentence pairs total. Each pair is {"input": "<noisy>", "target": "<clean>"}. Both sides are NFC-normalized. The clean target is the same sentence used as the target in nrl-ai/vn-diacritic-eval — spell correction is a strict superset of diacritic restoration, so we reuse the same registers-balanced corpus. Splits Two noise… See the full description on the dataset page: https://huggingface.co/datasets/nrl-ai/vn-spell-correction-eval.texttext-generation1K<n<10K0 likes28 downloads5mo agoHugging Face15torinriley /spell-correction Spell-Check Dataset This dataset consists of pairs of misspelled words and their corresponding correctly spelled words, designed for training and evaluating character-level spelling correction models. It is particularly useful for tasks such as: Spelling correction Character-level sequence-to-sequence modeling Error detection and correction in text Each data point in the dataset contains: misspelled: A misspelled version of a word. correct: The corrected spelling of the word.… See the full description on the dataset page: https://huggingface.co/datasets/torinriley/spell-correction.text10K<n<100K2 likes25 downloads2y agoHugging Face16SerejkaP /popular_names_spell_correctiontext1K<n<10K1 likes24 downloads1y agoHugging Face17rubin5341 /my-groq-spell-correction-mistake-datasettext1K<n<10K0 likes17 downloads1y agoHugging Face18NChechulin /russian-movie-spell-correctiontext1K<n<10K0 likes17 downloads1y agoHugging Face19rubin5341 /groq-spell-correction-mistake-datasettext1K<n<10K0 likes15 downloads1y agoHugging Face20nikfil /russian-spell-correction-datasettext1K<n<10K0 likes15 downloads11mo agoHugging Face21DenK-huggingFace /russian-spell-correction-dataset-2text1K<n<10K0 likes14 downloads1y agoHugging Face22rubin5341 /groq1441-spell-correction-mistake-datasettext1K<n<10K0 likes14 downloads1y agoHugging Face23Andrey32 /russian-spell-correction-groqtext1K<n<10K0 likes14 downloads11mo agoHugging Face24gfrslf /russian_spell_correction_groqtext1K<n<10K0 likes13 downloads1y agoHugging Face25polinOchka33 /spell_correction_rutext1K<n<10K0 likes13 downloads1d agoHugging Face26Lak1N /spell_correction_datasettext1K<n<10K0 likes11 downloads1y agoHugging Face27erickrribeiro /spell_correction_datasets_pt_brtextn<1K0 likes10 downloads3y agoHugging Face28minhbui /spell_correction_datasettext1M<n<10M0 likes10 downloads2y agoHugging Face29nesemenpolkov /syntetic-dataset-for-gigaam-spell-correctionaudio1K<n<10K0 likes9 downloads2y agoHugging Face30nlprunnerup /russian-spell-correction-datasettext1K<n<10K0 likes9 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.