datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vsec-vietnamese-spell-correction
VSEC: Vietnamese Spell Correction Dataset
Dataset Description
VSEC (Vietnamese Spell Correction) is a comprehensive dataset for Vietnamese spelling error detection and correction, containing 9,341 sentences with 11,202 human-made misspellings across 5,211 unique error types. This dataset represents the largest publicly available collection of Vietnamese spelling errors with syllable-level annotations, making it an invaluable resource for developing and evaluating… See the full description on the dataset page: https://huggingface.co/datasets/nguyenthanhasia/vsec-vietnamese-spell-correction.asr-spell-correction-ruasr_spell_correction_ruru-asr-spell-correctionru-asr-spell-correction-v2asr-spell-correction-ruasr_spell_correction_ruasr-spell-correction-ru-hw01
Russian ASR correction: homework 01
1020 pairs: 450 Groq-generated ASR-like inputs,
180 Groq-generated numeral-to-word pairs, and 390 identity
examples added by copying screened clean targets.
Model: openai/gpt-oss-120b. Generation: 8bae1a605e8506a1; prompt version: groq_asr_numbers_v2.
The ASR target names come from the existing Groq-generated pool targets.jsonl.
No Python character corruption is used. This is synthetic text, not real ASR output.
Generation and… See the full description on the dataset page: https://huggingface.co/datasets/sobadsodead/asr-spell-correction-ru-hw01.persian-spell-correction-dataset
Persian Spell Correction & Augmentation Dataset
This is a large-scale, parallel dataset for Persian spell correction, text normalization, and augmentation. It is designed to train and evaluate models for correcting a wide variety of common and synthetic errors in Persian text.
The dataset is built from two main components:
Natural Data: Text from diverse Persian corpora and its corresponding clean, corrected version (corrected_text) generated by an LLM.
Augmented Data: The… See the full description on the dataset page: https://huggingface.co/datasets/aligh4699/persian-spell-correction-dataset.russian-asr-spell-correctionspell_correction_rupopular_names_spell_correctionrussian-asr-spell-correction_v2my-groq-spell-correction-mistake-datasetgroq-spell-correction-mistake-datasetrussian-spell-correction-datasetrussian-asr-spell-correction_v3russian-spell-correction-dataset-2grok_spell_correction_datagroq1441-spell-correction-mistake-datasetrussian-spell-correction-groqrussian_spell_correction_groqspell_correction_datasetspell_correction_datasetsyntetic-dataset-for-gigaam-spell-correctionrussian-spell-correction-datasetrussian-spell-correction-datasetrussian-spell-correction-datasetrussian-spell-correction-datasetrussian-asr-spell-correction_vers2
