datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
librispeech_lmLanguage modeling resources to be used in conjunction with the LibriSpeech ASR corpus.openslr-sinhala-synthetic-spell-errors-quarter
Sinhala Dyslexic Spelling Correction Dataset
Dataset Description
This dataset contains Sinhala and code-mixed (Sinhala-English) text pairs for training spelling correction models, specifically designed to address dyslexia-like spelling errors.
Features
dyslexic_sentence: Input text with dyslexia-like spelling errors (string)
correct_sentence: Corrected output text (string)
Dataset Statistics
Split
Samples
Train
37,056
Test
9,265… See the full description on the dataset page: https://huggingface.co/datasets/SPEAK-PP/openslr-sinhala-synthetic-spell-errors-quarter.
