alphazhan/sauatai-ertegiler-kz-misspellings-kk-s170-len60-n6-m1-v1
SauatAI β Kazakh Misspelled Sentences from Ertegiler.kz SauatAI is a grammar-focused dataset built from 170 childrenβs stories scraped from ertegiler.kz on July 5, 2025. The dataset was designed to support Kazakh language grammar correction, error detection, and text augmentation research. π Dataset Details s170 β 170 unique stories were scraped and sentence-tokenized. len60 β Only sentences with β€60 characters were retained. n6 β Each correct sentence has 5β¦ See the full description on the dataset page: https://huggingface.co/datasets/alphazhan/sauatai-ertegiler-kz-misspellings-kk-s170-len60-n6-m1-v1.
SauatAI β Kazakh Misspelled Sentences from Ertegiler.kz
SauatAI is a grammar-focused dataset built from 170 childrenβs stories scraped from ertegiler.kz on July 5, 2025. The dataset was designed to support Kazakh language grammar correction, error detection, and text augmentation research.
π Dataset Details
- s170 β 170 unique stories were scraped and sentence-tokenized.
- len60 β Only sentences with β€60 characters were retained.
- n6 β Each correct sentence has 5 corresponding corrupted variants + 1 clean version.
- m1 β Each corrupted sentence contains exactly 1 mistake:
- The mistake is either a letter substitution (from a predefined spelling swap map) or a character omission.
- The type is chosen randomly per sentence.
- Ensures consistency and control over the number of errors per sample.
The final dataset consists of three CSV files:
train.csvval.csvtest.csv
Each file contains sentence-level pairs of grammatically correct and corrupted Kazakh text.
This version includes a total of 38,099 sentence pairs β 5 corrupted variants + 1 correct sentence per source (n6 setting).
π§ͺ Dataset Splits
π Column Descriptions
π§ Intended Use
- Grammar correction model pretraining/fine-tuning
- Corrupted-to-correct sequence generation
- Evaluation of large language models on noisy Kazakh text
- Robustness testing and error correction research
π License
This dataset is distributed under the Creative Commons Attribution 4.0 International (CC BY 4.0) license. You are free to use, share, and adapt the dataset β just provide attribution.
