CoolFace
Datasetpublic

alphazhan/sauatai-ertegiler-kz-misspellings-kk-s170-len60-n6-m1-v1

SauatAI β€” Kazakh Misspelled Sentences from Ertegiler.kz SauatAI is a grammar-focused dataset built from 170 children’s stories scraped from ertegiler.kz on July 5, 2025. The dataset was designed to support Kazakh language grammar correction, error detection, and text augmentation research. πŸ“Œ Dataset Details s170 β€” 170 unique stories were scraped and sentence-tokenized. len60 β€” Only sentences with ≀60 characters were retained. n6 β€” Each correct sentence has 5… See the full description on the dataset page: https://huggingface.co/datasets/alphazhan/sauatai-ertegiler-kz-misspellings-kk-s170-len60-n6-m1-v1.

sourceHugging Facecc-by-4.0updated 1y agoView on Hugging Face
0likes21downloads
Dataset Card

SauatAI β€” Kazakh Misspelled Sentences from Ertegiler.kz

SauatAI is a grammar-focused dataset built from 170 children’s stories scraped from ertegiler.kz on July 5, 2025. The dataset was designed to support Kazakh language grammar correction, error detection, and text augmentation research.

πŸ“Œ Dataset Details

  • β€”s170 β€” 170 unique stories were scraped and sentence-tokenized.
  • β€”len60 β€” Only sentences with ≀60 characters were retained.
  • β€”n6 β€” Each correct sentence has 5 corresponding corrupted variants + 1 clean version.
  • β€”m1 β€” Each corrupted sentence contains exactly 1 mistake:
  • β€”The mistake is either a letter substitution (from a predefined spelling swap map) or a character omission.
  • β€”The type is chosen randomly per sentence.
  • β€”Ensures consistency and control over the number of errors per sample.

The final dataset consists of three CSV files:

  • β€”train.csv
  • β€”val.csv
  • β€”test.csv

Each file contains sentence-level pairs of grammatically correct and corrupted Kazakh text.

This version includes a total of 38,099 sentence pairs β€” 5 corrupted variants + 1 correct sentence per source (n6 setting).

πŸ§ͺ Dataset Splits

SplitCountFile
Train25,399train.csv
Validation6,350val.csv
Test6,350test.csv

πŸ“„ Column Descriptions

Column NameDescription
story_idUnique ID of the source story
story_titleTitle of the story the sentence came from
sentence_idIndex of the sentence within the story
correct_sentence_lenLength (character count) of the correct sentence
correct_sentenceThe grammatically correct version of the sentence
corrupted_sentenceA noisy version with potential spelling or character-level errors
label0 if corrupted, 1 if correct

🧠 Intended Use

  • β€”Grammar correction model pretraining/fine-tuning
  • β€”Corrupted-to-correct sequence generation
  • β€”Evaluation of large language models on noisy Kazakh text
  • β€”Robustness testing and error correction research

πŸ“œ License

This dataset is distributed under the Creative Commons Attribution 4.0 International (CC BY 4.0) license. You are free to use, share, and adapt the dataset β€” just provide attribution.