datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GECTurk-generationHomepage: https://github.com/GGLAB-KU/gecturk/
russian_gec
📝 Dataset Card for Russian Grammar Error-Correction (25 362 sentence pairs)
A compact, high-quality corpus of Russian sentences with grammatical errors aligned to their human‑corrected counterparts. Ideal for training and benchmarking grammatical error‑correction (GEC) models, writing assistants, and translation post‑editing.
✨ Dataset Summary
Metric
Value
Sentence pairs
25 362
Avg. tokens / sentence
≈ 12
File size
~5 MB (CSV, UTF‑8)
Error types… See the full description on the dataset page: https://huggingface.co/datasets/dreuxx26/russian_gec.Turkish-OSCAR-GECFrench_GEC
[!NOTE]
Dataset origin: https://www.kaggle.com/datasets/isakbiderre/french-gec-dataset
Context
Wikipedia is a free encyclopedia where everyone can contribute and modify, delete, or add text to the articles.
Because of this, every day there is newly created text and, most importantly, new corrections made to preexisting sentences.
The idea is to find the corrections made to these sentences and create a dataset with X,y sentence pairs.
The data
This dataset was created… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/French_GEC.poc-gecTurkish-GPT-GECgec-cleanednepali_gec_data_v3istanbul_gecmis_havadurumunepali_gec_data_v4multilingual-gecgec-arabicTurkish-OSCAR-GEC
