datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ask-gec
Norwegian grammatical error correction (ASK)
This is the ASK-RAW dataset by Matias Jentoft (2023).
Cite
@mastersthesis{jentoft2023grammatical,
title={Grammatical Error Correction with byte-level language models},
author={Jentoft, Matias},
year={2023},
school={University of Oslo},
url={https://www.duo.uio.no/handle/10852/103885}
}
cs_gec
Introduction
This dataset is extracted by postprocessing data from Náplava et al., 2019. Specificially, we extracted gramatically incorrect sentences, and their respective corrections.
Then we convert task to binary detection of errorneous sentences. We downloaded the original dataset from LINDAT-Clarin repository.
Citation
@inproceedings{naplava-straka-2019-grammatical,
title = "Grammatical Error Correction in Low-Resource Scenarios",
author = "N{\'a}plava, Jakub… See the full description on the dataset page: https://huggingface.co/datasets/CZLC/cs_gec.galician-gec-corpora
Galician GEC Corpora
Dataset Summary
Galician GEC Corpora is a collection of Galician grammatical and orthographic correction datasets. The repository groups several sources of sentence-level correction pairs in a single Hugging Face dataset repository, with each source exposed as a separate configuration.
Each instance contains an incorrect or non-standard Galician sentence and its corrected version. Some subsets also include error labels, correction tags, edit… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/galician-gec-corpora.ua_gec_instruction_tuning
UA-GEC instruction tuning
This dataset contains prompts and expected outputs for the grammatical error
correction task in the Ukrainian language. It is based on the
CC-BY-4.0-licensed UA-GEC dataset. The
license of the original data is CC-BY-4.0.
This dataset contains 1,700 examples of fixing errors in long documents, and
~28,000 sentence-level examples.
The instructions ask to correct errors in the text. Sometimes the model outputs
the corrected text as is. At other times, it… See the full description on the dataset page: https://huggingface.co/datasets/osyvokon/ua_gec_instruction_tuning.ru-gec-v1ru_gec_v1
Dataset: RuGECv1 (Russian Grammatical Error Correction)
This dataset contains 707,261 parallel examples specifically curated for training Grammatical Error Correction (GEC) models for the Russian language. The dataset follows an instruction-tuning format, making it suitable for fine-tuning instruction-following language models.
Dataset Structure
Each example contains the following fields:
Field
Type
Description
input
string
Source text containing grammatical… See the full description on the dataset page: https://huggingface.co/datasets/alexanderpl/ru_gec_v1.est_gec_ged_gee_2025icelandic-sentences-gecru-gec-v1
Dataset: RuGECv1 (Russian Grammatical Error Correction)
This dataset contains 707,261 parallel examples specifically curated for training Grammatical Error Correction (GEC) models for the Russian language. The dataset follows an instruction-tuning format, making it suitable for fine-tuning instruction-following language models.
Dataset Structure
Each example contains the following fields:
Field
Type
Description
input
string
Source text containing grammatical… See the full description on the dataset page: https://huggingface.co/datasets/p1746-lingua/ru-gec-v1.gec_dala_tv2r_itsozkz-gec-synthetic-gpt4o-v1
SozKZ GEC Synthetic GPT-4o v1
Синтетикалық қазақ тілі грамматикалық қателерді түзету (GEC) датасеті. GPT-4o арқылы генерацияланған.
Synthetic Kazakh Grammatical Error Correction dataset generated with OpenAI GPT-4o.
Overview
Parameter
Value
Total examples
9,599
Synthetic error pairs
7,204
Identity (clean) examples
2,395 (25%)
Source model
GPT-4o
Seed source
Kazakh Wikipedia
Language
Kazakh (kk)
License
MIT
Data Format… See the full description on the dataset page: https://huggingface.co/datasets/stukenov/sozkz-gec-synthetic-gpt4o-v1.Group_project_gecbh_2025_trvGecko_Minden_nl_data_geco
