datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gec-test-setTest data for Icelandic spell and grammar checking, created as part of the Icelandic Language Technology Programme.
The test data is divided into three different formats, type 1, 2 and 3. For every original file corrected, three files are included in the test data when possible: _original, _corrected and _metadata. The original and metadata files are always .txt files, but the format of the corrected file differs between types.
Texts corrected are from the News2 subcorpus of the Icelandic… See the full description on the dataset page: https://huggingface.co/datasets/mideind/gec-test-set.nepali-gector-style-token-level-tag-for-ged
Nepali GEC (gector style) Token Tagging Dataset
This is a processed version of the sumitaryal/nepali_grammatical_error_correction dataset,
designed for training GEC-ToR-style sequence tagging models.
This dataset has been processed with a robust, multi-pass, content-aware alignment algorithm
to generate high-fidelity correction tags, including complex and adjacent SWAP operations.
Total Examples: 16,260,992
Training: 13,008,711
Validation: 2,439,231
test: 813,050… See the full description on the dataset page: https://huggingface.co/datasets/DipeshChaudhary/nepali-gector-style-token-level-tag-for-ged.muril-nepali-gector-style-token-level-tag-for-ged
Nepali GEC (gector style) Token Tagging Dataset
This is a processed version of the sumitaryal/nepali_grammatical_error_correction dataset,
designed for training GEC-ToR-style sequence tagging models.
This dataset has been processed with a robust, multi-pass, content-aware alignment algorithm
to generate high-fidelity correction tags, including complex and adjacent SWAP operations.
Total Examples: 16,260,992
Training: 13,008,711
Validation: 2,439,231
test: 813,050… See the full description on the dataset page: https://huggingface.co/datasets/DipeshChaudhary/muril-nepali-gector-style-token-level-tag-for-ged.geco2-assetsc4_200m-gec-train100k-test25k
Dataset Card for "c4_200m-gec-train100k-test25k"
More Information needed
ask-gec
Norwegian grammatical error correction (ASK)
This is the ASK-RAW dataset by Matias Jentoft (2023).
Cite
@mastersthesis{jentoft2023grammatical,
title={Grammatical Error Correction with byte-level language models},
author={Jentoft, Matias},
year={2023},
school={University of Oslo},
url={https://www.duo.uio.no/handle/10852/103885}
}
gec
Luganda Grammar Error Correction Dataset
A synthetic dataset for training and evaluating Luganda Grammar Error Correction (GEC) models.
Dataset Description
Overview
The dataset consists of pairs of:
src: a corrupted Luganda sentence
tgt: the corresponding original/correct Luganda sentence
Corruptions are generated from clean Luganda text using linguistically informed corruption operations.
The objective is to train models to transform an erroneous… See the full description on the dataset page: https://huggingface.co/datasets/sagepond/gec.lang-uk-fiction-gec-dialogs
Dataset Card for Ukrainian Fiction Grammatical Error Correction Dialogs
Dataset Description
Dataset Summary
This dataset is a processed version of fiction part of the lang-uk UberText Corpus. The goal for this dataset is to provide grammatical error correction knowledge grounding.
Languages
Ukrainian (uk)
Data Fields
instruction: Text containing task description
input: Processed text from the original text, including grammar errors
output: Correct text
task_type:… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/lang-uk-fiction-gec-dialogs.english-gec-tatoebamultilingual-gec
Dataset Card for Multilingual Grammar Error Correction
Dataset Summary
This dataset can be used to train a transformer model (we used T5) to correct grammar errors in simple sentences written in English, Spanish, French, or German.
This dataset was developed as a component for the Squidigies platform.
Supported Tasks and Leaderboards
Grammar Error Correction: By appending the prefix fix grammar: to the prrompt.
Language Detection: By appending the prefix:… See the full description on the dataset page: https://huggingface.co/datasets/juancavallotti/multilingual-gec.GECTurk-generationHomepage: https://github.com/GGLAB-KU/gecturk/
GEC25_C4english-gec-tatoeba-finetunegeco_data_generatorC4-200M-1M-GEC-Determinerdetails_NeuralNovel__Gecko-7B-v0.1
Dataset Card for Evaluation run of NeuralNovel/Gecko-7B-v0.1
Dataset automatically created during the evaluation run of model NeuralNovel/Gecko-7B-v0.1 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_NeuralNovel__Gecko-7B-v0.1.gec
Lekhya · Bangla GEC (synthetic)
(incorrect → correct) Bangla sentence pairs with typed, span-level edits, for
grammatical error correction. Built by Lekhya.
⚠️ Read this before training on it
These errors are synthetic. They were produced by rule-based injection into clean
newspaper prose. A model trained here learns to invert this noise function, which is
not the same as learning to correct Bangla. Expect strong scores on a test set built by
the same rules and… See the full description on the dataset page: https://huggingface.co/datasets/lekhya-ai/gec.russian_gec
📝 Dataset Card for Russian Grammar Error-Correction (25 362 sentence pairs)
A compact, high-quality corpus of Russian sentences with grammatical errors aligned to their human‑corrected counterparts. Ideal for training and benchmarking grammatical error‑correction (GEC) models, writing assistants, and translation post‑editing.
✨ Dataset Summary
Metric
Value
Sentence pairs
25 362
Avg. tokens / sentence
≈ 12
File size
~5 MB (CSV, UTF‑8)
Error types… See the full description on the dataset page: https://huggingface.co/datasets/dreuxx26/russian_gec.c4ai-takehome-gec-trainnewscorpus-for-gecHealthQAWikiCorrupted_spanish_to_GEC-GED_Lfluency-pairs-gec-only100-gecsThis dataset is designed to generate lyrics with HuggingArtists.cs_gec
Introduction
This dataset is extracted by postprocessing data from Náplava et al., 2019. Specificially, we extracted gramatically incorrect sentences, and their respective corrections.
Then we convert task to binary detection of errorneous sentences. We downloaded the original dataset from LINDAT-Clarin repository.
Citation
@inproceedings{naplava-straka-2019-grammatical,
title = "Grammatical Error Correction in Low-Resource Scenarios",
author = "N{\'a}plava, Jakub… See the full description on the dataset page: https://huggingface.co/datasets/CZLC/cs_gec.GECTurkCOWSL2H_GEC_cleanedgalician-gec-corpora
Galician GEC Corpora
Dataset Summary
Galician GEC Corpora is a collection of Galician grammatical and orthographic correction datasets. The repository groups several sources of sentence-level correction pairs in a single Hugging Face dataset repository, with each source exposed as a separate configuration.
Each instance contains an incorrect or non-standard Galician sentence and its corrected version. Some subsets also include error labels, correction tags, edit… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/galician-gec-corpora.gec-ro-textsREALEC_GEC_dataset_ACL_test
