datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
c4_200m-gec-train100k-test25k
Dataset Card for "c4_200m-gec-train100k-test25k"
More Information needed
gec
Luganda Grammar Error Correction Dataset
A synthetic dataset for training and evaluating Luganda Grammar Error Correction (GEC) models.
Dataset Description
Overview
The dataset consists of pairs of:
src: a corrupted Luganda sentence
tgt: the corresponding original/correct Luganda sentence
Corruptions are generated from clean Luganda text using linguistically informed corruption operations.
The objective is to train models to transform an erroneous… See the full description on the dataset page: https://huggingface.co/datasets/sagepond/gec.lang-uk-fiction-gec-dialogs
Dataset Card for Ukrainian Fiction Grammatical Error Correction Dialogs
Dataset Description
Dataset Summary
This dataset is a processed version of fiction part of the lang-uk UberText Corpus. The goal for this dataset is to provide grammatical error correction knowledge grounding.
Languages
Ukrainian (uk)
Data Fields
instruction: Text containing task description
input: Processed text from the original text, including grammar errors
output: Correct text
task_type:… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/lang-uk-fiction-gec-dialogs.english-gec-tatoebamultilingual-gec
Dataset Card for Multilingual Grammar Error Correction
Dataset Summary
This dataset can be used to train a transformer model (we used T5) to correct grammar errors in simple sentences written in English, Spanish, French, or German.
This dataset was developed as a component for the Squidigies platform.
Supported Tasks and Leaderboards
Grammar Error Correction: By appending the prefix fix grammar: to the prrompt.
Language Detection: By appending the prefix:… See the full description on the dataset page: https://huggingface.co/datasets/juancavallotti/multilingual-gec.GEC25_C4english-gec-tatoeba-finetuneC4-200M-1M-GEC-Determinergec
Lekhya · Bangla GEC (synthetic)
(incorrect → correct) Bangla sentence pairs with typed, span-level edits, for
grammatical error correction. Built by Lekhya.
⚠️ Read this before training on it
These errors are synthetic. They were produced by rule-based injection into clean
newspaper prose. A model trained here learns to invert this noise function, which is
not the same as learning to correct Bangla. Expect strong scores on a test set built by
the same rules and… See the full description on the dataset page: https://huggingface.co/datasets/lekhya-ai/gec.c4ai-takehome-gec-trainnewscorpus-for-gecWikiCorrupted_spanish_to_GEC-GED_LHealthQAfluency-pairs-gec-onlyGECTurkCOWSL2H_GEC_cleanedREALEC_GEC_dataset_ACL_testgec-coherence-coedit-synthWikiCorrupted-spanish_to_GEC-GEDREALEC_GEC_dataset_ACLC4-200M-1.55M-GEC-Determinercodeit_gecfr-gec-dataset
Dataset Card for "fr-gec-dataset"
More Information needed
gec-targeted-corrections-esl
GEC Targeted Corrections — ESL
An LLM-generated grammatical error correction (GEC) dataset targeting the specific error patterns that ESL learners most commonly produce. Each example is a (src, tgt) pair where src contains a realistic grammatical error and tgt is the minimally corrected version: only what is necessary is changed.
Dataset Summary
Split
Examples
train
2,037
Schema
{
"src": "She gave me some advices about the… See the full description on the dataset page: https://huggingface.co/datasets/martinsr/gec-targeted-corrections-esl.c4ai_csp-gec_datasetsfluency-pairs-gec-onklyREALEC_GEC_datasetgec-datasetnewscorpus-for-gec-smallSmolLM-135M-GEC-preference-dataset
