datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gec
Luganda Grammar Error Correction Dataset
A synthetic dataset for training and evaluating Luganda Grammar Error Correction (GEC) models.
Dataset Description
Overview
The dataset consists of pairs of:
src: a corrupted Luganda sentence
tgt: the corresponding original/correct Luganda sentence
Corruptions are generated from clean Luganda text using linguistically informed corruption operations.
The objective is to train models to transform an erroneous… See the full description on the dataset page: https://huggingface.co/datasets/sagepond/gec.galician-gec-corpora
Galician GEC Corpora
Dataset Summary
Galician GEC Corpora is a collection of Galician grammatical and orthographic correction datasets. The repository groups several sources of sentence-level correction pairs in a single Hugging Face dataset repository, with each source exposed as a separate configuration.
Each instance contains an incorrect or non-standard Galician sentence and its corrected version. Some subsets also include error labels, correction tags, edit… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/galician-gec-corpora.UberText-GEC
UberText-GEC Dataset
Overview
UberText-GEC is a large corpus of social media texts scraped from Ukrainian Telegram
(UberText 2.0. Dataset; Paper, Chaplynskyi, 2023)
and automatically corrected using the approach TBU.
Structure
uber_text_gec.csv - main data.
language - language of text;
text - original text;
correction - corrected text;
uber_uk_annotations.csv - contains human annotations for 1500 samples.
text - original text;
correction - corrected… See the full description on the dataset page: https://huggingface.co/datasets/lang-uk/UberText-GEC.gec-targeted-corrections-esl
GEC Targeted Corrections — ESL
An LLM-generated grammatical error correction (GEC) dataset targeting the specific error patterns that ESL learners most commonly produce. Each example is a (src, tgt) pair where src contains a realistic grammatical error and tgt is the minimally corrected version: only what is necessary is changed.
Dataset Summary
Split
Examples
train
2,037
Schema
{
"src": "She gave me some advices about the… See the full description on the dataset page: https://huggingface.co/datasets/martinsr/gec-targeted-corrections-esl.novel-f3-v0.1
novel-f3-v0.1
The web novel dataset contains only the first 3 chapters.
Copyright Notice:
The dataset “novel-f3-v0.1” may include third-party copyrighted materials.Before any use, please verify ownership and obtain permission from the original copyright holders.The maintainers assume no responsibility for unauthorized use.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by… See the full description on the dataset page: https://huggingface.co/datasets/GeCDU/novel-f3-v0.1.ru_gec_v1
Dataset: RuGECv1 (Russian Grammatical Error Correction)
This dataset contains 707,261 parallel examples specifically curated for training Grammatical Error Correction (GEC) models for the Russian language. The dataset follows an instruction-tuning format, making it suitable for fine-tuning instruction-following language models.
Dataset Structure
Each example contains the following fields:
Field
Type
Description
input
string
Source text containing grammatical… See the full description on the dataset page: https://huggingface.co/datasets/alexanderpl/ru_gec_v1.ro_gec_dataset
Dataset Card for Ro-GEC (Synthetic)
Ro-GEC is a synthetic dataset for Grammatical Error Correction (GEC) in Romanian. It contains approximately 100,000 pairs of clean and corrupted sentences generated using a hybrid pipeline of deterministic regex rules and Large Language Models (LLMs).
Dataset Details
Dataset Description
This dataset was created to address the scarcity of resources for Romanian Grammatical Error Correction. It takes clean sentences from the… See the full description on the dataset page: https://huggingface.co/datasets/andreidiaconu/ro_gec_dataset.ru-gec-v1
Dataset: RuGECv1 (Russian Grammatical Error Correction)
This dataset contains 707,261 parallel examples specifically curated for training Grammatical Error Correction (GEC) models for the Russian language. The dataset follows an instruction-tuning format, making it suitable for fine-tuning instruction-following language models.
Dataset Structure
Each example contains the following fields:
Field
Type
Description
input
string
Source text containing grammatical… See the full description on the dataset page: https://huggingface.co/datasets/p1746-lingua/ru-gec-v1.Til-GEC
DEPRECATED
Этот датасет не следует использовать для обучения GEC. 39% строк () —
это параллельный корпус переводов kk↔ru и kk↔en с задачей и пустыми
. Он попал в сборку из — претрейн-микса,
а не GEC-набора. Обучение без фильтра учит модель переводить вместо того, чтобы исправлять.
Внутри также лежат строки публичного бенчмарка (), что делает любую оценку
без явной деконтаминации недействительной.
Преемники:
— только
коррекция, без переводов и без бенчмарка (2 860 104 строки… See the full description on the dataset page: https://huggingface.co/datasets/TilQazyna/Til-GEC.Til-GEC-v2
Til-GEC-v2 — Kazakh grammatical error correction
Sentence pairs for correcting grammar and spelling in Kazakh: an input sentence that
contains an error and a target sentence that fixes it. Each pair is labelled with the type of
error and the subject domain it came from.
About 14% of the pairs (401 661) are identity pairs — input and target are the same
sentence, labelled identity. They are there on purpose: a corrector that cannot leave a
correct sentence alone is useless, and a… See the full description on the dataset page: https://huggingface.co/datasets/TilQazyna/Til-GEC-v2.sozkz-gec-synthetic-gpt4o-v1
SozKZ GEC Synthetic GPT-4o v1
Синтетикалық қазақ тілі грамматикалық қателерді түзету (GEC) датасеті. GPT-4o арқылы генерацияланған.
Synthetic Kazakh Grammatical Error Correction dataset generated with OpenAI GPT-4o.
Overview
Parameter
Value
Total examples
9,599
Synthetic error pairs
7,204
Identity (clean) examples
2,395 (25%)
Source model
GPT-4o
Seed source
Kazakh Wikipedia
Language
Kazakh (kk)
License
MIT
Data Format… See the full description on the dataset page: https://huggingface.co/datasets/stukenov/sozkz-gec-synthetic-gpt4o-v1.
