datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
galician-gec-corpora
Galician GEC Corpora
Dataset Summary
Galician GEC Corpora is a collection of Galician grammatical and orthographic correction datasets. The repository groups several sources of sentence-level correction pairs in a single Hugging Face dataset repository, with each source exposed as a separate configuration.
Each instance contains an incorrect or non-standard Galician sentence and its corrected version. Some subsets also include error labels, correction tags, edit… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/galician-gec-corpora.ru_gec_v1
Dataset: RuGECv1 (Russian Grammatical Error Correction)
This dataset contains 707,261 parallel examples specifically curated for training Grammatical Error Correction (GEC) models for the Russian language. The dataset follows an instruction-tuning format, making it suitable for fine-tuning instruction-following language models.
Dataset Structure
Each example contains the following fields:
Field
Type
Description
input
string
Source text containing grammatical… See the full description on the dataset page: https://huggingface.co/datasets/alexanderpl/ru_gec_v1.ru-gec-v1
Dataset: RuGECv1 (Russian Grammatical Error Correction)
This dataset contains 707,261 parallel examples specifically curated for training Grammatical Error Correction (GEC) models for the Russian language. The dataset follows an instruction-tuning format, making it suitable for fine-tuning instruction-following language models.
Dataset Structure
Each example contains the following fields:
Field
Type
Description
input
string
Source text containing grammatical… See the full description on the dataset page: https://huggingface.co/datasets/p1746-lingua/ru-gec-v1.sozkz-gec-synthetic-gpt4o-v1
SozKZ GEC Synthetic GPT-4o v1
Синтетикалық қазақ тілі грамматикалық қателерді түзету (GEC) датасеті. GPT-4o арқылы генерацияланған.
Synthetic Kazakh Grammatical Error Correction dataset generated with OpenAI GPT-4o.
Overview
Parameter
Value
Total examples
9,599
Synthetic error pairs
7,204
Identity (clean) examples
2,395 (25%)
Source model
GPT-4o
Seed source
Kazakh Wikipedia
Language
Kazakh (kk)
License
MIT
Data Format… See the full description on the dataset page: https://huggingface.co/datasets/stukenov/sozkz-gec-synthetic-gpt4o-v1.
