CoolFace
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sagepond /gecgated Luganda Grammar Error Correction Dataset A synthetic dataset for training and evaluating Luganda Grammar Error Correction (GEC) models. Dataset Description Overview The dataset consists of pairs of: src: a corrupted Luganda sentence tgt: the corresponding original/correct Luganda sentence Corruptions are generated from clean Luganda text using linguistically informed corruption operations. The objective is to train models to transform an erroneous… See the full description on the dataset page: https://huggingface.co/datasets/sagepond/gec.texttext-generation1M<n<10M0 likes287 downloads1d agoHugging Face02proxectonos /galician-gec-corpora Galician GEC Corpora Dataset Summary Galician GEC Corpora is a collection of Galician grammatical and orthographic correction datasets. The repository groups several sources of sentence-level correction pairs in a single Hugging Face dataset repository, with each source exposed as a separate configuration. Each instance contains an incorrect or non-standard Galician sentence and its corrected version. Some subsets also include error labels, correction tags, edit… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/galician-gec-corpora.texttext-generation100K<n<1M0 likes46 downloads4mo agoHugging Face03lang-uk /UberText-GEC UberText-GEC Dataset Overview UberText-GEC is a large corpus of social media texts scraped from Ukrainian Telegram (UberText 2.0. Dataset; Paper, Chaplynskyi, 2023) and automatically corrected using the approach TBU. Structure uber_text_gec.csv - main data. language - language of text; text - original text; correction - corrected text; uber_uk_annotations.csv - contains human annotations for 1500 samples. text - original text; correction - corrected… See the full description on the dataset page: https://huggingface.co/datasets/lang-uk/UberText-GEC.text-generation100K<n<1M1 likes43 downloads1y agoHugging Face04martinsr /gec-targeted-corrections-esl GEC Targeted Corrections — ESL An LLM-generated grammatical error correction (GEC) dataset targeting the specific error patterns that ESL learners most commonly produce. Each example is a (src, tgt) pair where src contains a realistic grammatical error and tgt is the minimally corrected version: only what is necessary is changed. Dataset Summary Split Examples train 2,037 Schema { "src": "She gave me some advices about the… See the full description on the dataset page: https://huggingface.co/datasets/martinsr/gec-targeted-corrections-esl.texttext-generation1K<n<10K0 likes37 downloads3mo agoHugging Face05GeCDU /novel-f3-v0.1 novel-f3-v0.1 The web novel dataset contains only the first 3 chapters. Copyright Notice: The dataset “novel-f3-v0.1” may include third-party copyrighted materials.Before any use, please verify ownership and obtain permission from the original copyright holders.The maintainers assume no responsibility for unauthorized use. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by… See the full description on the dataset page: https://huggingface.co/datasets/GeCDU/novel-f3-v0.1.texttext-generation100K<n<1M0 likes25 downloads11mo agoHugging Face06alexanderpl /ru_gec_v1 Dataset: RuGECv1 (Russian Grammatical Error Correction) This dataset contains 707,261 parallel examples specifically curated for training Grammatical Error Correction (GEC) models for the Russian language. The dataset follows an instruction-tuning format, making it suitable for fine-tuning instruction-following language models. Dataset Structure Each example contains the following fields: Field Type Description input string Source text containing grammatical… See the full description on the dataset page: https://huggingface.co/datasets/alexanderpl/ru_gec_v1.texttext-generation100K<n<1M1 likes22 downloads7mo agoHugging Face07andreidiaconu /ro_gec_dataset Dataset Card for Ro-GEC (Synthetic) Ro-GEC is a synthetic dataset for Grammatical Error Correction (GEC) in Romanian. It contains approximately 100,000 pairs of clean and corrupted sentences generated using a hybrid pipeline of deterministic regex rules and Large Language Models (LLMs). Dataset Details Dataset Description This dataset was created to address the scarcity of resources for Romanian Grammatical Error Correction. It takes clean sentences from the… See the full description on the dataset page: https://huggingface.co/datasets/andreidiaconu/ro_gec_dataset.texttext-classification100K<n<1M0 likes20 downloads9mo agoHugging Face08p1746-lingua /ru-gec-v1gated Dataset: RuGECv1 (Russian Grammatical Error Correction) This dataset contains 707,261 parallel examples specifically curated for training Grammatical Error Correction (GEC) models for the Russian language. The dataset follows an instruction-tuning format, making it suitable for fine-tuning instruction-following language models. Dataset Structure Each example contains the following fields: Field Type Description input string Source text containing grammatical… See the full description on the dataset page: https://huggingface.co/datasets/p1746-lingua/ru-gec-v1.texttext-generation100K<n<1M4 likes13 downloads8mo agoHugging Face09TilQazyna /Til-GECgated DEPRECATED Этот датасет не следует использовать для обучения GEC. 39% строк () — это параллельный корпус переводов kk↔ru и kk↔en с задачей и пустыми . Он попал в сборку из — претрейн-микса, а не GEC-набора. Обучение без фильтра учит модель переводить вместо того, чтобы исправлять. Внутри также лежат строки публичного бенчмарка (), что делает любую оценку без явной деконтаминации недействительной. Преемники: — только коррекция, без переводов и без бенчмарка (2 860 104 строки… See the full description on the dataset page: https://huggingface.co/datasets/TilQazyna/Til-GEC.texttext-generation1M<n<10M0 likes8 downloads2mo agoHugging Face10TilQazyna /Til-GEC-v2gated Til-GEC-v2 — Kazakh grammatical error correction Sentence pairs for correcting grammar and spelling in Kazakh: an input sentence that contains an error and a target sentence that fixes it. Each pair is labelled with the type of error and the subject domain it came from. About 14% of the pairs (401 661) are identity pairs — input and target are the same sentence, labelled identity. They are there on purpose: a corrector that cannot leave a correct sentence alone is useless, and a… See the full description on the dataset page: https://huggingface.co/datasets/TilQazyna/Til-GEC-v2.texttext-generation1M<n<10M0 likes8 downloads2mo agoHugging Face11stukenov /sozkz-gec-synthetic-gpt4o-v1gated SozKZ GEC Synthetic GPT-4o v1 Синтетикалық қазақ тілі грамматикалық қателерді түзету (GEC) датасеті. GPT-4o арқылы генерацияланған. Synthetic Kazakh Grammatical Error Correction dataset generated with OpenAI GPT-4o. Overview Parameter Value Total examples 9,599 Synthetic error pairs 7,204 Identity (clean) examples 2,395 (25%) Source model GPT-4o Seed source Kazakh Wikipedia Language Kazakh (kk) License MIT Data Format… See the full description on the dataset page: https://huggingface.co/datasets/stukenov/sozkz-gec-synthetic-gpt4o-v1.texttext-generation1K<n<10K0 likes6 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.