CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mideind /gec-test-setTest data for Icelandic spell and grammar checking, created as part of the Icelandic Language Technology Programme. The test data is divided into three different formats, type 1, 2 and 3. For every original file corrected, three files are included in the test data when possible: _original, _corrected and _metadata. The original and metadata files are always .txt files, but the format of the corrected file differs between types. Texts corrected are from the News2 subcorpus of the Icelandic… See the full description on the dataset page: https://huggingface.co/datasets/mideind/gec-test-set.0 likes1.7k downloads2y agoHugging Face02DipeshChaudhary /nepali-gector-style-token-level-tag-for-ged Nepali GEC (gector style) Token Tagging Dataset This is a processed version of the sumitaryal/nepali_grammatical_error_correction dataset, designed for training GEC-ToR-style sequence tagging models. This dataset has been processed with a robust, multi-pass, content-aware alignment algorithm to generate high-fidelity correction tags, including complex and adjacent SWAP operations. Total Examples: 16,260,992 Training: 13,008,711 Validation: 2,439,231 test: 813,050… See the full description on the dataset page: https://huggingface.co/datasets/DipeshChaudhary/nepali-gector-style-token-level-tag-for-ged.texttoken-classification10M<n<100M0 likes596 downloads11mo agoHugging Face03DipeshChaudhary /muril-nepali-gector-style-token-level-tag-for-ged Nepali GEC (gector style) Token Tagging Dataset This is a processed version of the sumitaryal/nepali_grammatical_error_correction dataset, designed for training GEC-ToR-style sequence tagging models. This dataset has been processed with a robust, multi-pass, content-aware alignment algorithm to generate high-fidelity correction tags, including complex and adjacent SWAP operations. Total Examples: 16,260,992 Training: 13,008,711 Validation: 2,439,231 test: 813,050… See the full description on the dataset page: https://huggingface.co/datasets/DipeshChaudhary/muril-nepali-gector-style-token-level-tag-for-ged.texttoken-classification10M<n<100M0 likes473 downloads11mo agoHugging Face04jerpelhan /geco2-assets0 likes436 downloads1mo agoHugging Face05hafidikhsan /c4_200m-gec-train100k-test25k Dataset Card for "c4_200m-gec-train100k-test25k" More Information needed tabular100K<n<1M0 likes419 downloads3y agoHugging Face06ltg /ask-gec Norwegian grammatical error correction (ASK) This is the ASK-RAW dataset by Matias Jentoft (2023). Cite @mastersthesis{jentoft2023grammatical, title={Grammatical Error Correction with byte-level language models}, author={Jentoft, Matias}, year={2023}, school={University of Oslo}, url={https://www.duo.uio.no/handle/10852/103885} } text10K<n<100K4 likes314 downloads3y agoHugging Face07sagepond /gecgated Luganda Grammar Error Correction Dataset A synthetic dataset for training and evaluating Luganda Grammar Error Correction (GEC) models. Dataset Description Overview The dataset consists of pairs of: src: a corrupted Luganda sentence tgt: the corresponding original/correct Luganda sentence Corruptions are generated from clean Luganda text using linguistically informed corruption operations. The objective is to train models to transform an erroneous… See the full description on the dataset page: https://huggingface.co/datasets/sagepond/gec.texttext-generation1M<n<10M0 likes277 downloads10h agoHugging Face08lapa-llm /lang-uk-fiction-gec-dialogs Dataset Card for Ukrainian Fiction Grammatical Error Correction Dialogs Dataset Description Dataset Summary This dataset is a processed version of fiction part of the lang-uk UberText Corpus. The goal for this dataset is to provide grammatical error correction knowledge grounding. Languages Ukrainian (uk) Data Fields instruction: Text containing task description input: Processed text from the original text, including grammar errors output: Correct text task_type:… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/lang-uk-fiction-gec-dialogs.textquestion-answering10K<n<100K0 likes251 downloads11mo agoHugging Face09juancavallotti /english-gec-tatoebatext10K<n<100K1 likes189 downloads4y agoHugging Face10juancavallotti /multilingual-gec Dataset Card for Multilingual Grammar Error Correction Dataset Summary This dataset can be used to train a transformer model (we used T5) to correct grammar errors in simple sentences written in English, Spanish, French, or German. This dataset was developed as a component for the Squidigies platform. Supported Tasks and Leaderboards Grammar Error Correction: By appending the prefix fix grammar: to the prrompt. Language Detection: By appending the prefix:… See the full description on the dataset page: https://huggingface.co/datasets/juancavallotti/multilingual-gec.texttranslation100K<n<1M11 likes153 downloads4y agoHugging Face11mcemilg /GECTurk-generationHomepage: https://github.com/GGLAB-KU/gecturk/ text100K<n<1M4 likes148 downloads3y agoHugging Face12bumblebearhug /GEC25_C4text10M<n<100M0 likes141 downloads1y agoHugging Face13juancavallotti /english-gec-tatoeba-finetunetabular10K<n<100K3 likes128 downloads4y agoHugging Face14bstds /geco_data_generator0 likes127 downloads4y agoHugging Face15abhinavsarkar /C4-200M-1M-GEC-Determinertext1M<n<10M1 likes124 downloads1y agoHugging Face16open-llm-leaderboard-old /details_NeuralNovel__Gecko-7B-v0.1 Dataset Card for Evaluation run of NeuralNovel/Gecko-7B-v0.1 Dataset automatically created during the evaluation run of model NeuralNovel/Gecko-7B-v0.1 on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_NeuralNovel__Gecko-7B-v0.1.0 likes112 downloads3y agoHugging Face17lekhya-ai /gec Lekhya · Bangla GEC (synthetic) (incorrect → correct) Bangla sentence pairs with typed, span-level edits, for grammatical error correction. Built by Lekhya. ⚠️ Read this before training on it These errors are synthetic. They were produced by rule-based injection into clean newspaper prose. A model trained here learns to invert this noise function, which is not the same as learning to correct Bangla. Expect strong scores on a test set built by the same rules and… See the full description on the dataset page: https://huggingface.co/datasets/lekhya-ai/gec.text100K<n<1M0 likes112 downloads19d agoHugging Face18dreuxx26 /russian_gec 📝 Dataset Card for Russian Grammar Error-Correction (25 362 sentence pairs) A compact, high-quality corpus of Russian sentences with grammatical errors aligned to their human‑corrected counterparts. Ideal for training and benchmarking grammatical error‑correction (GEC) models, writing assistants, and translation post‑editing. ✨ Dataset Summary Metric Value Sentence pairs 25 362 Avg. tokens / sentence ≈ 12 File size ~5 MB (CSV, UTF‑8) Error types… See the full description on the dataset page: https://huggingface.co/datasets/dreuxx26/russian_gec.text10K<n<100K3 likes105 downloads1y agoHugging Face19maxmyn /c4ai-takehome-gec-traintext10K<n<100K0 likes90 downloads2y agoHugging Face20rishikeshgautam /newscorpus-for-gectext1M<n<10M0 likes84 downloads2y agoHugging Face21gechim /HealthQAtext10K<n<100K0 likes82 downloads2y agoHugging Face22Ro551 /WikiCorrupted_spanish_to_GEC-GED_Ltext100K<n<1M0 likes74 downloads9d agoHugging Face23timonziegenbein /fluency-pairs-gec-onlytext10K<n<100K0 likes72 downloads11mo agoHugging Face24huggingartists /100-gecsThis dataset is designed to generate lyrics with HuggingArtists.textn<1K0 likes71 downloads4y agoHugging Face25CZLC /cs_gec Introduction This dataset is extracted by postprocessing data from Náplava et al., 2019. Specificially, we extracted gramatically incorrect sentences, and their respective corrections. Then we convert task to binary detection of errorneous sentences. We downloaded the original dataset from LINDAT-Clarin repository. Citation @inproceedings{naplava-straka-2019-grammatical, title = "Grammatical Error Correction in Low-Resource Scenarios", author = "N{\'a}plava, Jakub… See the full description on the dataset page: https://huggingface.co/datasets/CZLC/cs_gec.text10K<n<100K0 likes69 downloads2y agoHugging Face26GGLab /GECTurktext100K<n<1M3 likes68 downloads3y agoHugging Face27Ro551 /COWSL2H_GEC_cleanedtext10K<n<100K0 likes59 downloads5mo agoHugging Face28proxectonos /galician-gec-corpora Galician GEC Corpora Dataset Summary Galician GEC Corpora is a collection of Galician grammatical and orthographic correction datasets. The repository groups several sources of sentence-level correction pairs in a single Hugging Face dataset repository, with each source exposed as a separate configuration. Each instance contains an incorrect or non-standard Galician sentence and its corrected version. Some subsets also include error labels, correction tags, edit… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/galician-gec-corpora.texttext-generation100K<n<1M0 likes48 downloads4mo agoHugging Face29upb-nlp /gec-ro-textstext1M<n<10M1 likes47 downloads2y agoHugging Face30Zlovoblachko /REALEC_GEC_dataset_ACL_testtext10K<n<100K0 likes45 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.