CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01DipeshChaudhary /nepali-gector-style-token-level-tag-for-ged Nepali GEC (gector style) Token Tagging Dataset This is a processed version of the sumitaryal/nepali_grammatical_error_correction dataset, designed for training GEC-ToR-style sequence tagging models. This dataset has been processed with a robust, multi-pass, content-aware alignment algorithm to generate high-fidelity correction tags, including complex and adjacent SWAP operations. Total Examples: 16,260,992 Training: 13,008,711 Validation: 2,439,231 test: 813,050… See the full description on the dataset page: https://huggingface.co/datasets/DipeshChaudhary/nepali-gector-style-token-level-tag-for-ged.texttoken-classification10M<n<100M0 likes596 downloads11mo agoHugging Face02DipeshChaudhary /muril-nepali-gector-style-token-level-tag-for-ged Nepali GEC (gector style) Token Tagging Dataset This is a processed version of the sumitaryal/nepali_grammatical_error_correction dataset, designed for training GEC-ToR-style sequence tagging models. This dataset has been processed with a robust, multi-pass, content-aware alignment algorithm to generate high-fidelity correction tags, including complex and adjacent SWAP operations. Total Examples: 16,260,992 Training: 13,008,711 Validation: 2,439,231 test: 813,050… See the full description on the dataset page: https://huggingface.co/datasets/DipeshChaudhary/muril-nepali-gector-style-token-level-tag-for-ged.texttoken-classification10M<n<100M0 likes473 downloads11mo agoHugging Face03hafidikhsan /c4_200m-gec-train100k-test25k Dataset Card for "c4_200m-gec-train100k-test25k" More Information needed tabular100K<n<1M0 likes419 downloads3y agoHugging Face04ltg /ask-gec Norwegian grammatical error correction (ASK) This is the ASK-RAW dataset by Matias Jentoft (2023). Cite @mastersthesis{jentoft2023grammatical, title={Grammatical Error Correction with byte-level language models}, author={Jentoft, Matias}, year={2023}, school={University of Oslo}, url={https://www.duo.uio.no/handle/10852/103885} } text10K<n<100K4 likes314 downloads3y agoHugging Face05sagepond /gecgated Luganda Grammar Error Correction Dataset A synthetic dataset for training and evaluating Luganda Grammar Error Correction (GEC) models. Dataset Description Overview The dataset consists of pairs of: src: a corrupted Luganda sentence tgt: the corresponding original/correct Luganda sentence Corruptions are generated from clean Luganda text using linguistically informed corruption operations. The objective is to train models to transform an erroneous… See the full description on the dataset page: https://huggingface.co/datasets/sagepond/gec.texttext-generation1M<n<10M0 likes277 downloads22h agoHugging Face06lapa-llm /lang-uk-fiction-gec-dialogs Dataset Card for Ukrainian Fiction Grammatical Error Correction Dialogs Dataset Description Dataset Summary This dataset is a processed version of fiction part of the lang-uk UberText Corpus. The goal for this dataset is to provide grammatical error correction knowledge grounding. Languages Ukrainian (uk) Data Fields instruction: Text containing task description input: Processed text from the original text, including grammar errors output: Correct text task_type:… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/lang-uk-fiction-gec-dialogs.textquestion-answering10K<n<100K0 likes251 downloads11mo agoHugging Face07juancavallotti /english-gec-tatoebatext10K<n<100K1 likes189 downloads4y agoHugging Face08juancavallotti /multilingual-gec Dataset Card for Multilingual Grammar Error Correction Dataset Summary This dataset can be used to train a transformer model (we used T5) to correct grammar errors in simple sentences written in English, Spanish, French, or German. This dataset was developed as a component for the Squidigies platform. Supported Tasks and Leaderboards Grammar Error Correction: By appending the prefix fix grammar: to the prrompt. Language Detection: By appending the prefix:… See the full description on the dataset page: https://huggingface.co/datasets/juancavallotti/multilingual-gec.texttranslation100K<n<1M11 likes153 downloads4y agoHugging Face09mcemilg /GECTurk-generationHomepage: https://github.com/GGLAB-KU/gecturk/ text100K<n<1M4 likes148 downloads3y agoHugging Face10bumblebearhug /GEC25_C4text10M<n<100M0 likes141 downloads1y agoHugging Face11juancavallotti /english-gec-tatoeba-finetunetabular10K<n<100K3 likes128 downloads4y agoHugging Face12abhinavsarkar /C4-200M-1M-GEC-Determinertext1M<n<10M1 likes124 downloads1y agoHugging Face13lekhya-ai /gec Lekhya · Bangla GEC (synthetic) (incorrect → correct) Bangla sentence pairs with typed, span-level edits, for grammatical error correction. Built by Lekhya. ⚠️ Read this before training on it These errors are synthetic. They were produced by rule-based injection into clean newspaper prose. A model trained here learns to invert this noise function, which is not the same as learning to correct Bangla. Expect strong scores on a test set built by the same rules and… See the full description on the dataset page: https://huggingface.co/datasets/lekhya-ai/gec.text100K<n<1M0 likes112 downloads20d agoHugging Face14dreuxx26 /russian_gec 📝 Dataset Card for Russian Grammar Error-Correction (25 362 sentence pairs) A compact, high-quality corpus of Russian sentences with grammatical errors aligned to their human‑corrected counterparts. Ideal for training and benchmarking grammatical error‑correction (GEC) models, writing assistants, and translation post‑editing. ✨ Dataset Summary Metric Value Sentence pairs 25 362 Avg. tokens / sentence ≈ 12 File size ~5 MB (CSV, UTF‑8) Error types… See the full description on the dataset page: https://huggingface.co/datasets/dreuxx26/russian_gec.text10K<n<100K3 likes105 downloads1y agoHugging Face15maxmyn /c4ai-takehome-gec-traintext10K<n<100K0 likes90 downloads2y agoHugging Face16rishikeshgautam /newscorpus-for-gectext1M<n<10M0 likes84 downloads2y agoHugging Face17gechim /HealthQAtext10K<n<100K0 likes82 downloads2y agoHugging Face18Ro551 /WikiCorrupted_spanish_to_GEC-GED_Ltext100K<n<1M0 likes74 downloads10d agoHugging Face19timonziegenbein /fluency-pairs-gec-onlytext10K<n<100K0 likes72 downloads11mo agoHugging Face20huggingartists /100-gecsThis dataset is designed to generate lyrics with HuggingArtists.textn<1K0 likes71 downloads4y agoHugging Face21CZLC /cs_gec Introduction This dataset is extracted by postprocessing data from Náplava et al., 2019. Specificially, we extracted gramatically incorrect sentences, and their respective corrections. Then we convert task to binary detection of errorneous sentences. We downloaded the original dataset from LINDAT-Clarin repository. Citation @inproceedings{naplava-straka-2019-grammatical, title = "Grammatical Error Correction in Low-Resource Scenarios", author = "N{\'a}plava, Jakub… See the full description on the dataset page: https://huggingface.co/datasets/CZLC/cs_gec.text10K<n<100K0 likes69 downloads2y agoHugging Face22GGLab /GECTurktext100K<n<1M3 likes68 downloads3y agoHugging Face23Ro551 /COWSL2H_GEC_cleanedtext10K<n<100K0 likes59 downloads5mo agoHugging Face24proxectonos /galician-gec-corpora Galician GEC Corpora Dataset Summary Galician GEC Corpora is a collection of Galician grammatical and orthographic correction datasets. The repository groups several sources of sentence-level correction pairs in a single Hugging Face dataset repository, with each source exposed as a separate configuration. Each instance contains an incorrect or non-standard Galician sentence and its corrected version. Some subsets also include error labels, correction tags, edit… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/galician-gec-corpora.texttext-generation100K<n<1M0 likes48 downloads4mo agoHugging Face25upb-nlp /gec-ro-textstext1M<n<10M1 likes47 downloads2y agoHugging Face26Zlovoblachko /REALEC_GEC_dataset_ACL_testtext10K<n<100K0 likes45 downloads2mo agoHugging Face27JohnGorri /gec-coherence-coedit-synthtext10K<n<100K0 likes42 downloads3mo agoHugging Face28akufeldt /fr-gec-dataset Dataset Card for "fr-gec-dataset" More Information needed text10K<n<100K2 likes41 downloads3y agoHugging Face29Ro551 /WikiCorrupted_spanish_to_GEC-GED_largetext100K<n<1M0 likes41 downloads4mo agoHugging Face30asimokby /Turkish-OSCAR-GECtext1M<n<10M5 likes40 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.