CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hafidikhsan /c4_200m-gec-train100k-test25k Dataset Card for "c4_200m-gec-train100k-test25k" More Information needed tabular100K<n<1M0 likes420 downloads3y agoHugging Face02sagepond /gecgated Luganda Grammar Error Correction Dataset A synthetic dataset for training and evaluating Luganda Grammar Error Correction (GEC) models. Dataset Description Overview The dataset consists of pairs of: src: a corrupted Luganda sentence tgt: the corresponding original/correct Luganda sentence Corruptions are generated from clean Luganda text using linguistically informed corruption operations. The objective is to train models to transform an erroneous… See the full description on the dataset page: https://huggingface.co/datasets/sagepond/gec.texttext-generation1M<n<10M0 likes332 downloads2h agoHugging Face03lapa-llm /lang-uk-fiction-gec-dialogs Dataset Card for Ukrainian Fiction Grammatical Error Correction Dialogs Dataset Description Dataset Summary This dataset is a processed version of fiction part of the lang-uk UberText Corpus. The goal for this dataset is to provide grammatical error correction knowledge grounding. Languages Ukrainian (uk) Data Fields instruction: Text containing task description input: Processed text from the original text, including grammar errors output: Correct text task_type:… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/lang-uk-fiction-gec-dialogs.textquestion-answering10K<n<100K0 likes272 downloads11mo agoHugging Face04juancavallotti /english-gec-tatoebatext10K<n<100K1 likes189 downloads4y agoHugging Face05juancavallotti /multilingual-gec Dataset Card for Multilingual Grammar Error Correction Dataset Summary This dataset can be used to train a transformer model (we used T5) to correct grammar errors in simple sentences written in English, Spanish, French, or German. This dataset was developed as a component for the Squidigies platform. Supported Tasks and Leaderboards Grammar Error Correction: By appending the prefix fix grammar: to the prrompt. Language Detection: By appending the prefix:… See the full description on the dataset page: https://huggingface.co/datasets/juancavallotti/multilingual-gec.texttranslation100K<n<1M11 likes177 downloads4y agoHugging Face06bumblebearhug /GEC25_C4text10M<n<100M0 likes167 downloads1y agoHugging Face07juancavallotti /english-gec-tatoeba-finetunetabular10K<n<100K3 likes126 downloads4y agoHugging Face08abhinavsarkar /C4-200M-1M-GEC-Determinertext1M<n<10M1 likes125 downloads1y agoHugging Face09lekhya-ai /gec Lekhya · Bangla GEC (synthetic) (incorrect → correct) Bangla sentence pairs with typed, span-level edits, for grammatical error correction. Built by Lekhya. ⚠️ Read this before training on it These errors are synthetic. They were produced by rule-based injection into clean newspaper prose. A model trained here learns to invert this noise function, which is not the same as learning to correct Bangla. Expect strong scores on a test set built by the same rules and… See the full description on the dataset page: https://huggingface.co/datasets/lekhya-ai/gec.text100K<n<1M0 likes115 downloads23d agoHugging Face10maxmyn /c4ai-takehome-gec-traintext10K<n<100K0 likes93 downloads2y agoHugging Face11rishikeshgautam /newscorpus-for-gectext1M<n<10M0 likes91 downloads2y agoHugging Face12Ro551 /WikiCorrupted_spanish_to_GEC-GED_Ltext100K<n<1M0 likes82 downloads13d agoHugging Face13gechim /HealthQAtext10K<n<100K0 likes71 downloads2y agoHugging Face14timonziegenbein /fluency-pairs-gec-onlytext10K<n<100K0 likes68 downloads1y agoHugging Face15GGLab /GECTurktext100K<n<1M3 likes59 downloads3y agoHugging Face16Ro551 /COWSL2H_GEC_cleanedtext10K<n<100K0 likes55 downloads5mo agoHugging Face17Zlovoblachko /REALEC_GEC_dataset_ACL_testtext10K<n<100K0 likes52 downloads2mo agoHugging Face18JohnGorri /gec-coherence-coedit-synthtext10K<n<100K0 likes48 downloads3mo agoHugging Face19Ro551 /WikiCorrupted-spanish_to_GEC-GEDtext100K<n<1M1 likes44 downloads4mo agoHugging Face20Zlovoblachko /REALEC_GEC_dataset_ACLtext100K<n<1M0 likes43 downloads2mo agoHugging Face21abhinavsarkar /C4-200M-1.55M-GEC-Determinertext1M<n<10M0 likes42 downloads1y agoHugging Face22sirsam01 /codeit_gectabular10K<n<100K0 likes41 downloads2y agoHugging Face23akufeldt /fr-gec-dataset Dataset Card for "fr-gec-dataset" More Information needed text10K<n<100K2 likes40 downloads3y agoHugging Face24martinsr /gec-targeted-corrections-esl GEC Targeted Corrections — ESL An LLM-generated grammatical error correction (GEC) dataset targeting the specific error patterns that ESL learners most commonly produce. Each example is a (src, tgt) pair where src contains a realistic grammatical error and tgt is the minimally corrected version: only what is necessary is changed. Dataset Summary Split Examples train 2,037 Schema { "src": "She gave me some advices about the… See the full description on the dataset page: https://huggingface.co/datasets/martinsr/gec-targeted-corrections-esl.texttext-generation1K<n<10K0 likes37 downloads3mo agoHugging Face25kenza-ily /c4ai_csp-gec_datasetstext10K<n<100K0 likes36 downloads2y agoHugging Face26timonziegenbein /fluency-pairs-gec-onklytext10K<n<100K0 likes36 downloads1y agoHugging Face27Zlovoblachko /REALEC_GEC_datasettext10K<n<100K0 likes35 downloads1y agoHugging Face28judywq /gec-datasettext10K<n<100K0 likes33 downloads2y agoHugging Face29rishikeshgautam /newscorpus-for-gec-smalltext100K<n<1M0 likes31 downloads2y agoHugging Face30ammarnasr /SmolLM-135M-GEC-preference-datasettext10K<n<100K0 likes30 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.