CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ltg /ask-gec Norwegian grammatical error correction (ASK) This is the ASK-RAW dataset by Matias Jentoft (2023). Cite @mastersthesis{jentoft2023grammatical, title={Grammatical Error Correction with byte-level language models}, author={Jentoft, Matias}, year={2023}, school={University of Oslo}, url={https://www.duo.uio.no/handle/10852/103885} } text10K<n<100K4 likes314 downloads3y agoHugging Face02CZLC /cs_gec Introduction This dataset is extracted by postprocessing data from Náplava et al., 2019. Specificially, we extracted gramatically incorrect sentences, and their respective corrections. Then we convert task to binary detection of errorneous sentences. We downloaded the original dataset from LINDAT-Clarin repository. Citation @inproceedings{naplava-straka-2019-grammatical, title = "Grammatical Error Correction in Low-Resource Scenarios", author = "N{\'a}plava, Jakub… See the full description on the dataset page: https://huggingface.co/datasets/CZLC/cs_gec.text10K<n<100K0 likes69 downloads2y agoHugging Face03proxectonos /galician-gec-corpora Galician GEC Corpora Dataset Summary Galician GEC Corpora is a collection of Galician grammatical and orthographic correction datasets. The repository groups several sources of sentence-level correction pairs in a single Hugging Face dataset repository, with each source exposed as a separate configuration. Each instance contains an incorrect or non-standard Galician sentence and its corrected version. Some subsets also include error labels, correction tags, edit… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/galician-gec-corpora.texttext-generation100K<n<1M0 likes48 downloads4mo agoHugging Face04osyvokon /ua_gec_instruction_tuning UA-GEC instruction tuning This dataset contains prompts and expected outputs for the grammatical error correction task in the Ukrainian language. It is based on the CC-BY-4.0-licensed UA-GEC dataset. The license of the original data is CC-BY-4.0. This dataset contains 1,700 examples of fixing errors in long documents, and ~28,000 sentence-level examples. The instructions ask to correct errors in the text. Sometimes the model outputs the corrected text as is. At other times, it… See the full description on the dataset page: https://huggingface.co/datasets/osyvokon/ua_gec_instruction_tuning.text10K<n<100K2 likes36 downloads3y agoHugging Face05alexanderpl /ru-gec-v1text100K<n<1M0 likes23 downloads8mo agoHugging Face06alexanderpl /ru_gec_v1 Dataset: RuGECv1 (Russian Grammatical Error Correction) This dataset contains 707,261 parallel examples specifically curated for training Grammatical Error Correction (GEC) models for the Russian language. The dataset follows an instruction-tuning format, making it suitable for fine-tuning instruction-following language models. Dataset Structure Each example contains the following fields: Field Type Description input string Source text containing grammatical… See the full description on the dataset page: https://huggingface.co/datasets/alexanderpl/ru_gec_v1.texttext-generation100K<n<1M1 likes21 downloads7mo agoHugging Face07tartuNLP /est_gec_ged_gee_2025text100K<n<1M0 likes20 downloads1y agoHugging Face08mideind /icelandic-sentences-gectextn<1K0 likes18 downloads3y agoHugging Face09p1746-lingua /ru-gec-v1gated Dataset: RuGECv1 (Russian Grammatical Error Correction) This dataset contains 707,261 parallel examples specifically curated for training Grammatical Error Correction (GEC) models for the Russian language. The dataset follows an instruction-tuning format, making it suitable for fine-tuning instruction-following language models. Dataset Structure Each example contains the following fields: Field Type Description input string Source text containing grammatical… See the full description on the dataset page: https://huggingface.co/datasets/p1746-lingua/ru-gec-v1.texttext-generation100K<n<1M4 likes13 downloads8mo agoHugging Face10giannor /gec_dala_tv2r_ittext100K<n<1M0 likes12 downloads3mo agoHugging Face11stukenov /sozkz-gec-synthetic-gpt4o-v1gated SozKZ GEC Synthetic GPT-4o v1 Синтетикалық қазақ тілі грамматикалық қателерді түзету (GEC) датасеті. GPT-4o арқылы генерацияланған. Synthetic Kazakh Grammatical Error Correction dataset generated with OpenAI GPT-4o. Overview Parameter Value Total examples 9,599 Synthetic error pairs 7,204 Identity (clean) examples 2,395 (25%) Source model GPT-4o Seed source Kazakh Wikipedia Language Kazakh (kk) License MIT Data Format… See the full description on the dataset page: https://huggingface.co/datasets/stukenov/sozkz-gec-synthetic-gpt4o-v1.texttext-generation1K<n<10K0 likes5 downloads5mo agoHugging Face12Siva-Nandu-S /Group_project_gecbh_2025_trvtext1K<n<10K0 likes2 downloads2y agoHugging Face13Gecko51 /Gecko_Mindtextn<1K0 likes2 downloads11mo agoHugging Face14iwinther /en_nl_data_gecotext10K<n<100K0 likes1 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.