CoolFace
25 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01grammarly /coedit Dataset Card for CoEdIT: Text Editing via Instruction Tuning Paper: CoEdIT: Text Editing by Task-Specific Instruction Tuning Authors: Vipul Raheja, Dhruv Kumar, Ryan Koo, Dongyeop Kang Project Repo: https://github.com/vipulraheja/coedit Dataset Summary This is the dataset that was used to train the CoEdIT text editing models. Full details of the dataset can be found in our paper. Dataset Structure The dataset is in JSON format.… See the full description on the dataset page: https://huggingface.co/datasets/grammarly/coedit.texttext-generation10K<n<100K101 likes2k downloads3y agoHugging Face02NuBerea /allen-greenough-grammargated allen-greenough-grammar Allen and Greenough's New Latin Grammar for Schools and Colleges (J.B. Greenough, G.L. Kittredge, A.A. Howard, Benjamin L. D'Ooge, eds.; Boston: Ginn & Company, 1903 — public domain), in the Dickinson College Commentaries (DCC) digital re-edition (https://dcc.dickinson.edu/grammar/latin/), edited by Meagan Ayer under Chris Francese's direction, 2013-2016 ("A New Allen and Greenough"). This is a PROSE-witness t0 source repo, the standard reference grammar… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/allen-greenough-grammar.texttext-generationn<1K0 likes425 downloads1mo agoHugging Face03NuBerea /smyth-grammargated smyth-grammar Herbert Weir Smyth, A Greek Grammar for Colleges (New York: American Book Company, 1920 — public domain). PROSE-witness t0 source repo for Greek morphology and syntax, the sibling of allen-greenough-grammar (Latin) and gesenius-kautzsch-grammar (Hebrew) in the distributional-grammar programme: one row per numbered Smyth paragraph (§1–§3048, complete, plus the 213 "D"-suffixed dialect paragraphs), Greek examples in polytonic Unicode, Smyth's own cross-references as… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/smyth-grammar.tabulartext-generation1K<n<10K0 likes424 downloads1mo agoHugging Face04NuBerea /gesenius-kautzsch-grammargated gesenius-kautzsch-grammar Gesenius' Hebrew Grammar, as edited and enlarged by E. Kautzsch, 28th German edition (1909), translated by A. E. Cowley — Second English Edition (Oxford: Clarendon Press, 1910; public domain), in the Wikisource collaborative transcription (https://en.wikisource.org/wiki/Gesenius%27_Hebrew_Grammar). This is a PROSE-witness t0 source repo, the standard reference grammar for Biblical Hebrew morphology and syntax cited throughout the NuBerea Hebrew corpus… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/gesenius-kautzsch-grammar.texttext-generation1K<n<10K0 likes417 downloads1mo agoHugging Face05grammarly /medit Dataset Card for mEdIT: Multilingual Text Editing via Instruction Tuning Paper: mEdIT: Multilingual Text Editing via Instruction Tuning Authors: Vipul Raheja, Dimitris Alikaniotis, Vivek Kulkarni, Bashar Alhafni, Dhruv Kumar Project Repo: https://github.com/vipulraheja/medit Dataset Summary This is the dataset that was used to train the mEdIT text editing models. Full details of the dataset can be found in our paper. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/grammarly/medit.texttext-generation100K<n<1M14 likes296 downloads2y agoHugging Face06grammarly /spivavtor Dataset Card for Spivavtor Paper: Spivavtor: An Instruction Tuned Ukrainian Text Editing Model Authors: Aman Saini, Artem Chernodub, Vipul Raheja, Vivek Kulkarni Dataset Summary This is the dataset used to train all Spivavtor models. It contains data for 4 tasks - Grammatical Error Correction (GEC), Simplification, Coherence and Paraphrasing. The specific details are as follows: Task Examples in Training data Examples in Validation… See the full description on the dataset page: https://huggingface.co/datasets/grammarly/spivavtor.texttext-generation10K<n<100K5 likes125 downloads2y agoHugging Face07hoisoserious /formal-grammar-llm-benchmark A formal grammar benchmark to study learning and memorization of large language models Usage Load a specific grammar (config) and split: from datasets import load_dataset ds = load_dataset( "<username>/formal-grammar-llm-benchmark", name="pcfg_cfg3b_eq_len_skewed_prob", ) # available splits per grammar: # train_sequences, test_sequences, non_grammatical_sequences, # non_grammatical_*_grammar_edit_*, non_grammatical_*_edit_distance_*… See the full description on the dataset page: https://huggingface.co/datasets/hoisoserious/formal-grammar-llm-benchmark.tabulartext-generation100K<n<1M0 likes58 downloads5mo agoHugging Face08alex73 /GrammarDB GrammarDB: Comprehensive Belarusian Grammatical Database This dataset contains a massive grammatical database of the Belarusian language, featuring approximately 4.5 million entries. It provides exhaustive information on wordforms, their paradigms, morphological features, and accents. Description GrammarDB is the primary open-source resource for the morphological analysis of the Belarusian language. It serves as the backbone for systems like LanguageTool, spellcheckers… See the full description on the dataset page: https://huggingface.co/datasets/alex73/GrammarDB.texttoken-classification1M<n<10M0 likes56 downloads6mo agoHugging Face09ScoutieAutoML /scoutieDataset_russian_language_grammar_and_rules_vectorized Description in English: A dataset collected from 30 Russian-language Telegram channels on the topic of learning the Russian language. This dataset contains grammar, syntax, spelling and punctuation rules. The dataset was collected and marked automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link. Dataset fields: taskId - task identifier in the Scouti service. text - main text. url -… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/scoutieDataset_russian_language_grammar_and_rules_vectorized.tabulartext-classification10K<n<100K2 likes48 downloads2y agoHugging Face10alsubari /arabic-grammar-errorstexttext-classification100K<n<1M0 likes42 downloads6mo agoHugging Face11nassimjp /Pashto-grammar-100 🇦🇫 Pashto Grammar 100 Pashto Grammar 100 is a compact, focused dataset created to help AI models learn and understand fundamental Pashto grammar, sentence structure, grammatical concepts, and correct linguistic usage. The dataset contains carefully selected Pashto grammar examples designed for language learning, grammatical analysis, instruction tuning, and evaluation of Pashto language models. It is intended as a small but high-quality resource for researchers and developers… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-grammar-100.texttext-generationn<1K0 likes37 downloads10d agoHugging Face12CreativeAlloyYT /French_Grammar_Explanations This dataset contains 1500+ French grammar explanations. It's the one I used to train my finetuned LLM called FrenchLlama-3.2-1B-Instruct. You can use this dataset for your own training purposes & find the aforementioned model on my HuggingFace profile. textquestion-answering1K<n<10K0 likes36 downloads1y agoHugging Face13Nam-toon-studio /Punjabi-Gurmukhi-Grammar-Correction-Corpus ੴ Punjabi (Gurmukhi) Grammatical Error Correction Corpus ☬ ਪੰਜਾਬੀ (ਗੁਰਮੁਖੀ) ਵਿਆਕਰਣ ਸ਼ੁੱਧੀ ਅਤੇ ਸੁਧਾਰ ਡਾਟਾਸੈੱਟ (v1.0) 👨‍💻 Research & Engineering Lead Creator & Architect: Gurpreet Singh Dhillon (Nam-toon Studio) GitHub Profile: github.com/gurpreetsingh5523-source Flagship Innovation: AMRIT Research OS (100% Locally-Run Autonomous Medical AI) 📖 Overview / ਸੰਖੇਪ The Punjabi (Gurmukhi) Grammatical Error Correction… See the full description on the dataset page: https://huggingface.co/datasets/Nam-toon-studio/Punjabi-Gurmukhi-Grammar-Correction-Corpus.texttext-generation1K<n<10K0 likes34 downloads8h agoHugging Face14benzlokzik /distill-grammarcheck qwen-distill-dataset Bilingual (RU/EN) text-normalization pairs: typos, grammar, punctuation, and casual chat style fixed while preserving meaning, tone, and slang. Each example is a chat pair: user: "Fix the text: <dirty>" → assistant: <clean>. CSV mirrors (train.csv / valid.csv) use incorrect,correct columns. Usage from datasets import load_dataset # Chat format (messages) — train / validation ds = load_dataset("benzlokzik/distill-grammarcheck") # Flat CSV… See the full description on the dataset page: https://huggingface.co/datasets/benzlokzik/distill-grammarcheck.texttext-generation10K<n<100K0 likes33 downloads3mo agoHugging Face15nassimjp /pashto-stf-grammar-pairs Pashto SFT Grammar Pairs Dataset Description Pashto SFT Grammar Pairs is a native-speaker-curated collection of Pashto question–answer pairs focused on Pashto grammar (ګرامر), covering topics such as noun gender, number, case (فاعلي، مفعولي، اضافي), adjective agreement, pronouns, verb conjugation, sentence structure (SOV word order), and enclitics/suffixes. The dataset is formatted in the {"messages": [...]} chat-template style used by modern SFT pipelines (TRL… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-stf-grammar-pairs.texttext-generationn<1K0 likes21 downloads1mo agoHugging Face16huytd189 /japanese-grammar-correction Japanese Grammar Correction Dataset Description This dataset is designed to train language models to identify and correct a wide range of grammatical errors and stylistic issues in Japanese text. The data consists of pairs of incorrect and correct sentences, along with metadata that classifies the type of error and provides additional context. The dataset was created by both manual curation from discussions in Japanese learning communities and synthetically generated… See the full description on the dataset page: https://huggingface.co/datasets/huytd189/japanese-grammar-correction.texttext-generation1K<n<10K0 likes20 downloads1y agoHugging Face17FableMogger9000 /byzantine-synthetic-grammar Byzantine Synthetic Interval-Grammar — Neume ↔ Western Pitch A correct-by-construction supervised fine-tuning dataset for transcribing between Byzantine (Chrysanthine) neume notation and Western staff pitches, in both directions, over a diatonic interval grammar with ascending leaps up to an octave and rhythmic note durations. This dataset trains the model FableMogger9000/byzantine-synthetic-grammar-lora, which reaches 96% exact-match / 98% melodic equivalence on the held-out… See the full description on the dataset page: https://huggingface.co/datasets/FableMogger9000/byzantine-synthetic-grammar.texttranslation10K<n<100K0 likes20 downloads2mo agoHugging Face18nassimjp /english_grammar_sft_system 📘 Dataset Overview english_grammar_sft_system is a curated supervised fine‑tuning (SFT) dataset derived from the Teravee/1000_english-grammar-dataset.Each sample is converted into a three‑message conversation format: system — global behavior instruction user — instruction + question merged assistant — high‑quality answer This structure is ideal for training instruction‑following LLMs that must respond clearly, concisely, and accurately to grammar‑related queries.… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/english_grammar_sft_system.texttext-generation10K<n<100K0 likes19 downloads1mo agoHugging Face19ScoutieAutoML /scoutieDataset_chinese_russian_dictionary_grammar_spelling_vectorized Description in English: A dataset collected from 30 Russian-language Telegram channels on the topic of learning Chinese, this dataset contains grammar, syntax, spelling and punctuation rules, as well as Chinese words with Russian translations. The dataset was collected and marked automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link. Dataset fields: taskId - task identifier in the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/scoutieDataset_chinese_russian_dictionary_grammar_spelling_vectorized.tabulartext-classification1K<n<10K0 likes16 downloads2y agoHugging Face20ambrosfitz /2k_grammar_correctionstexttext-generation1K<n<10K1 likes14 downloads1y agoHugging Face21gplsi /alia_gva_grammar 📘 ALIA_GVA_Grammar Dataset The ALIA_GVA_Grammar dataset is a resource designed for text generation. The dataset consists of textual documents formatted in Markdown (.md), each provided as structured JSONL entries. Each entry includes information about the text's language, format, text, source, and metadata. 🧾 Column Descriptions Field Type Description format string Indicates the text format. All entries use "md" (Markdown). language string Language of… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_gva_grammar.texttext-generation1K<n<10K0 likes14 downloads4mo agoHugging Face22nassimjp /pashto-grammar-tutor Pashto Grammar Tutor A high-quality Pashto grammar instruction dataset designed for language learning, linguistic research, and supervised fine-tuning (SFT) of AI language models. The dataset focuses on grammatical analysis, verb conjugation, sentence structure, and teacher-style explanations written in Pashto. Dataset Summary Pashto Grammar Tutor is a specialized educational dataset containing grammar-focused instruction-response pairs. Each example presents a… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-grammar-tutor.texttext-generation1K<n<10K0 likes14 downloads4mo agoHugging Face23Jnx03 /kanitakorn-deepseek-v43-grammar-polarity-boundary-micro Kanitakorn DeepSeek v43 Grammar Polarity Boundary Micro Synthetic minimal-pair SFT shard for Thai grammar exact counting, polarity/exception wording, and close d/e option-boundary calibration. Rows: 600 Pairs: 300, two rows each Labels: a=b=c=d=e=120 Final answer contract: คำตอบ: (x) Sources: generated original synthetic prompts only; no benchmark prompt/gold/sample rows used Train SHA256: 863ea31d6039b9986f9647ae65bf0d4c382893361a59a6d70795fbd2d2e4739e Use only as a tiny… See the full description on the dataset page: https://huggingface.co/datasets/Jnx03/kanitakorn-deepseek-v43-grammar-polarity-boundary-micro.texttext-generationn<1K0 likes12 downloads3mo agoHugging Face24LorthGyu /indonesian-grammar-correction Koreksi Tata Bahasa Indonesia ✍️ Kumpulan 126 pasangan kalimat (asli → koreksi) untuk grammar error correction bahasa Indonesia. Kenapa dataset ini ada? Grammar error correction (GEC) untuk bahasa Indonesia belum ada di HF — padahal model GEC global (271 downloads) jadi salah satu kategori paling dicari. Dataset ini isi gap itu: dari ejaan ("Dimana" → "Di mana"), ragam ("gw udah" → "saya sudah"), pleonasme ("para siswa-siswa" → "para siswa"), sampai huruf kapital.… See the full description on the dataset page: https://huggingface.co/datasets/LorthGyu/indonesian-grammar-correction.texttext-generationn<1K0 likes12 downloads2mo agoHugging Face25ScoutieAutoML /scoutieDataset_english_russian_dictionary_grammar_spelling_vectorized Description in English: A dataset collected from 30 Russian-language Telegram channels on the topic of learning English, this dataset contains grammar, syntax, spelling and punctuation rules, as well as English words with Russian translations. The dataset was collected and marked automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link. Dataset fields: taskId - task identifier in the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/scoutieDataset_english_russian_dictionary_grammar_spelling_vectorized.tabulartext-classification10K<n<100K0 likes10 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.