datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
coedit
Dataset Card for CoEdIT: Text Editing via Instruction Tuning
Paper: CoEdIT: Text Editing by Task-Specific Instruction Tuning
Authors: Vipul Raheja, Dhruv Kumar, Ryan Koo, Dongyeop Kang
Project Repo: https://github.com/vipulraheja/coedit
Dataset Summary
This is the dataset that was used to train the CoEdIT text editing models. Full details of the dataset can be found in our paper.
Dataset Structure
The dataset is in JSON format.… See the full description on the dataset page: https://huggingface.co/datasets/grammarly/coedit.allen-greenough-grammar
allen-greenough-grammar
Allen and Greenough's New Latin Grammar for Schools and Colleges (J.B.
Greenough, G.L. Kittredge, A.A. Howard, Benjamin L. D'Ooge, eds.; Boston:
Ginn & Company, 1903 — public domain), in the Dickinson College
Commentaries (DCC) digital re-edition
(https://dcc.dickinson.edu/grammar/latin/), edited by Meagan Ayer under
Chris Francese's direction, 2013-2016 ("A New Allen and Greenough"). This is
a PROSE-witness t0 source repo, the standard reference grammar… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/allen-greenough-grammar.smyth-grammar
smyth-grammar
Herbert Weir Smyth, A Greek Grammar for Colleges (New York: American Book
Company, 1920 — public domain). PROSE-witness t0 source repo for Greek
morphology and syntax, the sibling of allen-greenough-grammar (Latin) and
gesenius-kautzsch-grammar (Hebrew) in the distributional-grammar programme:
one row per numbered Smyth paragraph (§1–§3048, complete, plus the 213
"D"-suffixed dialect paragraphs), Greek examples in polytonic Unicode, Smyth's
own cross-references as… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/smyth-grammar.gesenius-kautzsch-grammar
gesenius-kautzsch-grammar
Gesenius' Hebrew Grammar, as edited and enlarged by E. Kautzsch, 28th
German edition (1909), translated by A. E. Cowley — Second English
Edition (Oxford: Clarendon Press, 1910; public domain), in the
Wikisource collaborative transcription
(https://en.wikisource.org/wiki/Gesenius%27_Hebrew_Grammar). This is a
PROSE-witness t0 source repo, the standard reference grammar for Biblical
Hebrew morphology and syntax cited throughout the NuBerea Hebrew corpus… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/gesenius-kautzsch-grammar.medit
Dataset Card for mEdIT: Multilingual Text Editing via Instruction Tuning
Paper: mEdIT: Multilingual Text Editing via Instruction Tuning
Authors: Vipul Raheja, Dimitris Alikaniotis, Vivek Kulkarni, Bashar Alhafni, Dhruv Kumar
Project Repo: https://github.com/vipulraheja/medit
Dataset Summary
This is the dataset that was used to train the mEdIT text editing models. Full details of the dataset can be found in our paper.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/grammarly/medit.spivavtor
Dataset Card for Spivavtor
Paper: Spivavtor: An Instruction Tuned Ukrainian Text Editing Model
Authors: Aman Saini, Artem Chernodub, Vipul Raheja, Vivek Kulkarni
Dataset Summary
This is the dataset used to train all Spivavtor models. It contains data for 4 tasks - Grammatical Error Correction (GEC), Simplification, Coherence and Paraphrasing.
The specific details are as follows:
Task
Examples in Training data
Examples in Validation… See the full description on the dataset page: https://huggingface.co/datasets/grammarly/spivavtor.formal-grammar-llm-benchmark
A formal grammar benchmark to study learning and memorization of large language models
Usage
Load a specific grammar (config) and split:
from datasets import load_dataset
ds = load_dataset(
"<username>/formal-grammar-llm-benchmark",
name="pcfg_cfg3b_eq_len_skewed_prob",
)
# available splits per grammar:
# train_sequences, test_sequences, non_grammatical_sequences,
# non_grammatical_*_grammar_edit_*, non_grammatical_*_edit_distance_*… See the full description on the dataset page: https://huggingface.co/datasets/hoisoserious/formal-grammar-llm-benchmark.GrammarDB
GrammarDB: Comprehensive Belarusian Grammatical Database
This dataset contains a massive grammatical database of the Belarusian language, featuring approximately 4.5 million entries. It provides exhaustive information on wordforms, their paradigms, morphological features, and accents.
Description
GrammarDB is the primary open-source resource for the morphological analysis of the Belarusian language. It serves as the backbone for systems like LanguageTool, spellcheckers… See the full description on the dataset page: https://huggingface.co/datasets/alex73/GrammarDB.scoutieDataset_russian_language_grammar_and_rules_vectorized
Description in English:
A dataset collected from 30 Russian-language Telegram channels on the topic of learning the Russian language. This dataset contains grammar, syntax, spelling and punctuation rules.
The dataset was collected and marked automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link.
Dataset fields:
taskId - task identifier in the Scouti service. text - main text. url -… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/scoutieDataset_russian_language_grammar_and_rules_vectorized.arabic-grammar-errorsPashto-grammar-100
🇦🇫 Pashto Grammar 100
Pashto Grammar 100 is a compact, focused dataset created to help AI models learn and understand fundamental Pashto grammar, sentence structure, grammatical concepts, and correct linguistic usage.
The dataset contains carefully selected Pashto grammar examples designed for language learning, grammatical analysis, instruction tuning, and evaluation of Pashto language models.
It is intended as a small but high-quality resource for researchers and developers… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-grammar-100.French_Grammar_Explanations
This dataset contains 1500+ French grammar explanations. It's the one I used to train my finetuned LLM called FrenchLlama-3.2-1B-Instruct.
You can use this dataset for your own training purposes & find the aforementioned model on my HuggingFace profile.
Punjabi-Gurmukhi-Grammar-Correction-Corpus
ੴ Punjabi (Gurmukhi) Grammatical Error Correction Corpus
☬ ਪੰਜਾਬੀ (ਗੁਰਮੁਖੀ) ਵਿਆਕਰਣ ਸ਼ੁੱਧੀ ਅਤੇ ਸੁਧਾਰ ਡਾਟਾਸੈੱਟ (v1.0)
👨💻 Research & Engineering Lead
Creator & Architect: Gurpreet Singh Dhillon (Nam-toon Studio)
GitHub Profile: github.com/gurpreetsingh5523-source
Flagship Innovation: AMRIT Research OS (100% Locally-Run Autonomous Medical AI)
📖 Overview / ਸੰਖੇਪ
The Punjabi (Gurmukhi) Grammatical Error Correction… See the full description on the dataset page: https://huggingface.co/datasets/Nam-toon-studio/Punjabi-Gurmukhi-Grammar-Correction-Corpus.distill-grammarcheck
qwen-distill-dataset
Bilingual (RU/EN) text-normalization pairs: typos, grammar, punctuation, and
casual chat style fixed while preserving meaning, tone, and slang.
Each example is a chat pair: user: "Fix the text: <dirty>" → assistant: <clean>.
CSV mirrors (train.csv / valid.csv) use incorrect,correct columns.
Usage
from datasets import load_dataset
# Chat format (messages) — train / validation
ds = load_dataset("benzlokzik/distill-grammarcheck")
# Flat CSV… See the full description on the dataset page: https://huggingface.co/datasets/benzlokzik/distill-grammarcheck.pashto-stf-grammar-pairs
Pashto SFT Grammar Pairs
Dataset Description
Pashto SFT Grammar Pairs is a native-speaker-curated collection of Pashto question–answer pairs focused on Pashto grammar (ګرامر), covering topics such as noun gender, number, case (فاعلي، مفعولي، اضافي), adjective agreement, pronouns, verb conjugation, sentence structure (SOV word order), and enclitics/suffixes.
The dataset is formatted in the {"messages": [...]} chat-template style used by modern SFT pipelines (TRL… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-stf-grammar-pairs.japanese-grammar-correction
Japanese Grammar Correction Dataset
Description
This dataset is designed to train language models to identify and correct a wide range of grammatical errors and stylistic issues in Japanese text.
The data consists of pairs of incorrect and correct sentences, along with metadata that classifies the type of error and provides additional context.
The dataset was created by both manual curation from discussions in Japanese learning communities and synthetically generated… See the full description on the dataset page: https://huggingface.co/datasets/huytd189/japanese-grammar-correction.byzantine-synthetic-grammar
Byzantine Synthetic Interval-Grammar — Neume ↔ Western Pitch
A correct-by-construction supervised fine-tuning dataset for transcribing between Byzantine
(Chrysanthine) neume notation and Western staff pitches, in both directions, over a diatonic
interval grammar with ascending leaps up to an octave and rhythmic note durations.
This dataset trains the model
FableMogger9000/byzantine-synthetic-grammar-lora,
which reaches 96% exact-match / 98% melodic equivalence on the held-out… See the full description on the dataset page: https://huggingface.co/datasets/FableMogger9000/byzantine-synthetic-grammar.english_grammar_sft_system
📘 Dataset Overview
english_grammar_sft_system is a curated supervised fine‑tuning (SFT) dataset derived from the Teravee/1000_english-grammar-dataset.Each sample is converted into a three‑message conversation format:
system — global behavior instruction
user — instruction + question merged
assistant — high‑quality answer
This structure is ideal for training instruction‑following LLMs that must respond clearly, concisely, and accurately to grammar‑related queries.… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/english_grammar_sft_system.scoutieDataset_chinese_russian_dictionary_grammar_spelling_vectorized
Description in English:
A dataset collected from 30 Russian-language Telegram channels on the topic of learning Chinese, this dataset contains grammar, syntax, spelling and punctuation rules, as well as Chinese words with Russian translations.
The dataset was collected and marked automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link.
Dataset fields:
taskId - task identifier in the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/scoutieDataset_chinese_russian_dictionary_grammar_spelling_vectorized.2k_grammar_correctionsalia_gva_grammar
📘 ALIA_GVA_Grammar Dataset
The ALIA_GVA_Grammar dataset is a resource designed for text generation.
The dataset consists of textual documents formatted in Markdown (.md), each provided as structured JSONL entries.
Each entry includes information about the text's language, format, text, source, and metadata.
🧾 Column Descriptions
Field
Type
Description
format
string
Indicates the text format. All entries use "md" (Markdown).
language
string
Language of… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_gva_grammar.pashto-grammar-tutor
Pashto Grammar Tutor
A high-quality Pashto grammar instruction dataset designed for language learning, linguistic research, and supervised fine-tuning (SFT) of AI language models. The dataset focuses on grammatical analysis, verb conjugation, sentence structure, and teacher-style explanations written in Pashto.
Dataset Summary
Pashto Grammar Tutor is a specialized educational dataset containing grammar-focused instruction-response pairs. Each example presents a… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-grammar-tutor.kanitakorn-deepseek-v43-grammar-polarity-boundary-micro
Kanitakorn DeepSeek v43 Grammar Polarity Boundary Micro
Synthetic minimal-pair SFT shard for Thai grammar exact counting, polarity/exception wording, and close d/e option-boundary calibration.
Rows: 600
Pairs: 300, two rows each
Labels: a=b=c=d=e=120
Final answer contract: คำตอบ: (x)
Sources: generated original synthetic prompts only; no benchmark prompt/gold/sample rows used
Train SHA256: 863ea31d6039b9986f9647ae65bf0d4c382893361a59a6d70795fbd2d2e4739e
Use only as a tiny… See the full description on the dataset page: https://huggingface.co/datasets/Jnx03/kanitakorn-deepseek-v43-grammar-polarity-boundary-micro.indonesian-grammar-correction
Koreksi Tata Bahasa Indonesia ✍️
Kumpulan 126 pasangan kalimat (asli → koreksi) untuk grammar error correction bahasa Indonesia.
Kenapa dataset ini ada?
Grammar error correction (GEC) untuk bahasa Indonesia belum ada di HF — padahal model GEC global (271 downloads) jadi salah satu kategori paling dicari. Dataset ini isi gap itu: dari ejaan ("Dimana" → "Di mana"), ragam ("gw udah" → "saya sudah"), pleonasme ("para siswa-siswa" → "para siswa"), sampai huruf kapital.… See the full description on the dataset page: https://huggingface.co/datasets/LorthGyu/indonesian-grammar-correction.scoutieDataset_english_russian_dictionary_grammar_spelling_vectorized
Description in English:
A dataset collected from 30 Russian-language Telegram channels on the topic of learning English, this dataset contains grammar, syntax, spelling and punctuation rules, as well as English words with Russian translations.
The dataset was collected and marked automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link.
Dataset fields:
taskId - task identifier in the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/scoutieDataset_english_russian_dictionary_grammar_spelling_vectorized.
