datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
coedit
Dataset Card for CoEdIT: Text Editing via Instruction Tuning
Paper: CoEdIT: Text Editing by Task-Specific Instruction Tuning
Authors: Vipul Raheja, Dhruv Kumar, Ryan Koo, Dongyeop Kang
Project Repo: https://github.com/vipulraheja/coedit
Dataset Summary
This is the dataset that was used to train the CoEdIT text editing models. Full details of the dataset can be found in our paper.
Dataset Structure
The dataset is in JSON format.… See the full description on the dataset page: https://huggingface.co/datasets/grammarly/coedit.medit
Dataset Card for mEdIT: Multilingual Text Editing via Instruction Tuning
Paper: mEdIT: Multilingual Text Editing via Instruction Tuning
Authors: Vipul Raheja, Dimitris Alikaniotis, Vivek Kulkarni, Bashar Alhafni, Dhruv Kumar
Project Repo: https://github.com/vipulraheja/medit
Dataset Summary
This is the dataset that was used to train the mEdIT text editing models. Full details of the dataset can be found in our paper.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/grammarly/medit.spivavtor
Dataset Card for Spivavtor
Paper: Spivavtor: An Instruction Tuned Ukrainian Text Editing Model
Authors: Aman Saini, Artem Chernodub, Vipul Raheja, Vivek Kulkarni
Dataset Summary
This is the dataset used to train all Spivavtor models. It contains data for 4 tasks - Grammatical Error Correction (GEC), Simplification, Coherence and Paraphrasing.
The specific details are as follows:
Task
Examples in Training data
Examples in Validation… See the full description on the dataset page: https://huggingface.co/datasets/grammarly/spivavtor.Pashto-grammar-100
🇦🇫 Pashto Grammar 100
Pashto Grammar 100 is a compact, focused dataset created to help AI models learn and understand fundamental Pashto grammar, sentence structure, grammatical concepts, and correct linguistic usage.
The dataset contains carefully selected Pashto grammar examples designed for language learning, grammatical analysis, instruction tuning, and evaluation of Pashto language models.
It is intended as a small but high-quality resource for researchers and developers… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-grammar-100.Punjabi-Gurmukhi-Grammar-Correction-Corpus
ੴ Punjabi (Gurmukhi) Grammatical Error Correction Corpus
☬ ਪੰਜਾਬੀ (ਗੁਰਮੁਖੀ) ਵਿਆਕਰਣ ਸ਼ੁੱਧੀ ਅਤੇ ਸੁਧਾਰ ਡਾਟਾਸੈੱਟ (v1.0)
👨💻 Research & Engineering Lead
Creator & Architect: Gurpreet Singh Dhillon (Nam-toon Studio)
GitHub Profile: github.com/gurpreetsingh5523-source
Flagship Innovation: AMRIT Research OS (100% Locally-Run Autonomous Medical AI)
📖 Overview / ਸੰਖੇਪ
The Punjabi (Gurmukhi) Grammatical Error Correction… See the full description on the dataset page: https://huggingface.co/datasets/Nam-toon-studio/Punjabi-Gurmukhi-Grammar-Correction-Corpus.pashto-stf-grammar-pairs
Pashto SFT Grammar Pairs
Dataset Description
Pashto SFT Grammar Pairs is a native-speaker-curated collection of Pashto question–answer pairs focused on Pashto grammar (ګرامر), covering topics such as noun gender, number, case (فاعلي، مفعولي، اضافي), adjective agreement, pronouns, verb conjugation, sentence structure (SOV word order), and enclitics/suffixes.
The dataset is formatted in the {"messages": [...]} chat-template style used by modern SFT pipelines (TRL… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-stf-grammar-pairs.byzantine-synthetic-grammar
Byzantine Synthetic Interval-Grammar — Neume ↔ Western Pitch
A correct-by-construction supervised fine-tuning dataset for transcribing between Byzantine
(Chrysanthine) neume notation and Western staff pitches, in both directions, over a diatonic
interval grammar with ascending leaps up to an octave and rhythmic note durations.
This dataset trains the model
FableMogger9000/byzantine-synthetic-grammar-lora,
which reaches 96% exact-match / 98% melodic equivalence on the held-out… See the full description on the dataset page: https://huggingface.co/datasets/FableMogger9000/byzantine-synthetic-grammar.english_grammar_sft_system
📘 Dataset Overview
english_grammar_sft_system is a curated supervised fine‑tuning (SFT) dataset derived from the Teravee/1000_english-grammar-dataset.Each sample is converted into a three‑message conversation format:
system — global behavior instruction
user — instruction + question merged
assistant — high‑quality answer
This structure is ideal for training instruction‑following LLMs that must respond clearly, concisely, and accurately to grammar‑related queries.… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/english_grammar_sft_system.alia_gva_grammar
📘 ALIA_GVA_Grammar Dataset
The ALIA_GVA_Grammar dataset is a resource designed for text generation.
The dataset consists of textual documents formatted in Markdown (.md), each provided as structured JSONL entries.
Each entry includes information about the text's language, format, text, source, and metadata.
🧾 Column Descriptions
Field
Type
Description
format
string
Indicates the text format. All entries use "md" (Markdown).
language
string
Language of… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_gva_grammar.pashto-grammar-tutor
Pashto Grammar Tutor
A high-quality Pashto grammar instruction dataset designed for language learning, linguistic research, and supervised fine-tuning (SFT) of AI language models. The dataset focuses on grammatical analysis, verb conjugation, sentence structure, and teacher-style explanations written in Pashto.
Dataset Summary
Pashto Grammar Tutor is a specialized educational dataset containing grammar-focused instruction-response pairs. Each example presents a… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-grammar-tutor.kanitakorn-deepseek-v43-grammar-polarity-boundary-micro
Kanitakorn DeepSeek v43 Grammar Polarity Boundary Micro
Synthetic minimal-pair SFT shard for Thai grammar exact counting, polarity/exception wording, and close d/e option-boundary calibration.
Rows: 600
Pairs: 300, two rows each
Labels: a=b=c=d=e=120
Final answer contract: คำตอบ: (x)
Sources: generated original synthetic prompts only; no benchmark prompt/gold/sample rows used
Train SHA256: 863ea31d6039b9986f9647ae65bf0d4c382893361a59a6d70795fbd2d2e4739e
Use only as a tiny… See the full description on the dataset page: https://huggingface.co/datasets/Jnx03/kanitakorn-deepseek-v43-grammar-polarity-boundary-micro.indonesian-grammar-correction
Koreksi Tata Bahasa Indonesia ✍️
Kumpulan 126 pasangan kalimat (asli → koreksi) untuk grammar error correction bahasa Indonesia.
Kenapa dataset ini ada?
Grammar error correction (GEC) untuk bahasa Indonesia belum ada di HF — padahal model GEC global (271 downloads) jadi salah satu kategori paling dicari. Dataset ini isi gap itu: dari ejaan ("Dimana" → "Di mana"), ragam ("gw udah" → "saya sudah"), pleonasme ("para siswa-siswa" → "para siswa"), sampai huruf kapital.… See the full description on the dataset page: https://huggingface.co/datasets/LorthGyu/indonesian-grammar-correction.
