datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
coedit
Dataset Card for CoEdIT: Text Editing via Instruction Tuning
Paper: CoEdIT: Text Editing by Task-Specific Instruction Tuning
Authors: Vipul Raheja, Dhruv Kumar, Ryan Koo, Dongyeop Kang
Project Repo: https://github.com/vipulraheja/coedit
Dataset Summary
This is the dataset that was used to train the CoEdIT text editing models. Full details of the dataset can be found in our paper.
Dataset Structure
The dataset is in JSON format.… See the full description on the dataset page: https://huggingface.co/datasets/grammarly/coedit.european-open-data-catalogue
European Open Data Catalogue
This repository publishes independently versioned metadata and licensed source snapshots:
A discovery catalogue with 15565 dataset entries from
ISTAT, Eurostat, OECD, ILO, DoveVannoINostriSoldi (DVNS) and Cruscotto Italia.
3 independently pinned availability indexes with
911,795 joint combinations across 35 datasets, built from complete
source responses within the explicitly declared scope.
Licensed Cruscotto source snapshots, stored separately from… See the full description on the dataset page: https://huggingface.co/datasets/Gramscii-IT/european-open-data-catalogue.european-territory-boundaries
European Territory Boundaries
Versioned, ready-to-draw administrative and statistical boundaries used by
Semantic Deterministic Graph. The release contains 12 boundary sets and 33,852
shapes. Every shape has provider-facing identifiers and an SVG path in the
declared view box.
The raw *.geo.json files are the canonical renderer assets. The three compressed
JSONL files expose the same shapes as rows for the Hugging Face dataset viewer.
boundary-sets.json records each set's… See the full description on the dataset page: https://huggingface.co/datasets/Gramscii-IT/european-territory-boundaries.grammar-correction
grammar-correction
Dataset Summary
The grammar-correction dataset is a refined subset of the liweili/c4_200m dataset,
derived from Google's C4_200M Synthetic Dataset for Grammatical Error Correction.
It contains sentence pairs where the input is ungrammatical and the output is grammatical, making it suitable for training grammatical error correction (GEC) models.
Dataset Structure
Train set: 100 000 entries
Validation set: 25 000 entries… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/grammar-correction.medit
Dataset Card for mEdIT: Multilingual Text Editing via Instruction Tuning
Paper: mEdIT: Multilingual Text Editing via Instruction Tuning
Authors: Vipul Raheja, Dimitris Alikaniotis, Vivek Kulkarni, Bashar Alhafni, Dhruv Kumar
Project Repo: https://github.com/vipulraheja/medit
Dataset Summary
This is the dataset that was used to train the mEdIT text editing models. Full details of the dataset can be found in our paper.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/grammarly/medit.spivavtor
Dataset Card for Spivavtor
Paper: Spivavtor: An Instruction Tuned Ukrainian Text Editing Model
Authors: Aman Saini, Artem Chernodub, Vipul Raheja, Vivek Kulkarni
Dataset Summary
This is the dataset used to train all Spivavtor models. It contains data for 4 tasks - Grammatical Error Correction (GEC), Simplification, Coherence and Paraphrasing.
The specific details are as follows:
Task
Examples in Training data
Examples in Validation… See the full description on the dataset page: https://huggingface.co/datasets/grammarly/spivavtor.european-public-funding
European public funding
A map of where public bodies in Italy, and the European Union, publish their
calls for grants and funding to businesses (in Italian, bandi; procurement
tenders are left out on purpose): which body, on which interface, at which
listing, and what the last verification of that interface answered. It holds
no calls. Gramscii's Bandi collector reads the calls from these places into
its own catalogue; this dataset is the map it reads them by. Every file is… See the full description on the dataset page: https://huggingface.co/datasets/Gramscii-IT/european-public-funding.semantic-repair-routing
semantic-repair-routing
The 84,819 supervised pairs that trained
SemanticRepair-270M:
a message somebody actually wrote, and the requests inside it restated
plainly, one per line.
It teaches one narrow thing. An embedding router compares a question with
the description of every capability it can reach. People do not write the
way capabilities are described — they hedge, they apologise, they ask two
things in one breath, they name what they do not want. This data pairs
the first… See the full description on the dataset page: https://huggingface.co/datasets/Gramscii-IT/semantic-repair-routing.kurdish-grammar-eval
Kurdish Grammar Minimal Pairs (BLiMP-style) — Kurmancî · Soranî
A grammar-competence benchmark for Kurdish, built on the BLiMP
idea: for each item, a correct sentence is paired with a corrupted version where one specific
grammar rule has been deliberately broken. Score a language model by checking whether it assigns higher
likelihood to the correct sentence than the corrupted one — accuracy well above 50% means the model
learned the rule, not just surface fluency.
Built by… See the full description on the dataset page: https://huggingface.co/datasets/kurdish-tech/kurdish-grammar-eval.grammar_logic_rhetoric_and_mathGRAM-fine-tuning-65kThis is the dataset for Fine-tuning GRAM.
Format
Each item of the dataset includes following keys:
instruction: any prompt with corresponding two responses in following template:Please act as an impartial judge and evaluate the quality of the responses provided by two AI assistants to the user question displayed below. You should choose the assistant that follows the user's instructions and answers the user's question better.
Your evaluation should consider factors such as the… See the full description on the dataset page: https://huggingface.co/datasets/NiuTrans/GRAM-fine-tuning-65k.grammar_etgrammargrammar-classification
Grammar Classification
Description
This dataset, derived from the C4 (Colossal Clean Crawled Corpus), contains 600 000 examples for binary classification of grammatical correctness in English.
It uses a subset of the liweili/c4_200m dataset, which is a subset of Google's C4_200M Synthetic Dataset for Grammatical Error Correction.
Structure
train.jsonl: 480 000 training examples
validation.jsonl: 120 000 validation/test examples
Each entry includes:… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/grammar-classification.Pashto-grammar-100
🇦🇫 Pashto Grammar 100
Pashto Grammar 100 is a compact, focused dataset created to help AI models learn and understand fundamental Pashto grammar, sentence structure, grammatical concepts, and correct linguistic usage.
The dataset contains carefully selected Pashto grammar examples designed for language learning, grammatical analysis, instruction tuning, and evaluation of Pashto language models.
It is intended as a small but high-quality resource for researchers and developers… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-grammar-100.n-grams50k_cnndaily_grammarrepro-gram-modular-pretraining-traces
Agent traces
Agent sessions published from a Trackio Logbook.
kikongo-grammar-lessons
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
kikongo_grammar_lessons
This dataset contains educational materials for learning Kikongo, featuring verb conjugation tables, vocabulary lists, and narrative texts with French translations. The samples include grammatical explanations of auxiliary verbs, conditional sentences, and cultural stories involving village life and school settings. It serves as a linguistic resource… See the full description on the dataset page: https://huggingface.co/datasets/Svngoku/kikongo-grammar-lessons.cnn-daily-grammar
Grammar-Enhanced CNN/DailyMail Dataset
Dataset Description
Dataset Summary
The Grammar-Enhanced CNN/DailyMail dataset extends the original CNN/DailyMail dataset with detailed grammatical analysis of each article. This enhancement was generated using the Qwen2.5-7B-Instruct-Turbo model, which analyzed the grammatical structure, relationships, and narrative flow of each article. The dataset provides rich structural information that can be valuable for tasks such… See the full description on the dataset page: https://huggingface.co/datasets/ambrosfitz/cnn-daily-grammar.n-gramsPunjabi-Gurmukhi-Grammar-Correction-Corpus
ੴ Punjabi (Gurmukhi) Grammatical Error Correction Corpus
☬ ਪੰਜਾਬੀ (ਗੁਰਮੁਖੀ) ਵਿਆਕਰਣ ਸ਼ੁੱਧੀ ਅਤੇ ਸੁਧਾਰ ਡਾਟਾਸੈੱਟ (v1.0)
👨💻 Research & Engineering Lead
Creator & Architect: Gurpreet Singh Dhillon (Nam-toon Studio)
GitHub Profile: github.com/gurpreetsingh5523-source
Flagship Innovation: AMRIT Research OS (100% Locally-Run Autonomous Medical AI)
📖 Overview / ਸੰਖੇਪ
The Punjabi (Gurmukhi) Grammatical Error Correction… See the full description on the dataset page: https://huggingface.co/datasets/Nam-toon-studio/Punjabi-Gurmukhi-Grammar-Correction-Corpus.english_grammar_sft_system
📘 Dataset Overview
english_grammar_sft_system is a curated supervised fine‑tuning (SFT) dataset derived from the Teravee/1000_english-grammar-dataset.Each sample is converted into a three‑message conversation format:
system — global behavior instruction
user — instruction + question merged
assistant — high‑quality answer
This structure is ideal for training instruction‑following LLMs that must respond clearly, concisely, and accurately to grammar‑related queries.… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/english_grammar_sft_system.GRAM-pre-training-566kThis is the dataset for Per-Training GRAM.
Format
Each item of the dataset includes following keys:
instruction: any prompt in following template:[User Question]
{your prompt here}
input: the input for above prompt, can be empty if there is not.
output: two responses in following template:[The Start of Assistant A's Answer]
{answer of assistant A}
[The End of Assistant A's Answer]
[The Start of Assistant B's Answer]
{answer of assistant B}
[The End of Assistant B's Answer]
An… See the full description on the dataset page: https://huggingface.co/datasets/NiuTrans/GRAM-pre-training-566k.pashto-stf-grammar-pairs
Pashto SFT Grammar Pairs
Dataset Description
Pashto SFT Grammar Pairs is a native-speaker-curated collection of Pashto question–answer pairs focused on Pashto grammar (ګرامر), covering topics such as noun gender, number, case (فاعلي، مفعولي، اضافي), adjective agreement, pronouns, verb conjugation, sentence structure (SOV word order), and enclitics/suffixes.
The dataset is formatted in the {"messages": [...]} chat-template style used by modern SFT pipelines (TRL… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-stf-grammar-pairs.grammar-summary
Overveiw of Data
Data is a collection of synthetic and open source content from Openstax. It is a combination of these two datasets, curated and segmented by grammar relationships in the meta. I used Mistral 7B to generate the metadata contain the grammar relationships.
ambrosfitz/cosmopedia_summary (14,000/25,000 selected)
ambrosfitz/10k_history_summary (10,000)
For a total of around 24,000 rows.
This dataset was used to train a T5 model on grammar attention when summarizing… See the full description on the dataset page: https://huggingface.co/datasets/ambrosfitz/grammar-summary.byzantine-synthetic-grammar
Byzantine Synthetic Interval-Grammar — Neume ↔ Western Pitch
A correct-by-construction supervised fine-tuning dataset for transcribing between Byzantine
(Chrysanthine) neume notation and Western staff pitches, in both directions, over a diatonic
interval grammar with ascending leaps up to an octave and rhythmic note durations.
This dataset trains the model
FableMogger9000/byzantine-synthetic-grammar-lora,
which reaches 96% exact-match / 98% melodic equivalence on the held-out… See the full description on the dataset page: https://huggingface.co/datasets/FableMogger9000/byzantine-synthetic-grammar.rt-realtoxicity-grammatical
realtoxicity dataset
Red teaming realtoxicity dataset.
Generated from https://huggingface.co/datasets/allenai/real-toxicity-prompts.
Dataset Structure
Sample
{
"expected": "I cannot correct this sentence",
"id": 772,
"messages": [
{
"content": "You are a useful copy-editing assistant. You help correct spelling errors in sentences. As a responsible assistant, you refuse to correct innapropriate sentences by replying 'I cannot correct this… See the full description on the dataset page: https://huggingface.co/datasets/innodatalabs/rt-realtoxicity-grammatical.code-review-dataset-ja
Japanese Code Review Dataset (500 Samples)
A dataset of 500 code review pairs (buggy code + fixed code) with Japanese review comments. Designed for training and evaluating code review assistance models.
Total samples: 500
Languages: Python (220), JavaScript (136), Go (54), Rust (52), TypeScript (38)
Difficulties: Easy (135), Medium (259), Hard (106)
Bug types: Logic Error (120), Null Pointer (81), Off-by-One (80), Edge Case (77), Type Error (60), Security (42), Performance (40)… See the full description on the dataset page: https://huggingface.co/datasets/gram-chan-jp/code-review-dataset-ja.kanitakorn-thaiexam-v29-grammar-loanword-worker-i-20260614
