grammar
Datasets
All datasets matching “grammar”coedit
Dataset Card for CoEdIT: Text Editing via Instruction Tuning
Paper: CoEdIT: Text Editing by Task-Specific Instruction Tuning
Authors: Vipul Raheja, Dhruv Kumar, Ryan Koo, Dongyeop Kang
Project Repo: https://github.com/vipulraheja/coedit
Dataset Summary
This is the dataset that was used to train the CoEdIT text editing models. Full details of the dataset can be found in our paper.
Dataset Structure
The dataset is in JSON format.… See the full description on the dataset page: https://huggingface.co/datasets/grammarly/coedit.essay-grammar-range-qwen3.5-4b-trl-completions
TRL Completion logs
This dataset contains the completions generated during training using trl.
Find the trained model at https://huggingface.co/bihungba1101/essay-grammar-range-qwen3.5-4b-grpo.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion
completion: the… See the full description on the dataset page: https://huggingface.co/datasets/bihungba1101/essay-grammar-range-qwen3.5-4b-trl-completions.grammar-accuracy-qwen3.5-4b-trl-completions
TRL Completion logs
This dataset contains the completions generated during training using trl.
Find the trained model at https://huggingface.co/bihungba1101/grammar-accuracy-qwen3.5-4b-grpo.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion
completion:… See the full description on the dataset page: https://huggingface.co/datasets/bihungba1101/grammar-accuracy-qwen3.5-4b-trl-completions.unified-grammar
Unified Distributional Grammar (Greek + Latin + Hebrew)
NuBerea/unified-grammar — the j-layer construction × function × source × era matrix, mirroring the
shape of NuBerea/distributional-lexicon (lemma × sense × source × era) for grammar instead of
lexicon: every grammatical claim scoped, counted, basis-carrying (see METRIC SEMANTICS below), and traceable to corpus
instances, with traditional grammars (Smyth, Gesenius-Kautzsch-Cowley, Allen & Greenough) admitted only
as witness… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/unified-grammar.allen-greenough-grammar
allen-greenough-grammar
Allen and Greenough's New Latin Grammar for Schools and Colleges (J.B.
Greenough, G.L. Kittredge, A.A. Howard, Benjamin L. D'Ooge, eds.; Boston:
Ginn & Company, 1903 — public domain), in the Dickinson College
Commentaries (DCC) digital re-edition
(https://dcc.dickinson.edu/grammar/latin/), edited by Meagan Ayer under
Chris Francese's direction, 2013-2016 ("A New Allen and Greenough"). This is
a PROSE-witness t0 source repo, the standard reference grammar… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/allen-greenough-grammar.smyth-grammar
smyth-grammar
Herbert Weir Smyth, A Greek Grammar for Colleges (New York: American Book
Company, 1920 — public domain). PROSE-witness t0 source repo for Greek
morphology and syntax, the sibling of allen-greenough-grammar (Latin) and
gesenius-kautzsch-grammar (Hebrew) in the distributional-grammar programme:
one row per numbered Smyth paragraph (§1–§3048, complete, plus the 213
"D"-suffixed dialect paragraphs), Greek examples in polytonic Unicode, Smyth's
own cross-references as… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/smyth-grammar.
