datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
coedit-multilingual
CoEdIT Multilingual
A multilingual text-editing dataset for evaluating instruction-following edit models
(grammar correction, paraphrase, simplification, neutralization, coherence, clarity)
plus operational speech-transcript edits (translation, number/punctuation formatting,
filler removal, summarization, formalize/casualize, abbreviation expansion).
Construction
English: the original grammarly/coedit (train + validation).
Translated (de, es, fr, it, pt, nl, zh… See the full description on the dataset page: https://huggingface.co/datasets/jacekduszenko/coedit-multilingual.gec-coherence-coedit-syntho1o2o3_large_r2_coedit_with_human_pref_practice
Dataset Card for "o1o2o3_large_r2_coedit_with_human_pref_practice"
More Information needed
coedit-cot-reasoning
Dataset Card for CoEdIT-CoT-Reasoning: Text Editing with Step-by-Step (Chain of Thought) Reasoning
Dataset Description
This dataset extends the original CoEdIT dataset by adding detailed step-by-step reasoning traces that explain how to perform various text editing tasks. The reasoning traces simulate the thought process of an expert editor applying the given instructions.
Dataset Summary
CoEdIT-Reasoning augments the CoEdIT text editing dataset with… See the full description on the dataset page: https://huggingface.co/datasets/muzzz/coedit-cot-reasoning.BEE-spoke-data-coedit-reworded-dedupedr2_coedit
Dataset Card for "r2_coedit"
More Information needed
r2_coedit_iter
Dataset Card for "r2_coedit_iter"
More Information needed
coedit_phraseso1o2o3_large_r2_coedit_iter_with_human_pref_practice
Dataset Card for "o1o2o3_large_r2_coedit_iter_with_human_pref_practice"
More Information needed
coedit_llmr2_coedit_v2
Dataset Card for "r2_coedit_v2"
More Information needed
merged-coedit-a1-bonafid-2point5-milgrammarly_coedit
Dataset Card for "grammarly_coedit"
More Information needed
coedit-reworded
coedit-reworded
This is Grammarly's coedit dataset parsed into Alpaca-style instruction, input, and output rows, with the original instruction values replaced with a more diverse set of procedurally generated instructions. Contains 23930 unique values of instruction, as compared to the original 144. See coedit_reword.py for how these were generated.
All credit to the original authors of this dataset.
Citation
@article{raheja2023coedit,
title={CoEdIT: Text… See the full description on the dataset page: https://huggingface.co/datasets/chargoddard/coedit-reworded.o1o2o3_xl_r2_coedit_iter_with_human_pref_practice
Dataset Card for "o1o2o3_xl_r2_coedit_iter_with_human_pref_practice"
More Information needed
Vocab-CoEdIT
Vocab-Coedit
Made with ❤️ using 🦥 Unsloth Studio
Vocab-CoEdIT contains 59,949 records of vocabulary suggestions
🚀 Quick Start
from datasets import load_dataset
# Load the main dataset
dataset = load_dataset("bihungba1101/Vocab-CoEdIT", "data", split="train")
df = dataset.to_pandas()
📊 Dataset Summary
📈 Records: 59,949
📋 Columns: 7
✅ Completion: 99.9% (60,000 requested)
📋 Schema & Statistics
Column
Type
Column Type
Unique (%)
Null (%)… See the full description on the dataset page: https://huggingface.co/datasets/bihungba1101/Vocab-CoEdIT.coedit-koTranslated grammarly/coedit using nayohan/llama3-instrucTrans-enko-8b.
This dataset is a raw translated dataset and contains repetitive sentences generated by the model, so it needs to be filtered.
@article{raheja2023coedit,
title={CoEdIT: Text Editing by Task-Specific Instruction Tuning},
author={Vipul Raheja and Dhruv Kumar and Ryan Koo and Dongyeop Kang},
year={2023},
eprint={2305.09857},
archivePrefix={arXiv},
primaryClass={cs.CL}
}
r2_coedit_v2_in_outr1_coedit_v2
Dataset Card for "r1_coedit_v2"
More Information needed
r1_coedit
Dataset Card for "r1_coedit"
More Information needed
o1o2o3_xl_r2_coedit_with_human_pref_practice
Dataset Card for "o1o2o3_xl_r2_coedit_with_human_pref_practice"
More Information needed
o1o2o3_xl_r2_coedit
Dataset Card for "o1o2o3_xl_r2_coedit"
More Information needed
coedit-fluencyo1o2o3_xl_r2_coedit_with_human_pref
Dataset Card for "o1o2o3_xl_r2_coedit_with_human_pref"
More Information needed
r1_coedit_iter
Dataset Card for "r1_coedit_iter"
More Information needed
Coeditor-processed-demo2
Dataset Card for "Coeditor-processed-demo2"
More Information needed
r2_coedit_v2_convEssay-Vocab-CoEdITo1o2o3_large_r2_coedit
Dataset Card for "o1o2o3_large_r2_coedit"
More Information needed
Coeditor-processed-demo
Dataset Card for "Coeditor-processed-demo"
More Information needed
