datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
coedit
Dataset Card for CoEdIT: Text Editing via Instruction Tuning
Paper: CoEdIT: Text Editing by Task-Specific Instruction Tuning
Authors: Vipul Raheja, Dhruv Kumar, Ryan Koo, Dongyeop Kang
Project Repo: https://github.com/vipulraheja/coedit
Dataset Summary
This is the dataset that was used to train the CoEdIT text editing models. Full details of the dataset can be found in our paper.
Dataset Structure
The dataset is in JSON format.… See the full description on the dataset page: https://huggingface.co/datasets/grammarly/coedit.CoEdit-AlpacaAn Alpaca instruction conversion of Grammarly's CoEdIT dataset.
coedit
Dataset Card for CoEdIT: Text Editing via Instruction Tuning
Paper: CoEdIT: Text Editing by Task-Specific Instruction Tuning
Authors: Vipul Raheja, Dhruv Kumar, Ryan Koo, Dongyeop Kang
Project Repo: https://github.com/vipulraheja/coedit
Dataset Summary
This is the dataset that was used to train the CoEdIT text editing models. Full details of the dataset can be found in our paper.
Dataset Structure
The… See the full description on the dataset page: https://huggingface.co/datasets/billcavalieri/coedit.coedit-reworded-deduped-multiturn-sharegptEach sample contains 1 to 32 pairs.
pair_lengths Minimum: 44
pair_lengths Maximum: 2573
pair_lengths Average: 1053
turn_counts Minimum: 1
turn_counts Maximum: 32
turn_counts Average: 16.5
pair_lengths counted using metharme tags, and mistral tokenizer
turn_counts counted for pairs of human/gpt
CoEditLlama
