CoolFace
Datasetpublic

rodriguescarson/adaption-minimal-diff-proofreading-12k

Minimal-Diff Proofreading Proofread a sentence with the fewest possible edits (one to three tokens), explain the change, then give the corrected sentence. Rows 12,000 Domain writing and editing Format data.parquet, one row per example Licence apache-2.0 Built for supervised fine-tuning (SFT) experiments on Adaption AutoScientist Columns Column Description original_prompt The prompt (user turn) as uploaded. original_completion… See the full description on the dataset page: https://huggingface.co/datasets/rodriguescarson/adaption-minimal-diff-proofreading-12k.

sourceHugging Faceapache-2.0updated 13h agoView on Hugging Face
0likes6downloads
Dataset Card

Minimal-Diff Proofreading

Proofread a sentence with the fewest possible edits (one to three tokens), explain the change, then give the corrected sentence.

Rows12,000
Domainwriting and editing
Formatdata.parquet, one row per example
Licenceapache-2.0
Built forsupervised fine-tuning (SFT) experiments on Adaption AutoScientist

Columns

ColumnDescription
original_promptThe prompt (user turn) as uploaded.
original_completionThe target response as uploaded.
enhanced_promptEmpty in this dataset.
enhanced_completionEmpty in this dataset.

How it was built

Filtered from CoEdIT's grammar-correction split to pairs whose fix is 1 to 3 token edits; each explanation is derived mechanically from the diff, so nothing is invented.

Sources and licence

Notes

  • —Columns enhanced_prompt, enhanced_completion are empty in this dataset (the Adaption export reserves them for rewritten text).

Loading

python
from datasets import load_dataset
ds = load_dataset("rodriguescarson/adaption-minimal-diff-proofreading-12k", split="train")
python
import pandas as pd
df = pd.read_parquet("hf://datasets/rodriguescarson/adaption-minimal-diff-proofreading-12k/data.parquet")

Published by Carson Rodrigues (Hugging Face, Kaggle).