CoolFace
Datasetpublic

agentlans/grammar-correction

grammar-correction Dataset Summary The grammar-correction dataset is a refined subset of the liweili/c4_200m dataset, derived from Google's C4_200M Synthetic Dataset for Grammatical Error Correction. It contains sentence pairs where the input is ungrammatical and the output is grammatical, making it suitable for training grammatical error correction (GEC) models. Dataset Structure Train set: 100 000 entries Validation set: 25 000 entries… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/grammar-correction.

sourceHugging Faceupdated 2y agoView on Hugging Face
11likes386downloads

No commit history came back for main. The revision may not exist, or the source declined the request.