gec
Datasets
All datasets matching “gec”gec-test-setTest data for Icelandic spell and grammar checking, created as part of the Icelandic Language Technology Programme.
The test data is divided into three different formats, type 1, 2 and 3. For every original file corrected, three files are included in the test data when possible: _original, _corrected and _metadata. The original and metadata files are always .txt files, but the format of the corrected file differs between types.
Texts corrected are from the News2 subcorpus of the Icelandic… See the full description on the dataset page: https://huggingface.co/datasets/mideind/gec-test-set.nepali-gector-style-token-level-tag-for-ged
Nepali GEC (gector style) Token Tagging Dataset
This is a processed version of the sumitaryal/nepali_grammatical_error_correction dataset,
designed for training GEC-ToR-style sequence tagging models.
This dataset has been processed with a robust, multi-pass, content-aware alignment algorithm
to generate high-fidelity correction tags, including complex and adjacent SWAP operations.
Total Examples: 16,260,992
Training: 13,008,711
Validation: 2,439,231
test: 813,050… See the full description on the dataset page: https://huggingface.co/datasets/DipeshChaudhary/nepali-gector-style-token-level-tag-for-ged.muril-nepali-gector-style-token-level-tag-for-ged
Nepali GEC (gector style) Token Tagging Dataset
This is a processed version of the sumitaryal/nepali_grammatical_error_correction dataset,
designed for training GEC-ToR-style sequence tagging models.
This dataset has been processed with a robust, multi-pass, content-aware alignment algorithm
to generate high-fidelity correction tags, including complex and adjacent SWAP operations.
Total Examples: 16,260,992
Training: 13,008,711
Validation: 2,439,231
test: 813,050… See the full description on the dataset page: https://huggingface.co/datasets/DipeshChaudhary/muril-nepali-gector-style-token-level-tag-for-ged.geco2-assetsc4_200m-gec-train100k-test25k
Dataset Card for "c4_200m-gec-train100k-test25k"
More Information needed
ask-gec
Norwegian grammatical error correction (ASK)
This is the ASK-RAW dataset by Matias Jentoft (2023).
Cite
@mastersthesis{jentoft2023grammatical,
title={Grammatical Error Correction with byte-level language models},
author={Jentoft, Matias},
year={2023},
school={University of Oslo},
url={https://www.duo.uio.no/handle/10852/103885}
}
