CoolFace
20 results

gec

mideind /gec-test-setTest data for Icelandic spell and grammar checking, created as part of the Icelandic Language Technology Programme. The test data is divided into three different formats, type 1, 2 and 3. For every original file corrected, three files are included in the test data when possible: _original, _corrected and _metadata. The original and metadata files are always .txt files, but the format of the corrected file differs between types. Texts corrected are from the News2 subcorpus of the Icelandic… See the full description on the dataset page: https://huggingface.co/datasets/mideind/gec-test-set.0 likes1.7k downloads2y agoHugging FaceDipeshChaudhary /nepali-gector-style-token-level-tag-for-ged Nepali GEC (gector style) Token Tagging Dataset This is a processed version of the sumitaryal/nepali_grammatical_error_correction dataset, designed for training GEC-ToR-style sequence tagging models. This dataset has been processed with a robust, multi-pass, content-aware alignment algorithm to generate high-fidelity correction tags, including complex and adjacent SWAP operations. Total Examples: 16,260,992 Training: 13,008,711 Validation: 2,439,231 test: 813,050… See the full description on the dataset page: https://huggingface.co/datasets/DipeshChaudhary/nepali-gector-style-token-level-tag-for-ged.texttoken-classification10M<n<100M0 likes596 downloads11mo agoHugging FaceDipeshChaudhary /muril-nepali-gector-style-token-level-tag-for-ged Nepali GEC (gector style) Token Tagging Dataset This is a processed version of the sumitaryal/nepali_grammatical_error_correction dataset, designed for training GEC-ToR-style sequence tagging models. This dataset has been processed with a robust, multi-pass, content-aware alignment algorithm to generate high-fidelity correction tags, including complex and adjacent SWAP operations. Total Examples: 16,260,992 Training: 13,008,711 Validation: 2,439,231 test: 813,050… See the full description on the dataset page: https://huggingface.co/datasets/DipeshChaudhary/muril-nepali-gector-style-token-level-tag-for-ged.texttoken-classification10M<n<100M0 likes473 downloads11mo agoHugging Facejerpelhan /geco2-assets0 likes436 downloads1mo agoHugging Facehafidikhsan /c4_200m-gec-train100k-test25k Dataset Card for "c4_200m-gec-train100k-test25k" More Information needed tabular100K<n<1M0 likes419 downloads3y agoHugging Faceltg /ask-gec Norwegian grammatical error correction (ASK) This is the ASK-RAW dataset by Matias Jentoft (2023). Cite @mastersthesis{jentoft2023grammatical, title={Grammatical Error Correction with byte-level language models}, author={Jentoft, Matias}, year={2023}, school={University of Oslo}, url={https://www.duo.uio.no/handle/10852/103885} } text10K<n<100K4 likes314 downloads3y agoHugging Face