CoolFace
Datasetpublic

mideind/icelandic-error-corpus-IceEC

The Icelandic Error Corpus (IceEC) is a collection of texts in modern Icelandic annotated for mistakes related to spelling, grammar, and other issues. The texts are organized by genre. The current version includes sentences from student essays, online news texts and Wikipedia articles. Sentences within texts in the student essays had to be shuffled due to the license which they were originally published under, but neither the online news texts nor the Wikipedia articles needed to be shuffled.

sourceHugging Facecc-by-4.0updated 4y agoView on Hugging Face
1likes49downloads
Dataset Card

Icelandic Error Corpus

Refer to https://github.com/antonkarl/iceErrorCorpus for a description of the dataset.

Please cite the dataset as follows if you use it.

Anton Karl Ingason, Lilja Björk Stefánsdóttir, Þórunn Arnardóttir, and Xindan Xu. 2021. The Icelandic Error Corpus (IceEC). Version 1.1. (https://github.com/antonkarl/iceErrorCorpus)