CoolFace
Datasetpublic

astrideducation/cefr-combined-no-cefr-test

This dataset contains 3370555 sentences, which each have an assigned CEFR level derived from EFLLex (https://cental.uclouvain.be/cefrlex/efllex/download). The sentences comes from "the pile books3", which is available on Huggingface (https://huggingface.co/datasets/the_pile_books3). The CEFR levels used are A1, A2, B1, B2 and C1, and there are equals number of sentences for each level. Assigning each sentence a CEFR level followed is based on the concept of "shifted frequency distribution", introduced by David Alfter and his paper can be found at (https://gupea.ub.gu.se/bitstream/2077/66861/4/gupea_2077_66861_4.pdf). For each word in each sentence, take the CEFR level with the highest "shifted frequency distribution" in the EFLLex table. After all words have been processed, the sentence gets annotated with the most frequently appearing CEFR level from the whole senctence.

sourceHugging Faceupdated 5y agoView on Hugging Face
1likes175downloads
Dataset Card

No card is published for this repository, or it could not be fetched from Hugging Face right now.