astrideducation/cefr-combined-no-cefr-test
This dataset contains 3370555 sentences, which each have an assigned CEFR level derived from EFLLex (https://cental.uclouvain.be/cefrlex/efllex/download). The sentences comes from "the pile books3", which is available on Huggingface (https://huggingface.co/datasets/the_pile_books3). The CEFR levels used are A1, A2, B1, B2 and C1, and there are equals number of sentences for each level. Assigning each sentence a CEFR level followed is based on the concept of "shifted frequency distribution", introduced by David Alfter and his paper can be found at (https://gupea.ub.gu.se/bitstream/2077/66861/4/gupea_2077_66861_4.pdf). For each word in each sentence, take the CEFR level with the highest "shifted frequency distribution" in the EFLLex table. After all words have been processed, the sentence gets annotated with the most frequently appearing CEFR level from the whole senctence.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face