CoolFace
Datasetpublic

astrideducation/cefr-combined-no-cefr-test

This dataset contains 3370555 sentences, which each have an assigned CEFR level derived from EFLLex (https://cental.uclouvain.be/cefrlex/efllex/download). The sentences comes from "the pile books3", which is available on Huggingface (https://huggingface.co/datasets/the_pile_books3). The CEFR levels used are A1, A2, B1, B2 and C1, and there are equals number of sentences for each level. Assigning each sentence a CEFR level followed is based on the concept of "shifted frequency distribution", introduced by David Alfter and his paper can be found at (https://gupea.ub.gu.se/bitstream/2077/66861/4/gupea_2077_66861_4.pdf). For each word in each sentence, take the CEFR level with the highest "shifted frequency distribution" in the EFLLex table. After all words have been processed, the sentence gets annotated with the most frequently appearing CEFR level from the whole senctence.

sourceHugging Faceupdated 5y agoView on Hugging Face
1likes170downloads
8 commits on main
3b71e625y ago

reduces values to unpack from

vasilis
5b63f765y ago

Update cefr-combined-no-cefr-test.py

sebastiaan
da165735y ago

Upload test_dataset_wo_cefr.csv

sebastiaan
51b9da75y ago

Delete test_dataset_wo_cefr.csv

sebastiaan
3c911535y ago

Update cefr-combined-no-cefr-test.py

sebastiaan
4f1df9b5y ago

Upload test_dataset_wo_cefr.csv

sebastiaan
4ccb7d85y ago

Create cefr-combined-no-cefr-test.py

sebastiaan
25f22af5y ago

initial commit

system