datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
french_CEFRfrench_CEFRCEFR_Mixed_Dataset_1CEFR-Annotated-WordNet
CEFR-Annotated WordNet: LLM-Based Proficiency-Guided Semantic Database for Language Learning
Paper: CEFR-Annotated WordNet: LLM-Based Proficiency-Guided Semantic Database for Language Learning
Authors: Masato Kikuchi, Masatsugu Ono, Toshioki Soga, Tetsu Tanabe, Tadachika Ozono
Overview
CEFR-Annotated WordNet is a comprehensive semantic database that maps WordNet 3.0 senses and SemCor 3.0 instances to CEFR (Common European Framework of Reference for Languages)… See the full description on the dataset page: https://huggingface.co/datasets/star092304/CEFR-Annotated-WordNet.CEFR-Sentence-Level-Annotations
Dataset Card for Dataset Name
17k english sentences annotated by english education professionals. Original repo for CEFR-SP is located at this repo
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP):… See the full description on the dataset page: https://huggingface.co/datasets/edesaras/CEFR-Sentence-Level-Annotations.cefr_sp_enThis dataset has been indexed in the UniversalCEFR. The transformed version (in JSON format) retains the same license as the original dataset. Ownership and copyright remain with the original creators and/or dataset paper authors. If you use this transformed dataset, you must cite the following:
Dataset License: cc-by-nc-sa-4.0
Dataset Repository: https://github.com/yukiar/CEFR-SP/tree/main/CEFR-SP
Original Dataset Paper: Yuki Arase, Satoru Uchida, and Tomoyuki Kajiwara. 2022. CEFR-Based… See the full description on the dataset page: https://huggingface.co/datasets/UniversalCEFR/cefr_sp_en.cefr-combined-no-cefr-testThis dataset contains 3370555 sentences, which each have an assigned CEFR level derived from EFLLex (https://cental.uclouvain.be/cefrlex/efllex/download).
The sentences comes from "the pile books3", which is available on Huggingface (https://huggingface.co/datasets/the_pile_books3).
The CEFR levels used are A1, A2, B1, B2 and C1, and there are equals number of sentences for each level.
Assigning each sentence a CEFR level followed is based on the concept of "shifted frequency distribution", introduced by David Alfter and his paper can be found at (https://gupea.ub.gu.se/bitstream/2077/66861/4/gupea_2077_66861_4.pdf).
For each word in each sentence, take the CEFR level with the highest "shifted frequency distribution" in the EFLLex table.
After all words have been processed, the sentence gets annotated with the most frequently appearing CEFR level from the whole senctence.elg_cefr_enThis dataset has been indexed in the UniversalCEFR. The transformed version (in JSON format) retains the same license as the original dataset. Ownership and copyright remain with the original creators and/or dataset paper authors. If you use this transformed dataset, you must cite the following:
Dataset License: cc-by-nc-4.0
Dataset Repository: https://www.edia.nl/resources/elg/downloads
Original Dataset Paper: Breuker, M. (2023). CEFR Labelling and Assessment Services. In: Rehm, G. (eds)… See the full description on the dataset page: https://huggingface.co/datasets/UniversalCEFR/elg_cefr_en.CEFR_Mixed_Dataset_A1_A2_sisa
CEFR Dataset for A1 and A2
This dataset combines original CEFR-level sentences from training, validation, and test sets with synthetic sentences generated by a fine-tuned LLaMA-3-8B model for CEFR levels A1 (2000 sentences) and A2 (100 sentences). Synthetic sentences were validated using a fine-tuned MLP classifier (~93% accuracy) to ensure the predicted CEFR level is within 1 level of the intended level (e.g., A1 accepts A1, A2; A2 accepts A1, A2, B1). Duplicate sentences were… See the full description on the dataset page: https://huggingface.co/datasets/Mr-FineTuner/CEFR_Mixed_Dataset_A1_A2_sisa.cefr_asag_enThis dataset has been indexed in the UniversalCEFR. The transformed version (in JSON format) retains the same license as the original dataset. Ownership and copyright remain with the original creators and/or dataset paper authors. If you use this transformed dataset, you must cite the following:
Dataset License: cc-by-nc-sa-4.0
Dataset Repository: https://github.com/anaistack/cefr-asag-corpus
Original Dataset Paper: Anaïs Tack, Thomas François, Sophie Roekhaut, and Cédrick Fairon. 2017. Human… See the full description on the dataset page: https://huggingface.co/datasets/UniversalCEFR/cefr_asag_en.english_cefr_datasetCEFR_vocab_tokensDataset with English words classified along CEFR categories, tokenized forms based on sentencepiece tokenizer.
License based on foundational dataset, accessible at: http://www.englishprofile.org/wordlists/terms-of-use
Makxxx-french_CEFRCe répertoire est vide, il a été créé pour améliorer le référencement du jeu de données huggingface.co/datasets/Makxxx/french_CEFR.
vekkt-french_CEFRCe répertoire est vide, il a été créé pour améliorer le référencement du jeu de données huggingface.co/datasets/vekkt/french_CEFR.
melolingua-cefr-graded-multilingual-stories
MeloLingua CEFR-Graded Multilingual Stories Dataset
The MeloLingua CEFR-Graded Multilingual Stories Dataset is a citable educational corpus of 118 public A1–B2 language-learning stories in German, Spanish, French, Italian, Korean, and Russian. Records include target-language text, sentence-aligned English translations, contextual vocabulary, comprehension questions, sentence-building exercises, teaching metadata, provenance, and canonical links to original lessons on MeloLingua.… See the full description on the dataset page: https://huggingface.co/datasets/ismaelfi/melolingua-cefr-graded-multilingual-stories.swedish-cefr-text-complexity
Swedish CEFR Text Complexity Dataset
This dataset contains Swedish text examples labeled with approximate CEFR
reading levels from A1 to C2.
It was created for an information retrieval assignment about training text
classifiers with embeddings. The companion demo and classifier use
nicher92/saga-embed_v1 sentence embeddings and classical scikit-learn
classifiers.
The dataset is intended for Swedish text-complexity classification: given a
short Swedish sentence or paragraph, predict… See the full description on the dataset page: https://huggingface.co/datasets/kvest/swedish-cefr-text-complexity.CEFR_Mixed_Dataset_4090GPUCEFRLex
[!NOTE]
Dataset origin: https://pub.cl.uzh.ch/wiki/public/multiCEFRLex
Description
List of bilingual and trilingual matches with conditional translation probabilities.
Citation
@InProceedings{GraenAlfterSchneider2020,
author = {Gra\"{e}n, Johannes and Alfter, David and Schneider, Gerold},
title = {Using Multilingual Resources to Evaluate CEFRLex for Learner Applications},
booktitle = {Proceedings of the 12th Language Resources and… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/CEFRLex.welsh-cefrin-the-wild-jailbreak-prompts
In-The-Wild Jailbreak Prompts on LLMs
This is the official repository for the ACM CCS 2024 paper "Do Anything Now'': Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models by Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang.
In this project, employing our new framework JailbreakHub, we conduct the first measurement study on jailbreak prompts in the wild, with 15,140 prompts collected from December 2022 to December 2023 (including 1… See the full description on the dataset page: https://huggingface.co/datasets/Cefress/in-the-wild-jailbreak-prompts.test-cefrThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.ielts-cefr-5class-balancedCEFR_MIXED_dataset_60000CEFR_C2_Datasetcefr-cep-up-down-same-ABS-traincefr-lexical-balance-dataset-50-50-50
Dataset Card for "cefr-lexical-balance-dataset-50-50-50"
More Information needed
alpaca_format_CEFR_CEP_SS
Dataset Card for "alpaca_format_CEFR_CEP_SS"
More Information needed
elg_cefr_nlThis dataset has been indexed in the UniversalCEFR. The transformed version (in JSON format) retains the same license as the original dataset. Ownership and copyright remain with the original creators and/or dataset paper authors. If you use this transformed dataset, you must cite the following:
Dataset License: cc-by-nc-4.0
Dataset Repository: https://www.edia.nl/resources/elg/downloads
Original Dataset Paper: Breuker, M. (2023). CEFR Labelling and Assessment Services. In: Rehm, G. (eds)… See the full description on the dataset page: https://huggingface.co/datasets/UniversalCEFR/elg_cefr_nl.cefr-texts-10languages
Gold Standard CEFR Validation Dataset
Dataset Summary
This dataset is a high-quality synthetic validation set designed to evaluate models on CEFR (Common European Framework of Reference for Languages) Level Classification.
The dataset was generated using OpenAI's GPT-4o-mini. It contains approximately 3,000 examples balanced across 10 languages and 6 proficiency levels.
Dataset Structure
Data Fields
Each entry in the dataset consists of the… See the full description on the dataset page: https://huggingface.co/datasets/pinialt/cefr-texts-10languages.elg_cefr_deThis dataset has been indexed in the UniversalCEFR. The transformed version (in JSON format) retains the same license as the original dataset. Ownership and copyright remain with the original creators and/or dataset paper authors. If you use this transformed dataset, you must cite the following:
Dataset License: cc-by-nc-4.0
Dataset Repository: https://www.edia.nl/resources/elg/downloads
Original Dataset Paper: Breuker, M. (2023). CEFR Labelling and Assessment Services. In: Rehm, G. (eds)… See the full description on the dataset page: https://huggingface.co/datasets/UniversalCEFR/elg_cefr_de.
