datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CEFR-Annotated-WordNet
CEFR-Annotated WordNet: LLM-Based Proficiency-Guided Semantic Database for Language Learning
Paper: CEFR-Annotated WordNet: LLM-Based Proficiency-Guided Semantic Database for Language Learning
Authors: Masato Kikuchi, Masatsugu Ono, Toshioki Soga, Tetsu Tanabe, Tadachika Ozono
Overview
CEFR-Annotated WordNet is a comprehensive semantic database that maps WordNet 3.0 senses and SemCor 3.0 instances to CEFR (Common European Framework of Reference for Languages)… See the full description on the dataset page: https://huggingface.co/datasets/star092304/CEFR-Annotated-WordNet.cefr_sp_enThis dataset has been indexed in the UniversalCEFR. The transformed version (in JSON format) retains the same license as the original dataset. Ownership and copyright remain with the original creators and/or dataset paper authors. If you use this transformed dataset, you must cite the following:
Dataset License: cc-by-nc-sa-4.0
Dataset Repository: https://github.com/yukiar/CEFR-SP/tree/main/CEFR-SP
Original Dataset Paper: Yuki Arase, Satoru Uchida, and Tomoyuki Kajiwara. 2022. CEFR-Based… See the full description on the dataset page: https://huggingface.co/datasets/UniversalCEFR/cefr_sp_en.elg_cefr_enThis dataset has been indexed in the UniversalCEFR. The transformed version (in JSON format) retains the same license as the original dataset. Ownership and copyright remain with the original creators and/or dataset paper authors. If you use this transformed dataset, you must cite the following:
Dataset License: cc-by-nc-4.0
Dataset Repository: https://www.edia.nl/resources/elg/downloads
Original Dataset Paper: Breuker, M. (2023). CEFR Labelling and Assessment Services. In: Rehm, G. (eds)… See the full description on the dataset page: https://huggingface.co/datasets/UniversalCEFR/elg_cefr_en.cefr_asag_enThis dataset has been indexed in the UniversalCEFR. The transformed version (in JSON format) retains the same license as the original dataset. Ownership and copyright remain with the original creators and/or dataset paper authors. If you use this transformed dataset, you must cite the following:
Dataset License: cc-by-nc-sa-4.0
Dataset Repository: https://github.com/anaistack/cefr-asag-corpus
Original Dataset Paper: Anaïs Tack, Thomas François, Sophie Roekhaut, and Cédrick Fairon. 2017. Human… See the full description on the dataset page: https://huggingface.co/datasets/UniversalCEFR/cefr_asag_en.elg_cefr_nlThis dataset has been indexed in the UniversalCEFR. The transformed version (in JSON format) retains the same license as the original dataset. Ownership and copyright remain with the original creators and/or dataset paper authors. If you use this transformed dataset, you must cite the following:
Dataset License: cc-by-nc-4.0
Dataset Repository: https://www.edia.nl/resources/elg/downloads
Original Dataset Paper: Breuker, M. (2023). CEFR Labelling and Assessment Services. In: Rehm, G. (eds)… See the full description on the dataset page: https://huggingface.co/datasets/UniversalCEFR/elg_cefr_nl.elg_cefr_deThis dataset has been indexed in the UniversalCEFR. The transformed version (in JSON format) retains the same license as the original dataset. Ownership and copyright remain with the original creators and/or dataset paper authors. If you use this transformed dataset, you must cite the following:
Dataset License: cc-by-nc-4.0
Dataset Repository: https://www.edia.nl/resources/elg/downloads
Original Dataset Paper: Breuker, M. (2023). CEFR Labelling and Assessment Services. In: Rehm, G. (eds)… See the full description on the dataset page: https://huggingface.co/datasets/UniversalCEFR/elg_cefr_de.german-cefrFor further information, see the associated paper Link.
If you use this dataset in your research, please use the following citation:
@misc{ahlers2025classifyinggermanlanguageproficiency,
title={Classifying German Language Proficiency Levels Using Large Language Models},
author={Elias-Leander Ahlers and Malte Schilling},
year={2025},
eprint={2512.06483},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2512.06483},
}
English-CEFR-Explorer
🇬🇧 English CEFR Explorer Benchmark
This dataset is an automatically updated benchmark for testing LLM adherence to CEFR (Common European Framework of Reference for Languages) constraints.
Dataset Structure
The dataset is formatted as a JSONL file optimized for Instruction Tuning.
Data Fields
id: Unique identifier for the generation task.
messages: Standard chat format (System, User, Assistant).
System: Defines the persona (ESL Teacher).
User: The… See the full description on the dataset page: https://huggingface.co/datasets/yasincicek/English-CEFR-Explorer.BiLit-CEFR
Balanced Bilingual Literary Complexity Dataset (EN + ZH)
Description
This dataset contains paragraph-level literary segments labeled with CEFR-style difficulty levels (A/B/C).
Languages:
English (CommonLit-based scoring)
Chinese (heuristic labeling)
Data Fields
text: paragraph
difficulty: A/B/C
language: en/zh
open_source_books: original source
Labeling Method
English: CommonLit readability model,syntactic complexity, the proportion of rare… See the full description on the dataset page: https://huggingface.co/datasets/merriamtao/BiLit-CEFR.turkish-cefr-phrases
Turkish CEFR Phrases
A dataset of Turkish phrases extracted from YouTube videos, labeled with CEFR proficiency levels (A1–C2).
Dataset Details
Language: Turkish
Size: 174,526 phrases
Labels: A1, A2, B1, B2, C1, C2
Label Distribution
Level
Count
%
A2
70,626
40.5%
B1
57,758
33.1%
B2
34,728
19.9%
A1
7,016
4.0%
C1
4,304
2.5%
C2
94
0.05%
Data Fields
phrase_text — Turkish phrase
cefr_level — CEFR level label (A1 to C2)… See the full description on the dataset page: https://huggingface.co/datasets/crklih/turkish-cefr-phrases.cefr_sp_enThis dataset has been indexed in the UniversalCEFR. The transformed version (in JSON format) retains the same license as the original dataset. Ownership and copyright remain with the original creators and/or dataset paper authors. If you use this transformed dataset, you must cite the following:
Dataset License: cc-by-nc-sa-4.0
Dataset Repository: https://github.com/yukiar/CEFR-SP/tree/main/CEFR-SP
Original Dataset Paper: Yuki Arase, Satoru Uchida, and Tomoyuki Kajiwara. 2022. CEFR-Based… See the full description on the dataset page: https://huggingface.co/datasets/Ozodbek2/cefr_sp_en.North-Welsh-CEFR
