datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
french_CEFRenglish_cefr_datasetswedish-cefr-text-complexity
Swedish CEFR Text Complexity Dataset
This dataset contains Swedish text examples labeled with approximate CEFR
reading levels from A1 to C2.
It was created for an information retrieval assignment about training text
classifiers with embeddings. The companion demo and classifier use
nicher92/saga-embed_v1 sentence embeddings and classical scikit-learn
classifiers.
The dataset is intended for Swedish text-complexity classification: given a
short Swedish sentence or paragraph, predict… See the full description on the dataset page: https://huggingface.co/datasets/kvest/swedish-cefr-text-complexity.cefr-texts-10languages
Gold Standard CEFR Validation Dataset
Dataset Summary
This dataset is a high-quality synthetic validation set designed to evaluate models on CEFR (Common European Framework of Reference for Languages) Level Classification.
The dataset was generated using OpenAI's GPT-4o-mini. It contains approximately 3,000 examples balanced across 10 languages and 6 proficiency levels.
Dataset Structure
Data Fields
Each entry in the dataset consists of the… See the full description on the dataset page: https://huggingface.co/datasets/pinialt/cefr-texts-10languages.CEFR_expertCEFR_CER_expertSynth_CEFRSynth_gold_CEFRCEFR_mixtoenglish_cefr_datasetLLama2_expert_CEFRcefrdatasetbygbtcefr_pairs
