10languages
10-languages-samplesThis dataset is suitable for evaluating language-identificatioan model
cefr-texts-10languages
Gold Standard CEFR Validation Dataset
Dataset Summary
This dataset is a high-quality synthetic validation set designed to evaluate models on CEFR (Common European Framework of Reference for Languages) Level Classification.
The dataset was generated using OpenAI's GPT-4o-mini. It contains approximately 3,000 examples balanced across 10 languages and 6 proficiency levels.
Dataset Structure
Data Fields
Each entry in the dataset consists of the… See the full description on the dataset page: https://huggingface.co/datasets/pinialt/cefr-texts-10languages.99sentences-translate2-10languagesThis dataset includes 99 English sentences and translations to Arabic, Greek, Spanish, Italian, Indonesian, Urdu, Hindi, Korean, Vietnamese and Chinese
