datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
merlin
MERLIN corpus
Project URL: https://merlin-platform.eu/C_mcorpus.php
Dataset URL: https://clarin.eurac.edu/repository/xmlui/handle/20.500.12124/6
The MERLIN corpus is a written learner corpus for Czech, German, and Italian that has been designed to illustrate the Common European Framework of Reference for Languages (CEFR) with authentic learner data. The corpus contains learner texts produced in standardized language certifications covering CEFR levels A1-C1. The MERLIN annotation… See the full description on the dataset page: https://huggingface.co/datasets/aseifert/merlin.merlin_deThis dataset has been indexed in the UniversalCEFR. The transformed version (in JSON format) retains the same license as the original dataset. Ownership and copyright remain with the original creators and/or dataset paper authors. If you use this transformed dataset, you must cite the following:
Dataset License: cc-by-sa-4.0
Dataset Repository: https://www.merlin-platform.eu/C_mcorpus.php
Original Dataset Paper: Adriane Boyd, Jirka Hana, Lionel Nicolas, Detmar Meurers, Katrin Wisniewski… See the full description on the dataset page: https://huggingface.co/datasets/UniversalCEFR/merlin_de.ibm__merlinite-7b-details
Dataset Card for Evaluation run of ibm/merlinite-7b
Dataset automatically created during the evaluation run of model ibm/merlinite-7b
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ibm__merlinite-7b-details.merlin_itThis dataset has been indexed in the UniversalCEFR. The transformed version (in JSON format) retains the same license as the original dataset. Ownership and copyright remain with the original creators and/or dataset paper authors. If you use this transformed dataset, you must cite the following:
Dataset License: cc-by-sa-4.0
Dataset Repository: https://www.merlin-platform.eu/C_mcorpus.php
Original Dataset Paper: Adriane Boyd, Jirka Hana, Lionel Nicolas, Detmar Meurers, Katrin Wisniewski… See the full description on the dataset page: https://huggingface.co/datasets/UniversalCEFR/merlin_it.merlinmerlin_csThis dataset has been indexed in the UniversalCEFR. The transformed version (in JSON format) retains the same license as the original dataset. Ownership and copyright remain with the original creators and/or dataset paper authors. If you use this transformed dataset, you must cite the following:
Dataset License: cc-by-sa-4.0
Dataset Repository: https://www.merlin-platform.eu/C_mcorpus.php
Original Dataset Paper: Adriane Boyd, Jirka Hana, Lionel Nicolas, Detmar Meurers, Katrin Wisniewski… See the full description on the dataset page: https://huggingface.co/datasets/UniversalCEFR/merlin_cs.EuroAlign-1K
EuroAlign-1K
First systematic multilingual AI safety evaluation dataset covering 10 EU languages.
EuroAlign-1K measures alignment gaps in large language models across Central Eastern European and Nordic EU languages — a compliance concern under EU AI Act Article 14, which requires equal AI performance across all EU language groups.
Dataset Summary
Stat
Value
Total prompts
3,300
Languages
10
Prompts per language
330 (162 adversarial + 168… See the full description on the dataset page: https://huggingface.co/datasets/Merlin-Research/EuroAlign-1K.Merlin-v0.1Merlin-FR-v0.1idafalko_merlin_wikimerlin-intent-training-datamerlin_fewshotsa_new_full
