datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
melolingua-cefr-graded-multilingual-stories
MeloLingua CEFR-Graded Multilingual Stories Dataset
The MeloLingua CEFR-Graded Multilingual Stories Dataset is a citable educational corpus of 118 public A1–B2 language-learning stories in German, Spanish, French, Italian, Korean, and Russian. Records include target-language text, sentence-aligned English translations, contextual vocabulary, comprehension questions, sentence-building exercises, teaching metadata, provenance, and canonical links to original lessons on MeloLingua.… See the full description on the dataset page: https://huggingface.co/datasets/ismaelfi/melolingua-cefr-graded-multilingual-stories.in-the-wild-jailbreak-prompts
In-The-Wild Jailbreak Prompts on LLMs
This is the official repository for the ACM CCS 2024 paper "Do Anything Now'': Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models by Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang.
In this project, employing our new framework JailbreakHub, we conduct the first measurement study on jailbreak prompts in the wild, with 15,140 prompts collected from December 2022 to December 2023 (including 1… See the full description on the dataset page: https://huggingface.co/datasets/Cefress/in-the-wild-jailbreak-prompts.cefr-english-teaching-corpus
The English Classroom Corpus
A 600-hour, CEFR-graded English language teaching corpus (A1–C2) with aligned native-speaker audio. Single author. Clean rights. Available under commercial licence.
⚠️ The data is not hosted in this repository. This is a dataset card for discovery purposes. Licensing enquiries: the-english-classroom.com/licensing or jennifer@the-englishclassroom.com
Dataset summary
The English Classroom Corpus is a complete English language… See the full description on the dataset page: https://huggingface.co/datasets/jennifertec/cefr-english-teaching-corpus.English-CEFR-Explorer
🇬🇧 English CEFR Explorer Benchmark
This dataset is an automatically updated benchmark for testing LLM adherence to CEFR (Common European Framework of Reference for Languages) constraints.
Dataset Structure
The dataset is formatted as a JSONL file optimized for Instruction Tuning.
Data Fields
id: Unique identifier for the generation task.
messages: Standard chat format (System, User, Assistant).
System: Defines the persona (ESL Teacher).
User: The… See the full description on the dataset page: https://huggingface.co/datasets/yasincicek/English-CEFR-Explorer.
