datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Macaron
Macaron
[Paper]
Macaron is a controlled, human-written benchmark for multilingual and multicultural reasoning created with a template-first approach.Each example is scenario-aligned across English and a local language, enabling controlled comparison of reasoning under culturally grounded premises.
At a glance
Configuration
Rows
Description
MCQ
1,977
Bilingual multiple-choice questions (English + local language)
True-False
3,954
Bilingual verification… See the full description on the dataset page: https://huggingface.co/datasets/AlaaAhmed2444/Macaron.babylm-macaroni-corpus
babylm-macaroni training corpus
The training data for drooryck/babylm-macaroni,
the BabyLM 2026 Multilingual-track submission. All corpora are created from the official BabyLM 2026
multilingual corpora (babylm-{eng,nld,zho}) and embedded with code-switching using an instruction-tuned LLM.
The four corpora
Folder
Description
shuffled_cs/
Full code-switched corpus, all documents globally shuffled.
shuffled_nocs/
Matched unilingual twin (same documents… See the full description on the dataset page: https://huggingface.co/datasets/drooryck/babylm-macaroni-corpus.multilingual-macaroni-corpus
multilingual-macaroni training corpora
The 8 training corpora behind our BabyLM 2026 Multilingual-track study of code-switched
pretraining curricula (English / Dutch / Chinese). All are derived from the official BabyLM 2026
multilingual corpora (babylm-{eng,nld,zho}) — no external text — and held to the same 100M
byte-premium-adjusted-word budget, split across the three languages. Each corpus trains one model
condition in drooryck/multilingual-macaroni-models.
Subset
What… See the full description on the dataset page: https://huggingface.co/datasets/drooryck/multilingual-macaroni-corpus.
