CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sli016 /Macarons-plus-plus0 likes108 downloads24d agoHugging Face02AlaaAhmed2444 /Macaron Macaron [Paper] Macaron is a controlled, human-written benchmark for multilingual and multicultural reasoning created with a template-first approach.Each example is scenario-aligned across English and a local language, enabling controlled comparison of reasoning under culturally grounded premises. At a glance Configuration Rows Description MCQ 1,977 Bilingual multiple-choice questions (English + local language) True-False 3,954 Bilingual verification… See the full description on the dataset page: https://huggingface.co/datasets/AlaaAhmed2444/Macaron.tabularquestion-answering1K<n<10K2 likes97 downloads7mo agoHugging Face03open-llm-leaderboard-old /details_Test157t__Hex-Macaroniac-7b Dataset Card for Evaluation run of Test157t/Hex-Macaroniac-7b Dataset automatically created during the evaluation run of model Test157t/Hex-Macaroniac-7b on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Test157t__Hex-Macaroniac-7b.0 likes45 downloads3y agoHugging Face04open-llm-leaderboard-old /details_andrijdavid__Macaroni-7b-Tied Dataset Card for Evaluation run of andrijdavid/Macaroni-7b-Tied Dataset automatically created during the evaluation run of model andrijdavid/Macaroni-7b-Tied on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_andrijdavid__Macaroni-7b-Tied.0 likes44 downloads3y agoHugging Face05open-llm-leaderboard-old /details_SanjiWatsuki__Loyal-Macaroni-Maid-7B Dataset Card for Evaluation run of SanjiWatsuki/Loyal-Macaroni-Maid-7B Dataset automatically created during the evaluation run of model SanjiWatsuki/Loyal-Macaroni-Maid-7B on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_SanjiWatsuki__Loyal-Macaroni-Maid-7B.0 likes42 downloads3y agoHugging Face06open-llm-leaderboard-old /details_andrijdavid__macaroni-7b Dataset Card for Evaluation run of andrijdavid/macaroni-7b Dataset automatically created during the evaluation run of model andrijdavid/macaroni-7b on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_andrijdavid__macaroni-7b.0 likes40 downloads3y agoHugging Face07open-llm-leaderboard-old /details_andrijdavid__Macaroni-v2-7b Dataset Card for Evaluation run of andrijdavid/Macaroni-v2-7b Dataset automatically created during the evaluation run of model andrijdavid/Macaroni-v2-7b on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_andrijdavid__Macaroni-v2-7b.0 likes37 downloads3y agoHugging Face08drooryck /multilingual-macaroni-corpus multilingual-macaroni training corpora The 8 training corpora behind our BabyLM 2026 Multilingual-track study of code-switched pretraining curricula (English / Dutch / Chinese). All are derived from the official BabyLM 2026 multilingual corpora (babylm-{eng,nld,zho}) — no external text — and held to the same 100M byte-premium-adjusted-word budget, split across the three languages. Each corpus trains one model condition in drooryck/multilingual-macaroni-models. Subset What… See the full description on the dataset page: https://huggingface.co/datasets/drooryck/multilingual-macaroni-corpus.text1M<n<10M0 likes29 downloads2mo agoHugging Face09open-llm-leaderboard-old /details_Vasanth__Valor_Macaroni_moe Dataset Card for Evaluation run of Vasanth/Valor_Macaroni_moe Dataset automatically created during the evaluation run of model Vasanth/Valor_Macaroni_moe on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Vasanth__Valor_Macaroni_moe.0 likes25 downloads3y agoHugging Face10drooryck /babylm-macaroni-corpus babylm-macaroni training corpus The training data for drooryck/babylm-macaroni, the BabyLM 2026 Multilingual-track submission. All corpora are created from the official BabyLM 2026 multilingual corpora (babylm-{eng,nld,zho}) and embedded with code-switching using an instruction-tuned LLM. The four corpora Folder Description shuffled_cs/ Full code-switched corpus, all documents globally shuffled. shuffled_nocs/ Matched unilingual twin (same documents… See the full description on the dataset page: https://huggingface.co/datasets/drooryck/babylm-macaroni-corpus.text100K<n<1M0 likes25 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.