datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ai2_arc-pt
marlosb/ai2-arc-pt
This dataset is a Portuguese translation of the original AI2 ARC (AI2 Reasoning Challenge) dataset released by AllenAI.
Original Dataset
Hugging Face: allenai/ai2_arc
Homepage: https://allenai.org/data/arc
Paper: Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Dataset Summary
AI2 ARC consists of 7,787 genuine grade-school level, multiple-choice science questions, assembled to encourage research in advanced… See the full description on the dataset page: https://huggingface.co/datasets/marlosb/ai2_arc-pt.mmlu-pt
marlosb/mmlu-pt
This dataset is a Portuguese translation of the original MMLU (Measuring Massive Multitask Language Understanding) dataset.
Original Dataset
Hugging Face: cais/mmlu
Repository: https://github.com/hendrycks/test
Paper: Measuring Massive Multitask Language Understanding
Dataset Summary
MMLU contains ~15,000 multiple-choice questions across 57 subjects (elementary mathematics, history, computer science, law, medicine, etc.), designed to evaluate… See the full description on the dataset page: https://huggingface.co/datasets/marlosb/mmlu-pt.
