lighteval
Test_Llama3-lighteval-finetune-mergedQwen_Qwen2.5-0.5B-Instruct-DigitalLearningGmbH_MATH-lighteval-RL_mergedQwen_Qwen2.5-0.5B-Instruct-DigitalLearningGmbH_MATH-lighteval-RL-fullllm-jp-3-13b-instruct2-grpo-MATH-lighteval_step1000Qwen2.5-1.5B-Instruct_MATH-lighteval_epoch_1_bs_8_lr_2e-05_length_2048_G_16Qwen2.5-7B-Open-R1-GRPO-math-lighteval-nonscale-weightedQwen2.5-7B-Open-R1-GRPO-math-lighteval-1epochstop-withformatQwen2.5-7B-Open-R1-GRPO-math-lighteval-v2
siqa
Dataset Card for "siqa"
More Information needed
piqa
Dataset Card for "Physical Interaction: Question Answering"
Dataset Summary
To apply eyeshadow without a brush, should I use a cotton swab or a toothpick?
Questions requiring this kind of physical commonsense pose a challenge to state-of-the-art
natural language understanding systems. The PIQA dataset introduces the task of physical commonsense reasoning
and a corresponding benchmark dataset Physical Interaction: Question Answering or PIQA.
Physical commonsense knowledge… See the full description on the dataset page: https://huggingface.co/datasets/lighteval/piqa.MATH-lighteval
Dataset Card for Mathematics Aptitude Test of Heuristics (MATH) dataset in lighteval format
Dataset Summary
The Mathematics Aptitude Test of Heuristics (MATH) dataset consists of problems
from mathematics competitions, including the AMC 10, AMC 12, AIME, and more.
Each problem in MATH has a full step-by-step solution, which can be used to teach
models to generate answer derivations and explanations. This version of the dataset
contains appropriate builder configs s.t. it… See the full description on the dataset page: https://huggingface.co/datasets/DigitalLearningGmbH/MATH-lighteval.mmlu
Dataset Card for MMLU
Dataset Summary
Measuring Massive Multitask Language Understanding by Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt (ICLR 2021).
This is a massive multitask test consisting of multiple-choice questions from various branches of knowledge. The test spans subjects in the humanities, social sciences, hard sciences, and other areas that are important for some people to learn. This covers 57 tasks… See the full description on the dataset page: https://huggingface.co/datasets/lighteval/mmlu.mutualbbh
