datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PopMCQ
🎯 PopMCQ
Does your model pick the famous answer or the correct one?
📌 Overview
PopMCQ renders the same question six ways. The question and the correct answer never change — only how popular the three distractors are. That makes option popularity an independent variable, so an accuracy swing across S1–S6 is attributable to popularity rather than to question difficulty.
The swings are large. Under the hardest setting (S2), models pick a popular-but-wrong… See the full description on the dataset page: https://huggingface.co/datasets/DataScience-UIBK/PopMCQ.SOAR
SOAR: System Operations and Reasoning Benchmark
System Operations and Reasoning (SOAR) is a human-curated benchmark designed to evaluate the ability of large language models (LLMs) to interpret and answer questions over IT operational data, with a primary focus on Unix/Linux-style system administration contexts.It focuses on how well models can interpret, extract, and reason over operational data — including logs, configurations, and system outputs — and on how question… See the full description on the dataset page: https://huggingface.co/datasets/LAIA-UIB/SOAR.
