datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ClimaQA
ClimaQA: An Automated Evaluation Framework for Climate Question Answering Models (ICLR 2025)
Check the paper's webpage and GitHub for more info!
The ClimaQA benchmark is designed to evaluate Large Language Models (LLMs) on climate science question-answering tasks by ensuring scientific rigor and complexity. It is built from graduate-level climate science textbooks, which provide a reliable foundation for generating questions with precise terminology and complex scientific theories.… See the full description on the dataset page: https://huggingface.co/datasets/Rose-STL-Lab/ClimaQA.agentic-reasoning-benchmark
Agentic & Reasoning Benchmark (ARB) – Expanded
Ein synthetischer Benchmark mit 2.550 Fragen und Lösungen, optimiert für die Evaluation von Agentic Capabilities und Reasoning.
Überblick
Eigenschaft
Wert
Anzahl Beispiele
2.550
Kategorien
8
Schwierigkeitsgrade
easy / medium / hard
Formate
CSV + JSON
Reproduzierbarkeit
Generator-Skript (seed=42) enthalten
Lizenz
CC-BY-4.0
Kategorien
Kategorie
Anzahl
Beschreibung… See the full description on the dataset page: https://huggingface.co/datasets/roskosmos19/agentic-reasoning-benchmark.Legal_advice_czech
Legal Advice Dataset
Dataset Description
This dataset contains scraped legal questions (but also a simply informational content without any question present) and answers from Bezplatná Právní Poradna. The data consists of legal inquiries submitted by users and expert responses provided on the website. It is structured for ease of use in natural language processing (NLP) tasks related to legal text classification, question-answering models, and text summarization. However… See the full description on the dataset page: https://huggingface.co/datasets/roslein/Legal_advice_czech.MMLU-Phrasing-Benchmark
MMLU Phrasing Benchmark
This dataset is a phrasing variant of cais/mmlu, put together by Roscommon Systems to see whether the way a question is worded affects how accurately language models answer it.
Each of the 2,650 questions appears four ways: the original text from MMLU, a polite version, a formal academic version, and an angry/demanding version. The answer choices and correct answers are identical to the source dataset in all cases.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/RoscommonSystems/MMLU-Phrasing-Benchmark.
