datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mt_bench_human_judgments
Content
This dataset contains 3.3K expert-level pairwise human preferences for model responses generated by 6 models in response to 80 MT-bench questions.
The 6 models are GPT-4, GPT-3.5, Claud-v1, Vicuna-13B, Alpaca-13B, and LLaMA-13B. The annotators are mostly graduate students with expertise in the topic areas of each of the questions. The details of data collection can be found in our paper.
Agreement Calculation
This Colab notebook shows how to compute the… See the full description on the dataset page: https://huggingface.co/datasets/lmsys/mt_bench_human_judgments.IPHO2026
IPhO 2026 Curated Problems
This repository packages the official English problem, solution, and marking
materials for the LVI International Physics Olympiad (Bucaramanga, Colombia,
2026) as machine-readable, subquestion-level records.
Contents
Configuration
Rows
Description
all
41
All curated subquestions
theory
23
Theory papers T1–T3
experiment
18
Experimental paper E1
formalization_ready
29
Subset selected for theorem formalization… See the full description on the dataset page: https://huggingface.co/datasets/humanfia-lab/IPHO2026.Corp_FinanceLongContextReasoning
Corp_Finance Long-Context Reasoning — Expert-Authored Credit Agreement Benchmark (Showcase Sample)
A five-record public sample from Corp_Finance Long-Context Reasoning, a subject-matter-expert benchmarking dataset built by Human Edge for evaluating frontier model reasoning over leveraged finance and syndicated credit documentation.
Every question, reasoning trace, and answer in this dataset was authored by a practicing finance professional and independently reviewed by 2-3… See the full description on the dataset page: https://huggingface.co/datasets/HumanEdgeAI/Corp_FinanceLongContextReasoning.MoA_Long_HumanQA
MoA: Mixture of Sparse Attention for Automatic Large Language Model Compression
This is the dataset used by the automatic sparse attention compression method MoA.
It enhances the calibration dataset by integrating long-range dependencies and model alignment.
MoA utilizes long-contextual datasets, which include question-answer pairs heavily dependent on long-range content.
The question-answer pairs are written by human in this dataset repository. Large language Models (LLMs) should… See the full description on the dataset page: https://huggingface.co/datasets/nics-efc/MoA_Long_HumanQA.LegalReasoning
Human Edge — Legal Reasoning Evaluation (Showcase Sample)
A public 6-task sample from a rubric-based legal reasoning evaluation dataset built by
Human Edge (humanedgetech.ai). Each task is authored and reviewed by
practicing senior lawyers and is designed to produce a verifiable, per-criterion reward signal
for post-training and evaluation of frontier language models on high-complexity legal work.
The sample contains one task per legal subdomain, drawn from a larger internal… See the full description on the dataset page: https://huggingface.co/datasets/HumanEdgeAI/LegalReasoning.mind2web-subset-human
Mind2Web Subset - Human Demonstrations
A collection of human-demonstrated web navigation tasks with detailed interaction traces. This dataset captures real browser interactions including clicks, typing, scrolling, DOM states, screenshots, and HTTP requests for web agent training and evaluation.
Overview
This dataset contains tasks performed by humans in real web environments, capturing:
Golden trajectories: Step-by-step sequences of actions (clicks, typing, navigation)… See the full description on the dataset page: https://huggingface.co/datasets/josancamon/mind2web-subset-human.MedicalReasoning
Project Aletheia — Expert-Grounded Medical QA Evaluation (6-Case Illustrative Sample)
A 6-record sample drawn from the complete Project Aletheia dataset — a 52-case, subject-matter-expert evaluation pilot built by Human Edge to ground medical question-answering preference data in licensed clinical judgment.
Every gold-standard response, segment agreement count, pairwise preference, and reasoning trace was authored or adjudicated by a licensed medical doctor, through a pipeline… See the full description on the dataset page: https://huggingface.co/datasets/HumanEdgeAI/MedicalReasoning.HumanAgencyBench_results
Full HumanAgencyBench results using GPT 4.1 generated prompts and o3 as evaluator
Dataset Description
This dataset contains comprehensive evaluation results from testing 25 different language models across 6 areas of behaviours critical for human agency support. Each model was evaluated on 3,000 prompts (500 per category), resulting in 75,000 total evaluations designed to assess model behavior in scenarios relevant to human agency.
Dataset Summary
Total… See the full description on the dataset page: https://huggingface.co/datasets/ferocious-aardvark/HumanAgencyBench_results.mt_bench_human_judgments
Content
This dataset contains 3.3K expert-level pairwise human preferences for model responses generated by 6 models in response to 80 MT-bench questions.
The 6 models are GPT-4, GPT-3.5, Claud-v1, Vicuna-13B, Alpaca-13B, and LLaMA-13B. The annotators are mostly graduate students with expertise in the topic areas of each of the questions. The details of data collection can be found in our paper.
Agreement Calculation
This Colab notebook shows how to compute the… See the full description on the dataset page: https://huggingface.co/datasets/yzygalaxy/mt_bench_human_judgments.mt_bench_human_judgments
Content
This dataset contains 3.3K expert-level pairwise human preferences for model responses generated by 6 models in response to 80 MT-bench questions.
The 6 models are GPT-4, GPT-3.5, Claud-v1, Vicuna-13B, Alpaca-13B, and LLaMA-13B. The annotators are mostly graduate students with expertise in the topic areas of each of the questions. The details of data collection can be found in our paper.
Agreement Calculation
This Colab notebook shows how to compute the… See the full description on the dataset page: https://huggingface.co/datasets/shehbaz0101/mt_bench_human_judgments.HumanAgencyEval_results
Full HumanAgencyEval results using GPT 4.1 generated prompts and o3 as evaluator
Dataset Description
This dataset contains comprehensive evaluation results from testing 25 different language models across 6 areas of behaviours critical for human agency support. Each model was evaluated on 3,000 prompts (500 per category), resulting in 75,000 total evaluations designed to assess model behavior in scenarios relevant to human agency.
Dataset Summary
Total… See the full description on the dataset page: https://huggingface.co/datasets/anon34957/HumanAgencyEval_results.
