CoolFace
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lmsys /mt_bench_human_judgments Content This dataset contains 3.3K expert-level pairwise human preferences for model responses generated by 6 models in response to 80 MT-bench questions. The 6 models are GPT-4, GPT-3.5, Claud-v1, Vicuna-13B, Alpaca-13B, and LLaMA-13B. The annotators are mostly graduate students with expertise in the topic areas of each of the questions. The details of data collection can be found in our paper. Agreement Calculation This Colab notebook shows how to compute the… See the full description on the dataset page: https://huggingface.co/datasets/lmsys/mt_bench_human_judgments.tabularquestion-answering1K<n<10K148 likes2k downloads3y agoHugging Face02humanfia-lab /IPHO2026 IPhO 2026 Curated Problems This repository packages the official English problem, solution, and marking materials for the LVI International Physics Olympiad (Bucaramanga, Colombia, 2026) as machine-readable, subquestion-level records. Contents Configuration Rows Description all 41 All curated subquestions theory 23 Theory papers T1–T3 experiment 18 Experimental paper E1 formalization_ready 29 Subset selected for theorem formalization… See the full description on the dataset page: https://huggingface.co/datasets/humanfia-lab/IPHO2026.imagequestion-answeringn<1K0 likes204 downloads2mo agoHugging Face03HumanEdgeAI /Corp_FinanceLongContextReasoning Corp_Finance Long-Context Reasoning — Expert-Authored Credit Agreement Benchmark (Showcase Sample) A five-record public sample from Corp_Finance Long-Context Reasoning, a subject-matter-expert benchmarking dataset built by Human Edge for evaluating frontier model reasoning over leveraged finance and syndicated credit documentation. Every question, reasoning trace, and answer in this dataset was authored by a practicing finance professional and independently reviewed by 2-3… See the full description on the dataset page: https://huggingface.co/datasets/HumanEdgeAI/Corp_FinanceLongContextReasoning.tabularquestion-answeringn<1K0 likes62 downloads20d agoHugging Face04nics-efc /MoA_Long_HumanQA MoA: Mixture of Sparse Attention for Automatic Large Language Model Compression This is the dataset used by the automatic sparse attention compression method MoA. It enhances the calibration dataset by integrating long-range dependencies and model alignment. MoA utilizes long-contextual datasets, which include question-answer pairs heavily dependent on long-range content. The question-answer pairs are written by human in this dataset repository. Large language Models (LLMs) should… See the full description on the dataset page: https://huggingface.co/datasets/nics-efc/MoA_Long_HumanQA.tabularquestion-answering1K<n<10K4 likes47 downloads2y agoHugging Face05HumanEdgeAI /LegalReasoning Human Edge — Legal Reasoning Evaluation (Showcase Sample) A public 6-task sample from a rubric-based legal reasoning evaluation dataset built by Human Edge (humanedgetech.ai). Each task is authored and reviewed by practicing senior lawyers and is designed to produce a verifiable, per-criterion reward signal for post-training and evaluation of frontier language models on high-complexity legal work. The sample contains one task per legal subdomain, drawn from a larger internal… See the full description on the dataset page: https://huggingface.co/datasets/HumanEdgeAI/LegalReasoning.tabulartext-generationn<1K0 likes32 downloads2mo agoHugging Face06josancamon /mind2web-subset-human Mind2Web Subset - Human Demonstrations A collection of human-demonstrated web navigation tasks with detailed interaction traces. This dataset captures real browser interactions including clicks, typing, scrolling, DOM states, screenshots, and HTTP requests for web agent training and evaluation. Overview This dataset contains tasks performed by humans in real web environments, capturing: Golden trajectories: Step-by-step sequences of actions (clicks, typing, navigation)… See the full description on the dataset page: https://huggingface.co/datasets/josancamon/mind2web-subset-human.tabularreinforcement-learningn<1K0 likes24 downloads1y agoHugging Face07HumanEdgeAI /MedicalReasoning Project Aletheia — Expert-Grounded Medical QA Evaluation (6-Case Illustrative Sample) A 6-record sample drawn from the complete Project Aletheia dataset — a 52-case, subject-matter-expert evaluation pilot built by Human Edge to ground medical question-answering preference data in licensed clinical judgment. Every gold-standard response, segment agreement count, pairwise preference, and reasoning trace was authored or adjudicated by a licensed medical doctor, through a pipeline… See the full description on the dataset page: https://huggingface.co/datasets/HumanEdgeAI/MedicalReasoning.tabularquestion-answeringn<1K0 likes21 downloads2mo agoHugging Face08ferocious-aardvark /HumanAgencyBench_results Full HumanAgencyBench results using GPT 4.1 generated prompts and o3 as evaluator Dataset Description This dataset contains comprehensive evaluation results from testing 25 different language models across 6 areas of behaviours critical for human agency support. Each model was evaluated on 3,000 prompts (500 per category), resulting in 75,000 total evaluations designed to assess model behavior in scenarios relevant to human agency. Dataset Summary Total… See the full description on the dataset page: https://huggingface.co/datasets/ferocious-aardvark/HumanAgencyBench_results.tabularquestion-answering10K<n<100K0 likes12 downloads1y agoHugging Face09yzygalaxy /mt_bench_human_judgments Content This dataset contains 3.3K expert-level pairwise human preferences for model responses generated by 6 models in response to 80 MT-bench questions. The 6 models are GPT-4, GPT-3.5, Claud-v1, Vicuna-13B, Alpaca-13B, and LLaMA-13B. The annotators are mostly graduate students with expertise in the topic areas of each of the questions. The details of data collection can be found in our paper. Agreement Calculation This Colab notebook shows how to compute the… See the full description on the dataset page: https://huggingface.co/datasets/yzygalaxy/mt_bench_human_judgments.tabularquestion-answering1K<n<10K0 likes12 downloads6mo agoHugging Face10shehbaz0101 /mt_bench_human_judgments Content This dataset contains 3.3K expert-level pairwise human preferences for model responses generated by 6 models in response to 80 MT-bench questions. The 6 models are GPT-4, GPT-3.5, Claud-v1, Vicuna-13B, Alpaca-13B, and LLaMA-13B. The annotators are mostly graduate students with expertise in the topic areas of each of the questions. The details of data collection can be found in our paper. Agreement Calculation This Colab notebook shows how to compute the… See the full description on the dataset page: https://huggingface.co/datasets/shehbaz0101/mt_bench_human_judgments.tabularquestion-answering1K<n<10K0 likes10 downloads2mo agoHugging Face11anon34957 /HumanAgencyEval_results Full HumanAgencyEval results using GPT 4.1 generated prompts and o3 as evaluator Dataset Description This dataset contains comprehensive evaluation results from testing 25 different language models across 6 areas of behaviours critical for human agency support. Each model was evaluated on 3,000 prompts (500 per category), resulting in 75,000 total evaluations designed to assess model behavior in scenarios relevant to human agency. Dataset Summary Total… See the full description on the dataset page: https://huggingface.co/datasets/anon34957/HumanAgencyEval_results.tabularquestion-answering10K<n<100K0 likes4 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.