datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
UnsolvedMath🌐 Browse UnsolvedMath online
✅ Paper: Open Mathematical Problems as an AI Reasoning Benchmark
UnsolvedMath Dataset
A comprehensive curated collection of 15,458 open and partially solved mathematics problems across all domains and difficulty levels, including the largest collection of Erdős problems available in machine-readable format. Available for browsing at unsolvedmath.com.
Paper: "Open Mathematical Problems as an AI Reasoning Benchmark"
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/ulamai/UnsolvedMath.Agentic-Long-Context-Understanding-QA 📖 Agentic Long Context Understanding 📖
Self-Taught Agentic Long Context Understanding (Arxiv).
AgenticLU refines complex, long-context queries through self-clarifications and contextual grounding, enabling robust long-document understanding in a single pass.
Installation Requirements
This codebase is largely based on OpenRLHF and Helmet, kudos to them.
The requirements are the same
pip install openrlhf
pip install -r ./HELMET/requirements.txt… See the full description on the dataset page: https://huggingface.co/datasets/yzhuang/Agentic-Long-Context-Understanding-QA.Multi-Turn-Insurance-Underwriting
Dataset Card for Multi-Turn-Insurance-Underwriting
Dataset Summary
This dataset includes sample traces and associated metadata from multi-turn interactions between a commercial underwriter and AI assistant. We built the system in langgraph with model context protocol and ReAct agents. In each sample, the underwriter has a specific task to solve related to a recent application for insurance by a small business. We created a diverse sample dataset covering 6 distinct types… See the full description on the dataset page: https://huggingface.co/datasets/snorkelai/Multi-Turn-Insurance-Underwriting.DoW-UFO-UAP-1
Department of War UFO/UAP Release 01 OCR + Metadata
This repository is intended as the canonical machine-readable Hugging Face dataset for public Department of War / PURSUE UFO-UAP Release 01 records.
It uses one dataset repo with internal sharding, not one repo per source file. Users can load only the table they need via named configs: pages, packets, sources, classification_markings, triage, or media_assets.
Agent-native access
This dataset ships with an importable… See the full description on the dataset page: https://huggingface.co/datasets/unmodeled-tyler/DoW-UFO-UAP-1.Reverse-baseline-bias-unbiasorc-bench
ORC-bench
Task 1: Topological Path Finding
Task 2: Topological Connectivity
Task 3: Linear Power Flow
Task 4: Contingency Analysis
Task 5: Power Grid ControlTask 6: Power Flow Optimization
Task 1: Topological Path Finding
Problem Formulation
This task assesses the spatial reasoning ability of the model by asking it to determine the shortest path between two specific buses in a given power grid state. The grid state… See the full description on the dataset page: https://huggingface.co/datasets/serval-uni-lu/orc-bench.un-digital-library
United Nations Digital Library (UNDL) Comprehensive Master Dataset
1. Executive Summary
Welcome to the United Nations Digital Library (UNDL) Comprehensive Master Dataset repository. This dataset represents a monumental effort to harvest, normalize, enrich, and democratize access to the vast archives of the United Nations. By leveraging advanced web harvesting techniques, robust state management, and modern big-data formats, this repository provides researchers… See the full description on the dataset page: https://huggingface.co/datasets/AdhyanshVerma/un-digital-library.Original-baseline-bias-unbiasmedquad-retrieval-pretriage
MedQuAD Retrieval Pre-Triage Dataset
Dataset Description
This repository contains a processed, retrieval-oriented derivative of the MedQuAD medical question-answering dataset.
It was prepared for contextual medical information retrieval in SortMed, an academic medical pre-triage assistant.
The corpus is not used to train the SortMed triage classifiers. It is used by a separate semantic retrieval component that identifies medically related question-answer entries… See the full description on the dataset page: https://huggingface.co/datasets/cristian-untaru/medquad-retrieval-pretriage.BEEP_eval
🚗 BEst DrivEr’s License Performer (BEEP) Dataset
BEEP is a challenge benchmark designed to evaluate large language models (LLMs) through a simulation of the Italian driver’s license exam. This dataset focuses on understanding traffic laws and reasoning through driving situations, replicating the complexity of the Italian licensing process.
📁 Dataset Structure
Column
Data Type
Description
Categorisation Structure
[String]
Hierarchical categorisation of major… See the full description on the dataset page: https://huggingface.co/datasets/Crisp-Unimib/BEEP_eval.NeurIPS26_Precise-but-Uncoupled
Precise but Uncoupled
Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning
Accepted to NeurIPS 2026 — Main Track
Protocol traces, process metrics and derived tables for the paper. Reviewer
detection quality and successful critique uptake are empirically separable:
a multi-agent protocol can identify errors accurately and still fail to change
the answer it carries forward.
Resource
Link
📄 Paper
arXiv:2607.15388 · PDF
🌐… See the full description on the dataset page: https://huggingface.co/datasets/AgentsSci/NeurIPS26_Precise-but-Uncoupled.openm3chest-labels
OpenM3Chest Labels (OM3C)
JSON label files and Series UIDs from the OpenM3Chest dataset, prepared for fine-tuning medical vision-language models such as MedGemma.
Raw imaging data (DICOM) can be downloaded from IDC (Imaging Data Commons) using the Series Instance UIDs provided in unique_keys.txt.
Dataset Summary
OpenM3Chest is a medical multimodal multitask dataset for diagnosing chest abnormalities with a focus on lung cancer screening. The original raw data comes… See the full description on the dataset page: https://huggingface.co/datasets/UngLong/openm3chest-labels.unsolved-math-clean
🧠 Unsolved Math — Clean
8,626 curated open research problems in mathematics and CS — including 122 Millennium Prize Problems — deduplicated, schema-flattened, and packaged as proper parquet configs with an eval-only benchmark view.
A reasoning frontier dataset: every problem here is actually unsolved or partially solved — ideal for honest capability probing instead of contaminated benchmarks.
Clean derivative of ulamai/UnsolvedMath (8,785 problems). License unchanged:… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/unsolved-math-clean.mmlu_hinted_questions
MMLU Hinted Questions
Dataset Description
This dataset contains multiple-choice questions derived from MMLU and augmented with misleading hints. The misleading hints are intentionally designed to point to an incorrect answer.
The dataset was developed as part of the UnfaithRL project, which studies cue-following and unfaithful reasoning under reinforcement learning with verifiable rewards.
Specifically, it was used to investigate whether language models follow… See the full description on the dataset page: https://huggingface.co/datasets/UnfaithRL/mmlu_hinted_questions.aletheia_code_problems
Aletheia Code Problems with Misleading Hints
Dataset Description
This dataset contains multiple-choice code-reasoning problems derived from Aletheia-Bench and augmented with misleading textual hints. The misleading hints are intentionally designed to point to an incorrect answer.
The dataset was developed as part of the UnfaithRL project, which studies cue-following and unfaithful reasoning under reinforcement learning with verifiable rewards.
Specifically, it was… See the full description on the dataset page: https://huggingface.co/datasets/UnfaithRL/aletheia_code_problems.MoralTextManipulation
📊 Exploring LLMs’ Ability to Spontaneously and Conditionally Modify Moral Expressions through Text Manipulation
Morality serves as the foundation of societal structure, guiding legal systems, shaping cultural values, and influencing individual self-perception. With the rise and pervasiveness of generative AI tools, and particularly Large Language Models (LLMs), concerns arise regarding how these tools capture and potentially alter moral dimensions through machine-generated text… See the full description on the dataset page: https://huggingface.co/datasets/MLNTeam-Unical/MoralTextManipulation.Phunny
Phunny: A Humor-Based QA Benchmark for Evaluating LLM Generalization
Welcome to Phunny, a humor-based question answering (QA) benchmark designed to evaluate the reasoning and generalization abilities of large language models (LLMs) through structured puns.
This repository accompanies our ACL 2025 main track paper:"What do you call a dog that is incontrovertibly true? Dogma: Testing LLM Generalization through Humor"
To reproduce our experiments: Code available on GitHub… See the full description on the dataset page: https://huggingface.co/datasets/disi-unibo-nlp/Phunny.earth-love-united-climate-knowledge
🌍 Earth Love United Climate Knowledge Dataset
The most comprehensive open climate science knowledge dataset.
10,128 text chunks + 124 structured facts + 4.54B year geological memory + 10 tipping points.
Built to power GAIA — an AI that embodies the living consciousness of Earth.
Dataset Overview
This dataset gives an AI system authoritative, sourced knowledge about climate change,
carbon, Earth science, and solutions. It has four layers:
Layer 1: Text Knowledge… See the full description on the dataset page: https://huggingface.co/datasets/ego0op/earth-love-united-climate-knowledge.mmlu_mixed_questions
MMLU Mixed Hinted and Unhinted Questions
Dataset Description
This dataset contains multiple-choice questions derived from MMLU and augmented with misleading hints. The misleading hints are intentionally designed to point to an incorrect answer.
The dataset contains a random mixture of:
Hinted examples, where a misleading cue points toward an incorrect answer.
Unhinted examples, where no misleading cue is provided.
The dataset was developed as part of the… See the full description on the dataset page: https://huggingface.co/datasets/UnfaithRL/mmlu_mixed_questions.Multi-Turn-Insurance-Underwriting
Dataset Card for Multi-Turn-Insurance-Underwriting
Dataset Summary
This dataset includes sample traces and associated metadata from multi-turn interactions between a commercial underwriter and AI assistant. We built the system in langgraph with model context protocol and ReAct agents. In each sample, the underwriter has a specific task to solve related to a recent application for insurance by a small business. We created a diverse sample dataset covering 6 distinct types… See the full description on the dataset page: https://huggingface.co/datasets/dennis-panos/Multi-Turn-Insurance-Underwriting.Multi-Turn-Insurance-Underwriting
Dataset Card for Multi-Turn-Insurance-Underwriting
Dataset Summary
This dataset includes sample traces and associated metadata from multi-turn interactions between a commercial underwriter and AI assistant. We built the system in langgraph with model context protocol and ReAct agents. In each sample, the underwriter has a specific task to solve related to a recent application for insurance by a small business. We created a diverse sample dataset covering 6 distinct types… See the full description on the dataset page: https://huggingface.co/datasets/AliDjl/Multi-Turn-Insurance-Underwriting.dodf-saude-qa
DODF Public Health QA
Dataset Summary
DODF Public Health QA is a synthetic question-answering (QA) dataset for evaluating Retrieval-Augmented Generation (RAG) systems over official public health publications from the Diário Oficial do Distrito Federal (DODF), the Government Gazette of the Federal District, Brazil.
The dataset focuses on location-aware questions about public health facilities and services — such as Basic Health Units (UBS), Emergency Care Units… See the full description on the dataset page: https://huggingface.co/datasets/gvic-unb/dodf-saude-qa.ArabicMMLU_undiac
Fajri Koto, Haonan Li, Sara Shatnawi, Jad Doughman, Abdelrahman Boda Sadallah, Aisha Alraeesi, Khalid Almubarak, Zaid Alyafeai, Neha Sengupta, Shady Shehata, Nizar Habash, Preslav Nakov, and Timothy Baldwin
MBZUAI, Prince Sattam bin Abdulaziz University, KFUPM, Core42, NYU Abu Dhabi, The University of Melbourne
Introduction
We present ArabicMMLU, the first multi-task language understanding benchmark for Arabic language, sourced from school exams across diverse… See the full description on the dataset page: https://huggingface.co/datasets/go-inoue/ArabicMMLU_undiac.simson-unified-knowledge-graph
🧠 Simson Unified Knowledge Graph
173 Nodes × 348 Edges – der Klebstoff zwischen allen Simson-Datasets.
Was das ist
Ein maschinenlesbarer Graph, der alle 6 Datasets miteinander verknüpft:
Dataset
Status
Nodes
racing-planet-simson-traces
Diagnose-Traces
15
simson-forum-qa-pairs
Forum-Wissen
30
simson-repair-manual
Technische Daten
14
racing-planet-product-catalog
Teilekatalog
37
simson-youtube-tutorials
Video-Tutorials
20… See the full description on the dataset page: https://huggingface.co/datasets/jmp1987/simson-unified-knowledge-graph.UN16_Peace-Justice
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/maitri-vv/UN16_Peace-Justice.covidqa-unique-context
Dataset Card for "covidqa-unique-context"
More Information needed
ccisd-unified-master-2024
CCISD Unified School Master (2024)
School-level records for Clear Creek Independent School District (Texas), compiled from
the district's public school pages and Texas Education Agency accountability reports.
Covers 39 schools with principal names, contact details, enrollment, and accountability
ratings.
Loading
from datasets import load_dataset
ds = load_dataset("robworks-software/ccisd-unified-master-2024")
all_schools = ds["full"] # all 39 schools… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/ccisd-unified-master-2024.
