datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
worldcup2026
⚽ WorldCup Arena
A Leakage-Free Forecasting Benchmark on a Live Tournament
Can a language model forecast a match — when the match had not been played at the moment it was asked?
🌐 Language / 语言 : 中文 ▾
📊 四张表
点开本页顶部的 Data Studio 标签即可浏览,也可以直接按名字加载。
Config
行数
内容
fixtures
104
基准本体 —— 喂给模型的头部信息,以及结算后的 90 分钟赛果,七个盘口全部推导好(outcome_1x2、over_2_5、both_score、odd_total)
dossiers
2,208
简报索引 —— 46 快照 × 48… See the full description on the dataset page: https://huggingface.co/datasets/Social-AI-2026/worldcup2026.docinsights-2026-shared-task-data
DocInsights 2026 Shared Task: DocSem
Document-grounded quantitative reasoning with evidence attribution
DocSem is the shared task of DocInsights 2026, the Workshop on Document Intelligence and Understanding co-located with EMNLP 2026 in Budapest, Hungary. The workshop theme is Beyond Plain Text: Bridging NLP and Document AI.
Workshop shared task | Source repository | Submission portal | Participant guide
Participants receive a PDF document and a paraphrased user_query. Systems… See the full description on the dataset page: https://huggingface.co/datasets/amitbcp/docinsights-2026-shared-task-data.US_Regulation_ECFR_20260101
US Regulation eCFR 2026-01-01 Dataset
This repository contains a structured, machine-readable version of the Electronic Code of Federal Regulations (eCFR), captured as of January 1, 2026. Unlike the annual CFR snapshots, this dataset reflects the editorialized, near real-time version of federal regulations.
Dataset Description
The eCFR is a daily updated editorial compilation of CFR material and Federal Register amendments. This dataset captures a specific point-in-time… See the full description on the dataset page: https://huggingface.co/datasets/Azzindani/US_Regulation_ECFR_20260101.US_Regulation_CFR_20260101paper-baselines-20260920
Paper baseline evaluation archive
Public archive of 108 selected configurations / 20,808 graded outcomes on BrowseComp-Plus150, DR-9K256, MuSiQue300 and FRAMES150. It includes complete run traces, raw predictions, grading correspondence, frozen retrieval assets and the evaluated runtime image. These are fixed-corpus local evaluation slices, not official full online benchmark results.
Code and reproduction guide: ys-2020/miles, paper/baselines. Frozen source commit: db33363ab57b.… See the full description on the dataset page: https://huggingface.co/datasets/CharlieLLL/paper-baselines-20260920.DocVQA-2026
DocVQA 2026 | ICDAR2026 Competition on Multimodal Reasoning over Documents in Multiple Domains
Building upon previous DocVQA benchmarks, this evaluation dataset introduces challenging reasoning questions over a diverse collection of documents spanning eight domains, including business reports, scientific papers, slides, posters, maps, comics, infographics, and engineering drawings.
By expanding coverage to new document domains and… See the full description on the dataset page: https://huggingface.co/datasets/VLR-CVC/DocVQA-2026.AIME_2000_2026_Kimi_K3
AIME 2000–2026 — Kimi K3 reasoning traces
🔄 Changelog
2026-08-08 — full re-generation. All reasoning traces were regenerated from scratch and re-verified against the official answer key.
New schema — added gen_attempts_low, gen_attempts_high; renamed gen_parsed_answer → gen_answer_int and answer_note → problem_note; removed gen_effort, gen_pass1.
New generation — only use the bare problem (v1 appended an "ANSWER:" format instruction), so traces are cleaner.… See the full description on the dataset page: https://huggingface.co/datasets/bevangelista/AIME_2000_2026_Kimi_K3.persona-drift-contextecho
ContextEcho — Released Dataset
Per-cell evaluation corpus and donated session prefixes for the ContextEcho
benchmark. This Hugging Face repository hosts the released dataset artifacts.
The canonical project page, latest README, code, reproduction instructions, and
donation workflow are maintained on GitHub:
https://github.com/Accenture/ContextEcho
Donate a coding-agent session: https://accenture.github.io/ContextEcho/donate/
For the formal datasheet, see DATASHEET.md.… See the full description on the dataset page: https://huggingface.co/datasets/contextecho2026/persona-drift-contextecho.SocietyBench
🔮 SocietyBench
Forecasting Counterfactual Social-World Evolution
Can a language model forecast how a real social event unfolds — when it cannot tell which event it is?
🌐 Language / 语言 : 中文 ▾
📊 纵览表
点开本页顶部的 Data Studio 标签即可浏览:一行一个事件,五行看完整个榜的规模与构成。
Config
行数
内容
overview
5
一眼看懂这个榜 —— 一行一个事件:领域、时间线节点数、25 个截止点的首尾与跨度、题量与每点均值、真假比、时间事件数与 90 天内占比、A/B/C/D 题型配比、难易分布
from datasets import load_dataset
ov… See the full description on the dataset page: https://huggingface.co/datasets/Social-AI-2026/SocietyBench.100k-corpus-2026
MAST 100K Corpus 2026
This dataset contains the fixed English document corpus used for MAST @ FIRE 2026, the Multilingual Agentic Search Track. MAST evaluates whether multilingual agentic search systems can answer complex questions posed in different languages by retrieving English evidence and producing short, correct English answers.
This corpus is copied from BrowseComp-Plus, a benchmark for Deep-Research systems that isolates the effect of the retriever and the LLM agent to… See the full description on the dataset page: https://huggingface.co/datasets/mast-benchmark/100k-corpus-2026.Nurisk-ICRA2026
Nurisk: VQA for Risk Assessment in Autonomous Driving
Nurisk is a visual question answering dataset focusing on risk assessment for autonomous driving. Each row contains:
image: a BEV image
question: a driving-related question
answer: the ground truth answer
Paper
NuRisk: A Visual Question Answering Dataset for Agent-Level Risk Assessment in Autonomous Driving — see the paper on arXiv:2509.25944 .
Framework
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/TUM-AVS/Nurisk-ICRA2026.razavi-benchRazavi-bench
An expert-curated benchmark for analog-design reasoning.
Razavi-bench packages the question-answer assessments from Behzad Razavi's
Analog Design Experiments With AI Part 1 and Part 2 into a clean
one-task-per-directory benchmark. The tasks probe whether a model can reason
about MOS devices, small-signal circuits, feedback, oscillators, comparators,
dividers, LNAs, TIAs, and LC oscillators.
Each task directory keeps only the benchmark prompt, figure, and curated… See the full description on the dataset page: https://huggingface.co/datasets/Arcadia-2026/razavi-bench.meddialbench
MedDialBench
A controlled factorial benchmark for evaluating LLM diagnostic robustness under parametric adversarial patient behaviors.
Anonymous submission to NeurIPS 2026 Datasets and Benchmarks Track.
The companion paper is currently under double-blind review. Author identities and institutional affiliations are intentionally omitted. After publication, this repository will be transferred to a permanent (de-anonymized) location and this README updated accordingly.… See the full description on the dataset page: https://huggingface.co/datasets/anon-meddial-2026/meddialbench.gaokao_math_2026
gaokao_math_2026 Dataset
gaokao_math_2026 contains 100 questions from five Chinese National College Entrance Examination (Gaokao) mathematics papers administered in 2026. The benchmark preserves the original exam weighting, for a total of 750 points, and covers single-choice, multiple-choice, fill-in-the-blank, and written-response questions.
Dataset Development
gaokao_math_2026 was compiled from five publicly available 2026 Gaokao mathematics papers: National… See the full description on the dataset page: https://huggingface.co/datasets/XHToken/gaokao_math_2026.multilingual-queries-2026
MAST Multilingual Queries 2026
This dataset contains the multilingual query set for MAST @ FIRE 2026, the Multilingual Agentic Search Track. MAST evaluates whether multilingual agentic search systems can answer complex questions posed in different languages.
MAST builds on BrowseComp-Plus (ACL 2026), a reproducible and verifiable extension of BrowseComp with challenging English queries, a verified English corpus of roughly 100K web-sourced documents, and human judgments. In the… See the full description on the dataset page: https://huggingface.co/datasets/mast-benchmark/multilingual-queries-2026.indic-queries-2026
MAST Indic Queries 2026
This dataset contains the Indic query set for MAST @ FIRE 2026, the Multilingual Agentic Search Track. MAST evaluates whether multilingual agentic search systems can answer complex questions posed in different languages.
MAST builds on BrowseComp-Plus (ACL 2026), a reproducible and verifiable extension of BrowseComp with challenging English queries, a verified English corpus of roughly 100K web-sourced documents, and human judgments. In the 2026 MAST… See the full description on the dataset page: https://huggingface.co/datasets/mast-benchmark/indic-queries-2026.recube-dataThis dataset serves as the official data repository for the ReCUBE benchmark; the corresponding evaluation codebase and execution instructions can be found at https://anonymous.4open.science/r/ReCUBE-E0FB/README.md.
Data
This directory contains all benchmark data for ReCUBE.
Download
All data files are hosted on Hugging Face and can be downloaded using:
# Install huggingface_hub if not already installed
pip install huggingface_hub
# Download the entire dataset… See the full description on the dataset page: https://huggingface.co/datasets/recube-anon-2026/recube-data.cybersecurity-soc-threat-hunting-sft-dpo-2026
🛡️ Enterprise Cybersecurity AI, SOC Tier-3 & Threat Hunting SFT/DPO Dataset (2026)
High-precision multi-turn instruction tuning and preference optimization dataset with step-by-step SOC Tier-3 Chain-of-Thought (<thought>) kill-chain diagnostic trees for fine-tuning LLMs (Llama-3.3, Qwen-2.5-Coder, DeepSeek-R1-Distill, Mistral) into Senior SOC Threat Hunters, Incident Responders, and Red-Team Defense Architects.
📊 Dataset Architecture & Highlights… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/cybersecurity-soc-threat-hunting-sft-dpo-2026.financial-statement-modeling-sft-dpo-2026
📈 Enterprise Financial AI, SEC 10-K & Valuation Modeling SFT/DPO Dataset (2026)
High-precision multi-turn instruction tuning and preference optimization dataset with step-by-step arithmetic Chain-of-Thought (<thought>) reasoning chains for fine-tuning LLMs (Llama-3.3, Qwen-2.5-Coder, DeepSeek-R1-Distill, Mistral) into Wall Street Equity Research Associates, M&A Valuation Modelers, and Senior Forensic Auditors.
📊 Dataset Architecture & Highlights
Multi-Turn… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/financial-statement-modeling-sft-dpo-2026.wikipedia-en-2026-07-01-passages
English Wikipedia Passages, Chunked (2026-07-01)
Every English Wikipedia article split into retrieval-sized passages with title
and section attached. A clean, dated corpus for RAG — embed it yourself, or use
the ready-made vectors and indexes in the companion repos:
embeddings
·
faiss.
Contents
17,473,199 passages, 21 GB, JSON Lines (one passage per line).
Fields: id (<pageid>#<n>), title, section, text.
Line order matches ids.txt / vector row order in the… See the full description on the dataset page: https://huggingface.co/datasets/Sherlock-Comms/wikipedia-en-2026-07-01-passages.Agent-ValueBench
Agent-ValueBench
Agent-ValueBench constitutes the first comprehensive benchmark dedicated to evaluating the underlying values of autonomous agents. It features 394 executable environments across 16 domains, offering 4,335 value-conflict tasks that span 28 value systems (332 dimensions).
This Hugging Face release contains both structured JSONL tables for dataset viewing and Croissant metadata generation, and the original raw benchmark artifacts.
Repository Structure… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-nips2026/Agent-ValueBench.swe-bench-multi-file-refactoring-sft-dpo-2026
💻 Enterprise Autonomous SWE-bench AI & Multi-File Code Refactoring SFT/DPO Dataset (2026)
High-precision multi-turn instruction tuning and preference optimization dataset with step-by-step call-stack Chain-of-Thought (<thought>) reasoning trees for fine-tuning LLMs (Qwen-2.5-Coder, Llama-3.3, DeepSeek-R1-Distill, Mistral) into Autonomous Software Engineers and SWE-bench Benchmark Agents.
📊 Dataset Architecture & Highlights
Multi-Turn Code Reviews:… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/swe-bench-multi-file-refactoring-sft-dpo-2026.submission-eval-artifacts
NeurIPS ED 2026 Anonymous Evaluation Artifacts
This dataset repo contains sanitized evaluation artifacts for an anonymous NeurIPS ED 2026 submission. It is metadata-focused: normalized benchmark JSON, selected small paper-facing summaries, reviewer indexes, and manifests.
Checkpoint artifacts are referenced through neurips-ed2026-anon-checkpoints/submission-checkpoints. This dataset repo does not contain model checkpoints or model weights.
Anonymous code artifact:… See the full description on the dataset page: https://huggingface.co/datasets/neurips-ed2026-anon-checkpoints/submission-eval-artifacts.IPHO2026
IPhO 2026 Curated Problems
This repository packages the official English problem, solution, and marking
materials for the LVI International Physics Olympiad (Bucaramanga, Colombia,
2026) as machine-readable, subquestion-level records.
Contents
Configuration
Rows
Description
all
41
All curated subquestions
theory
23
Theory papers T1–T3
experiment
18
Experimental paper E1
formalization_ready
29
Subset selected for theorem formalization… See the full description on the dataset page: https://huggingface.co/datasets/humanfia-lab/IPHO2026.cybersec-master-dataset
Cybersecurity Master Instruction Dataset
Overview
A large-scale cybersecurity instruction-tuning dataset in ShareGPT conversational format,
assembled from multiple authoritative open sources and deduplicated.
At 1,807,941 deduplicated records, this appears to be one of the larger cybersecurity
LLM fine-tuning / instruction-style datasets on Hugging Face, and likely among the larger
broad vulnerability-intelligence corpora in conversational/instruction format. It is… See the full description on the dataset page: https://huggingface.co/datasets/Voidreaper2026/cybersec-master-dataset.Leandata
LEANDATA
A collection of Lean-formalized STEM problem-solving examples across physics, chemistry, calculus, probability, and related domains.
Dataset summary
Dataset page: https://huggingface.co/datasets/anon-ed-2026/Leandata
Total examples: 580
Loading with datasets
from datasets import load_dataset
ds = load_dataset("anon-ed-2026/Leandata", "atkins")
print(ds["train"][0]["problem_id"])
openai-terra-batch-wiki-brazil-1000-partial-20260724-01
OpenAI Terra Batch — Wikipédia PT-BR (run parcial)
Checkpoint publicável de uma execução real e interrompida do fluxo
document_task_matrix. A execução planejou gerar uma matriz de 1.000
documentos da Wikipédia em português por 25 tasks canônicas usando a Responses
API Batch e o modelo gpt-5.6-terra.
Este repositório não representa a conclusão dos 25.000 pares planejados. Ele
contém somente os 1.282 candidatos aceitos após a reconciliação offline de
todos os resultados Batch já… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/openai-terra-batch-wiki-brazil-1000-partial-20260724-01.enwikivoyage-retrieval-202605
English Wikivoyage Retrieval 2026-05
I like travel datasets because they are about real places, real constraints,
and the small practical questions people ask before they go somewhere. I am
sharing this English Wikivoyage retrieval corpus in that spirit: as honest
work from a researcher-builder who wants to explore the world, make the
pipeline inspectable, and let other people reuse or challenge the choices.
This is not presented as a finished travel product or a private… See the full description on the dataset page: https://huggingface.co/datasets/vincentss/enwikivoyage-retrieval-202605.Theory_of_Mind_CoMMET
CoMMET
CoMMET is a multi-turn, multimodal benchmark designed to evaluate the Theory of Mind (ToM) capabilities of multimodal large language models (MLLMs).
Unlike conventional single-turn Theory of Mind benchmarks, CoMMET represents each scenario as a sequence of interconnected turns. Models are required to reason about stories, previous interactions, feedback, questions, and, when necessary, visual information.
The benchmark covers multiple types of mental-state reasoning and… See the full description on the dataset page: https://huggingface.co/datasets/rrchen2026/Theory_of_Mind_CoMMET.ipho2026-formalized-results
IPhO 2026 Lean formalization results
This dataset publishes two independent Lean 4 solution sets for the 23
theoretical subproblems of IPhO 2026:
Result set
Theory targets
Compiles
Active proof placeholders
Codex v2
23/23
yes
0
Kimi K3
23/23
yes
0
The six selected experimental targets are outside the requested accuracy
scope and are not included here. Each result directory is a standalone Lean
project containing the theorem files and pinned build dependencies.… See the full description on the dataset page: https://huggingface.co/datasets/humanfia-lab/ipho2026-formalized-results.
