datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
newswire
Dataset Card for NewsWire
Dataset Summary
NewsWire contains 2.7 million unique public domain U.S. news wire articles, written between 1878 and 1977. Locations in these articles are georeferenced, topics are tagged using customized neural topic classification, named entities are recognized, and individuals are disambiguated to Wikipedia using a novel entity disambiguation model.
Languages
English (en)
Dataset Structure
Each year in the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/dell-research-harvard/newswire.scorio-lite
Scorio Lite contains 1,211,520 sampled attempts from four model configurations and
six reasoning benchmarks. Each model was run 80 times on every question.
The five competition-math splits contain 186 questions. The superGPQA split contains a
frozen, field-balanced sample of 3,600 questions. Each row includes the generation,
rule-based grading, scores from CompassVerifier-3B and a reference-free verifier, and
aggregate token statistics. Per-model configs also include token strings, log… See the full description on the dataset page: https://huggingface.co/datasets/harimo/scorio-lite.Terminal-Bench-Hard
Terminal-Bench Hard
Terminal-Bench Hard is a set of 100 challenging terminal-based agent tasks.
The tasks cover software engineering, debugging, data processing, system
administration, security, scientific computing, and related command-line
workflows.
Contents
tasks/: runnable tasks in Harbor format.
metadata/tasks.parquet: searchable task metadata and instructions.
Each task directory contains task.toml, instruction.md, an
environment/ directory, and verifier… See the full description on the dataset page: https://huggingface.co/datasets/Zhongzhi1228/Terminal-Bench-Hard.korean-bar-exam-hard-current-law-precedent-sft-1000
Korean Current-Law Bar Exam Hard SFT 1000
대한민국 현행 법령을 기준으로 만든 변호사시험 선택형 고난도 스타일 SFT 데이터 1,000문항입니다.
초기 직접 조문확인형 생성본은 실제 제14ㆍ15회 변호사시험보다 쉬워서, 이 버전은 다음 기준으로 다시 만들었습니다.
ㄱ/ㄴ/ㄷ/ㄹ 복합정오형 중심
甲/乙/丙, 검사ㆍ사법경찰관ㆍ행정청ㆍ회사ㆍ소송당사자 등이 등장하는 사례형 비중 확대
단순 근거 조문 선택형 제거
정답뿐 아니라 각 지문별 O/X 이유와 참고 법령 조문 제공
제15회 변호사시험 data/questions.csv와 높은 유사도 문항 제외
Files
data/questions.csv: Hugging Face preview용 메인 CSV입니다.
sft/train.jsonl: messages 형식 SFT용 JSONL입니다.
metadata/qa_report.json: 생성 수량, 난도 관련… See the full description on the dataset page: https://huggingface.co/datasets/gyung/korean-bar-exam-hard-current-law-precedent-sft-1000.GroundCocoa
Dataset Card for Dataset Name
GroundCocoa is a benchmark to evaluate conditional and compositional reasoning in large language models through a flight-booking task presented in multiple-choice format.
Dataset Details
The test set consists of 4849 samples consisting of 728 unique user requirements. User requirements may be repeated with varying options. In additon, we also provide a small validation set that may be used for certain parameter tuning. It consists of 52… See the full description on the dataset page: https://huggingface.co/datasets/harsh147/GroundCocoa.harmonic-reasoning-v1
Harmonic Reasoning v1
Support This Work
I'm a PhD student in visual neuroscience at the University of Toronto who also happens to spend way too much time fine-tuning, merging, and quantizing open-weight models on rented H100s and a local DGX Spark. All training compute is self-funded — balancing GPU costs against a student budget. If my open-weight models or datasets have been useful to you, consider supporting future releases.
Support on Ko-fi
Harmonic Reasoning v1 is a… See the full description on the dataset page: https://huggingface.co/datasets/DJLougen/harmonic-reasoning-v1.diffusion-vs-ar-hard-sudoku
Diffusion vs AR Hard Sudoku
This repository packages 8,148,696 Sudoku examples in the CSV format expected
by HKUNLP/diffusion-vs-ar, plus its original 100k/1k easy baseline.
Every processed file has these columns:
column
meaning
quizzes
81 row-major digits; 0 is an empty cell
solutions
complete 81-digit solution
source
original collection
dataset
normalized dataset family
official_rating
rating supplied by the source
rating_type
semantics of that rating… See the full description on the dataset page: https://huggingface.co/datasets/fhyfhy/diffusion-vs-ar-hard-sudoku.SimpleQA-verified-Hard-Qwen3-8B
SimpleQA Verified Hard for Qwen3-8B
Dataset Summary
This dataset contains the 866 questions that
Qwen/Qwen3-8B failed to answer correctly in up to eight attempts from the
official
google/simpleqa-verified
benchmark.
Each source question was scheduled for eight stochastic generations. As soon
as one generation was graded CORRECT, sampling stopped and the question was
excluded. Questions retained here therefore have pass@8 = 0 under the
model, prompt, sampling, and… See the full description on the dataset page: https://huggingface.co/datasets/AmirMohseni/SimpleQA-verified-Hard-Qwen3-8B.quantum-hardware-device-physics
Neura Parse — Quantum Hardware Device Physics: Qubit Design, Coherence, Control & Scaling
A physics- and engineering-deep vertical on how qubits are built, controlled, and scaled across superconducting, trapped-ion, neutral-atom, and spin modalities (plus emerging erasure/biased-noise qubits). Device-physics derivations, coherence-limit analyses, control-stack engineering, and 2025-2026 scaling/interconnect work, with QuTiP/scqubits simulation context — expanding the general… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-hardware-device-physics.hard-layer-v3-epistemic-honesty
VMTI Hard Layer v3: Epistemic Honesty Benchmark for Biomedical LLMs
Dataset Description
The VMTI-Trust Index (VTI) Hard Layer v3 benchmark evaluates large language models' ability to detect numerical contradictions and physiological impossibilities in clinical trial data. Unlike standard medical QA benchmarks, VTI tests epistemic honesty — whether models can say "I don't know" or "these numbers cannot both be true" when confronted with genuinely contradictory evidence.… See the full description on the dataset page: https://huggingface.co/datasets/Synho/hard-layer-v3-epistemic-honesty.harmonic-reasoning-v1
Harmonic Reasoning v1
Support This Work
I'm a PhD student in visual neuroscience at the University of Toronto who also happens to spend way too much time fine-tuning, merging, and quantizing open-weight models on rented H100s and a local DGX Spark. All training compute is self-funded — balancing GPU costs against a student budget. If my open-weight models or datasets have been useful to you, consider supporting future releases.
Support on Ko-fi
Harmonic Reasoning v1 is a… See the full description on the dataset page: https://huggingface.co/datasets/PhantomG27249/harmonic-reasoning-v1.gt-harmbench
GT-HarmBench
GT-HarmBench is a game-theoretic AI safety benchmark that evaluates whether large language models can reason strategically in realistic, AI-risk–grounded scenarios.
Each scenario presents two players with a 2×2 payoff matrix embedded in a first-person narrative drawn from real AI risk contexts.
Models are evaluated on their ability to:
identify and play Nash equilibria (individual rationality),
select actions that maximise utilitarian welfare (sum of payoffs)… See the full description on the dataset page: https://huggingface.co/datasets/Jinesis/gt-harmbench.harmonic-reasoning-v1
Harmonic Reasoning v1
Support This Work
I'm a PhD student in visual neuroscience at the University of Toronto who also happens to spend way too much time fine-tuning, merging, and quantizing open-weight models on rented H100s and a local DGX Spark. All training compute is self-funded — balancing GPU costs against a student budget. If my open-weight models or datasets have been useful to you, consider supporting future releases.
Support on Ko-fi
Harmonic Reasoning v1 is a… See the full description on the dataset page: https://huggingface.co/datasets/Testing333555/harmonic-reasoning-v1.thinking-benchmark-hard-but-doable
Thinking Benchmark — Hard-but-Doable Panel
Eight competition-math problems selected for the Cost of Overthinking study's controlled trace-length comparison. These are the "hold the problem constant" panel: problems that every tested frontier model (GPT-5, GPT-5.4, o3) solves reliably (≥7/8 at k=8) but still has to genuinely reason about (no instant one-shots).
The goal is to observe how mean and variance of reasoning-trace length differ across models on identical, non-trivial… See the full description on the dataset page: https://huggingface.co/datasets/tyrtleli/thinking-benchmark-hard-but-doable.sciknoweval-v2-hard-autogradable-512-2026-04-28
SciKnowEval v2 Hard Autogradable 512 - 2026-04-28
A 512-example sanity subset sampled from hicai-zju/SciKnowEval (v2, test) for Plan-CRL scientific reasoning evals.
Selection seed: 20260428.
Filtering and balancing:
excludes L1
keeps L2, L3, L4
keeps autogradable types: mcq-4-choices, mcq-2-choices, true_or_false, filling
requires answerKey or answer
balances domains at 128 examples each: Biology, Chemistry, Material, Physics
per domain: 32 L2, 48 L3, 48 L4
Useful fields for… See the full description on the dataset page: https://huggingface.co/datasets/1337xyz1337xyz/sciknoweval-v2-hard-autogradable-512-2026-04-28.cybersec-jsonschemabench-cloudtrail-objective-natural-hard-v5
CybersecJSONSchemaBench CloudTrail Objective Natural Hard v5
This is a 100-problem natural-prompt long-context cybersecurity reasoning subset built from the full flAWS CloudTrail corpus.
Each row contains a natural analyst-style question, a large CloudTrail JSONL context, and the JSON schema the answer must match. Gold answers are deterministic hidden-oracle results over the serialized slice and are not included in this public export.
Families
actor_recon_to_change: 22… See the full description on the dataset page: https://huggingface.co/datasets/achinta3/cybersec-jsonschemabench-cloudtrail-objective-natural-hard-v5.ChemRD-Hard-100
ChemRD-Hard-100
100 bilingual (English / Chinese) PhD-level chemistry items, selected by measured difficulty from a
474-item verified pool. Every item is self-contained: everything needed to answer it is in the item.
Leaderboard
18 arms, each one isolated process or request per item, both languages, no shared context between
items. Reasoning is off wherever the endpoint allows it, and the last column reports the reasoning
tokens actually observed, so the claim is… See the full description on the dataset page: https://huggingface.co/datasets/EscheWang/ChemRD-Hard-100.indian-government-schemes-2025
Indian Government Schemes Dataset 2026
Dataset Description
The most comprehensive structured dataset of Indian central and state government schemes — 4,693 schemes across all ministries and states, with machine-readable eligibility fields.
Maintained by SmartDuke Technologies · Coimbatore, Tamil Nadu, India
This dataset powers SchemeFit — India's government scheme finder for citizens and businesses.
What Makes This Different
Most existing Indian… See the full description on the dataset page: https://huggingface.co/datasets/har123ish/indian-government-schemes-2025.cleand_HARDMath元データ: https://github.com/sarahmart/HARDMath?tab=readme-ov-file
dataフォルダのMARDMath.jsonより
データ件数: 1,054
平均トークン数: 2112
最大トークン数: 28,886
合計トークン数: 2,226,401
ファイル形式: JSONL
ファイルサイズ: 5.4 MB
cybersec-jsonschemabench-cloudtrail-hard-v2-400
CybersecJSONSchemaBench CloudTrail Hard v2 400
This is a 100-problem synthetic long-context cybersecurity reasoning subset built from the full flAWS CloudTrail corpus.
The benchmark asks models to return JSON matching the provided answer schema. Each row contains a short analyst request, a large CloudTrail JSONL context, and hidden deterministic evaluation metadata.
This variant uses shorter 400-record contexts than the full CloudTrail Hard v2 export so direct API evaluation is less… See the full description on the dataset page: https://huggingface.co/datasets/achinta3/cybersec-jsonschemabench-cloudtrail-hard-v2-400.cybersec-jsonschemabench-cloudtrail-objective-hard-v3
CybersecJSONSchemaBench CloudTrail Objective Hard v3
This is a 100-problem objective long-context cybersecurity reasoning subset built from the full flAWS CloudTrail corpus.
Each row contains an objective query prompt, a large CloudTrail JSONL context, and the JSON schema the answer must match. Gold answers are deterministic query results over the serialized slice and are not included in this public export.
Families
apigateway_restapi_event_profile: 10… See the full description on the dataset page: https://huggingface.co/datasets/achinta3/cybersec-jsonschemabench-cloudtrail-objective-hard-v3.cybersec-jsonschemabench-cloudtrail-natural-hard-v4
CybersecJSONSchemaBench CloudTrail Natural Hard v4
This is a 100-problem natural-prompt long-context cybersecurity reasoning subset built from the full flAWS CloudTrail corpus.
Each row contains a natural analyst-style question, a large CloudTrail JSONL context, and the JSON schema the answer must match. Gold answers are deterministic hidden-oracle results over the serialized slice and are not included in this public export.
Families
actor_recon_to_change: 20… See the full description on the dataset page: https://huggingface.co/datasets/achinta3/cybersec-jsonschemabench-cloudtrail-natural-hard-v4.MMLU-Pro-json
MMLU-Pro json
This is a reupload of MMLU-Pro in json format. Please, refer to the original dataset for details.
