CoolFace
23 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01dell-research-harvard /newswire Dataset Card for NewsWire Dataset Summary NewsWire contains 2.7 million unique public domain U.S. news wire articles, written between 1878 and 1977. Locations in these articles are georeferenced, topics are tagged using customized neural topic classification, named entities are recognized, and individuals are disambiguated to Wikipedia using a novel entity disambiguation model. Languages English (en) Dataset Structure Each year in the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/dell-research-harvard/newswire.tabulartext-classification1M<n<10M92 likes4.1k downloads1y agoHugging Face02harimo /scorio-lite Scorio Lite contains 1,211,520 sampled attempts from four model configurations and six reasoning benchmarks. Each model was run 80 times on every question. The five competition-math splits contain 186 questions. The superGPQA split contains a frozen, field-balanced sample of 3,600 questions. Each row includes the generation, rule-based grading, scores from CompassVerifier-3B and a reference-free verifier, and aggregate token statistics. Per-model configs also include token strings, log… See the full description on the dataset page: https://huggingface.co/datasets/harimo/scorio-lite.tabulartext-generation1M<n<10M0 likes2.5k downloads1mo agoHugging Face03Zhongzhi1228 /Terminal-Bench-Hard Terminal-Bench Hard Terminal-Bench Hard is a set of 100 challenging terminal-based agent tasks. The tasks cover software engineering, debugging, data processing, system administration, security, scientific computing, and related command-line workflows. Contents tasks/: runnable tasks in Harbor format. metadata/tasks.parquet: searchable task metadata and instructions. Each task directory contains task.toml, instruction.md, an environment/ directory, and verifier… See the full description on the dataset page: https://huggingface.co/datasets/Zhongzhi1228/Terminal-Bench-Hard.imagequestion-answeringn<1K0 likes2k downloads2mo agoHugging Face04gyung /korean-bar-exam-hard-current-law-precedent-sft-1000 Korean Current-Law Bar Exam Hard SFT 1000 대한민국 현행 법령을 기준으로 만든 변호사시험 선택형 고난도 스타일 SFT 데이터 1,000문항입니다. 초기 직접 조문확인형 생성본은 실제 제14ㆍ15회 변호사시험보다 쉬워서, 이 버전은 다음 기준으로 다시 만들었습니다. ㄱ/ㄴ/ㄷ/ㄹ 복합정오형 중심 甲/乙/丙, 검사ㆍ사법경찰관ㆍ행정청ㆍ회사ㆍ소송당사자 등이 등장하는 사례형 비중 확대 단순 근거 조문 선택형 제거 정답뿐 아니라 각 지문별 O/X 이유와 참고 법령 조문 제공 제15회 변호사시험 data/questions.csv와 높은 유사도 문항 제외 Files data/questions.csv: Hugging Face preview용 메인 CSV입니다. sft/train.jsonl: messages 형식 SFT용 JSONL입니다. metadata/qa_report.json: 생성 수량, 난도 관련… See the full description on the dataset page: https://huggingface.co/datasets/gyung/korean-bar-exam-hard-current-law-precedent-sft-1000.tabularquestion-answering1K<n<10K0 likes306 downloads4mo agoHugging Face05harsh147 /GroundCocoa Dataset Card for Dataset Name GroundCocoa is a benchmark to evaluate conditional and compositional reasoning in large language models through a flight-booking task presented in multiple-choice format. Dataset Details The test set consists of 4849 samples consisting of 728 unique user requirements. User requirements may be repeated with varying options. In additon, we also provide a small validation set that may be used for certain parameter tuning. It consists of 52… See the full description on the dataset page: https://huggingface.co/datasets/harsh147/GroundCocoa.tabularquestion-answering1K<n<10K1 likes290 downloads2y agoHugging Face06DJLougen /harmonic-reasoning-v1 Harmonic Reasoning v1 Support This Work I'm a PhD student in visual neuroscience at the University of Toronto who also happens to spend way too much time fine-tuning, merging, and quantizing open-weight models on rented H100s and a local DGX Spark. All training compute is self-funded — balancing GPU costs against a student budget. If my open-weight models or datasets have been useful to you, consider supporting future releases. Support on Ko-fi Harmonic Reasoning v1 is a… See the full description on the dataset page: https://huggingface.co/datasets/DJLougen/harmonic-reasoning-v1.tabulartext-generationn<1K27 likes94 downloads6mo agoHugging Face07fhyfhy /diffusion-vs-ar-hard-sudoku Diffusion vs AR Hard Sudoku This repository packages 8,148,696 Sudoku examples in the CSV format expected by HKUNLP/diffusion-vs-ar, plus its original 100k/1k easy baseline. Every processed file has these columns: column meaning quizzes 81 row-major digits; 0 is an empty cell solutions complete 81-digit solution source original collection dataset normalized dataset family official_rating rating supplied by the source rating_type semantics of that rating… See the full description on the dataset page: https://huggingface.co/datasets/fhyfhy/diffusion-vs-ar-hard-sudoku.tabularquestion-answering1M<n<10M0 likes74 downloads2mo agoHugging Face08AmirMohseni /SimpleQA-verified-Hard-Qwen3-8B SimpleQA Verified Hard for Qwen3-8B Dataset Summary This dataset contains the 866 questions that Qwen/Qwen3-8B failed to answer correctly in up to eight attempts from the official google/simpleqa-verified benchmark. Each source question was scheduled for eight stochastic generations. As soon as one generation was graded CORRECT, sampling stopped and the question was excluded. Questions retained here therefore have pass@8 = 0 under the model, prompt, sampling, and… See the full description on the dataset page: https://huggingface.co/datasets/AmirMohseni/SimpleQA-verified-Hard-Qwen3-8B.tabularquestion-answeringn<1K0 likes42 downloads2mo agoHugging Face09Neura-parse /quantum-hardware-device-physics Neura Parse — Quantum Hardware Device Physics: Qubit Design, Coherence, Control & Scaling A physics- and engineering-deep vertical on how qubits are built, controlled, and scaled across superconducting, trapped-ion, neutral-atom, and spin modalities (plus emerging erasure/biased-noise qubits). Device-physics derivations, coherence-limit analyses, control-stack engineering, and 2025-2026 scaling/interconnect work, with QuTiP/scqubits simulation context — expanding the general… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-hardware-device-physics.tabulartext-generation100K<n<1M0 likes39 downloads3mo agoHugging Face10Synho /hard-layer-v3-epistemic-honesty VMTI Hard Layer v3: Epistemic Honesty Benchmark for Biomedical LLMs Dataset Description The VMTI-Trust Index (VTI) Hard Layer v3 benchmark evaluates large language models' ability to detect numerical contradictions and physiological impossibilities in clinical trial data. Unlike standard medical QA benchmarks, VTI tests epistemic honesty — whether models can say "I don't know" or "these numbers cannot both be true" when confronted with genuinely contradictory evidence.… See the full description on the dataset page: https://huggingface.co/datasets/Synho/hard-layer-v3-epistemic-honesty.tabularquestion-answering1K<n<10K0 likes38 downloads5mo agoHugging Face11PhantomG27249 /harmonic-reasoning-v1 Harmonic Reasoning v1 Support This Work I'm a PhD student in visual neuroscience at the University of Toronto who also happens to spend way too much time fine-tuning, merging, and quantizing open-weight models on rented H100s and a local DGX Spark. All training compute is self-funded — balancing GPU costs against a student budget. If my open-weight models or datasets have been useful to you, consider supporting future releases. Support on Ko-fi Harmonic Reasoning v1 is a… See the full description on the dataset page: https://huggingface.co/datasets/PhantomG27249/harmonic-reasoning-v1.tabulartext-generationn<1K0 likes34 downloads5mo agoHugging Face12Jinesis /gt-harmbench GT-HarmBench GT-HarmBench is a game-theoretic AI safety benchmark that evaluates whether large language models can reason strategically in realistic, AI-risk–grounded scenarios. Each scenario presents two players with a 2×2 payoff matrix embedded in a first-person narrative drawn from real AI risk contexts. Models are evaluated on their ability to: identify and play Nash equilibria (individual rationality), select actions that maximise utilitarian welfare (sum of payoffs)… See the full description on the dataset page: https://huggingface.co/datasets/Jinesis/gt-harmbench.tabularquestion-answering1K<n<10K0 likes29 downloads2mo agoHugging Face13Testing333555 /harmonic-reasoning-v1 Harmonic Reasoning v1 Support This Work I'm a PhD student in visual neuroscience at the University of Toronto who also happens to spend way too much time fine-tuning, merging, and quantizing open-weight models on rented H100s and a local DGX Spark. All training compute is self-funded — balancing GPU costs against a student budget. If my open-weight models or datasets have been useful to you, consider supporting future releases. Support on Ko-fi Harmonic Reasoning v1 is a… See the full description on the dataset page: https://huggingface.co/datasets/Testing333555/harmonic-reasoning-v1.tabulartext-generationn<1K0 likes28 downloads5mo agoHugging Face14tyrtleli /thinking-benchmark-hard-but-doable Thinking Benchmark — Hard-but-Doable Panel Eight competition-math problems selected for the Cost of Overthinking study's controlled trace-length comparison. These are the "hold the problem constant" panel: problems that every tested frontier model (GPT-5, GPT-5.4, o3) solves reliably (≥7/8 at k=8) but still has to genuinely reason about (no instant one-shots). The goal is to observe how mean and variance of reasoning-trace length differ across models on identical, non-trivial… See the full description on the dataset page: https://huggingface.co/datasets/tyrtleli/thinking-benchmark-hard-but-doable.tabularquestion-answeringn<1K0 likes16 downloads4mo agoHugging Face151337xyz1337xyz /sciknoweval-v2-hard-autogradable-512-2026-04-28 SciKnowEval v2 Hard Autogradable 512 - 2026-04-28 A 512-example sanity subset sampled from hicai-zju/SciKnowEval (v2, test) for Plan-CRL scientific reasoning evals. Selection seed: 20260428. Filtering and balancing: excludes L1 keeps L2, L3, L4 keeps autogradable types: mcq-4-choices, mcq-2-choices, true_or_false, filling requires answerKey or answer balances domains at 128 examples each: Biology, Chemistry, Material, Physics per domain: 32 L2, 48 L3, 48 L4 Useful fields for… See the full description on the dataset page: https://huggingface.co/datasets/1337xyz1337xyz/sciknoweval-v2-hard-autogradable-512-2026-04-28.tabularquestion-answeringn<1K0 likes15 downloads5mo agoHugging Face16achinta3 /cybersec-jsonschemabench-cloudtrail-objective-natural-hard-v5 CybersecJSONSchemaBench CloudTrail Objective Natural Hard v5 This is a 100-problem natural-prompt long-context cybersecurity reasoning subset built from the full flAWS CloudTrail corpus. Each row contains a natural analyst-style question, a large CloudTrail JSONL context, and the JSON schema the answer must match. Gold answers are deterministic hidden-oracle results over the serialized slice and are not included in this public export. Families actor_recon_to_change: 22… See the full description on the dataset page: https://huggingface.co/datasets/achinta3/cybersec-jsonschemabench-cloudtrail-objective-natural-hard-v5.tabularquestion-answeringn<1K0 likes14 downloads5mo agoHugging Face17EscheWang /ChemRD-Hard-100gated ChemRD-Hard-100 100 bilingual (English / Chinese) PhD-level chemistry items, selected by measured difficulty from a 474-item verified pool. Every item is self-contained: everything needed to answer it is in the item. Leaderboard 18 arms, each one isolated process or request per item, both languages, no shared context between items. Reasoning is off wherever the endpoint allows it, and the last column reports the reasoning tokens actually observed, so the claim is… See the full description on the dataset page: https://huggingface.co/datasets/EscheWang/ChemRD-Hard-100.tabularquestion-answeringn<1K0 likes11 downloads1mo agoHugging Face18har123ish /indian-government-schemes-2025 Indian Government Schemes Dataset 2026 Dataset Description The most comprehensive structured dataset of Indian central and state government schemes — 4,693 schemes across all ministries and states, with machine-readable eligibility fields. Maintained by SmartDuke Technologies · Coimbatore, Tamil Nadu, India This dataset powers SchemeFit — India's government scheme finder for citizens and businesses. What Makes This Different Most existing Indian… See the full description on the dataset page: https://huggingface.co/datasets/har123ish/indian-government-schemes-2025.tabulartext-classification1K<n<10K0 likes10 downloads1mo agoHugging Face19LLMTeamAkiyama /cleand_HARDMath元データ: https://github.com/sarahmart/HARDMath?tab=readme-ov-file dataフォルダのMARDMath.jsonより データ件数: 1,054 平均トークン数: 2112 最大トークン数: 28,886 合計トークン数: 2,226,401 ファイル形式: JSONL ファイルサイズ: 5.4 MB tabularquestion-answering1K<n<10K0 likes9 downloads1y agoHugging Face20achinta3 /cybersec-jsonschemabench-cloudtrail-hard-v2-400 CybersecJSONSchemaBench CloudTrail Hard v2 400 This is a 100-problem synthetic long-context cybersecurity reasoning subset built from the full flAWS CloudTrail corpus. The benchmark asks models to return JSON matching the provided answer schema. Each row contains a short analyst request, a large CloudTrail JSONL context, and hidden deterministic evaluation metadata. This variant uses shorter 400-record contexts than the full CloudTrail Hard v2 export so direct API evaluation is less… See the full description on the dataset page: https://huggingface.co/datasets/achinta3/cybersec-jsonschemabench-cloudtrail-hard-v2-400.tabularquestion-answeringn<1K0 likes8 downloads5mo agoHugging Face21achinta3 /cybersec-jsonschemabench-cloudtrail-objective-hard-v3 CybersecJSONSchemaBench CloudTrail Objective Hard v3 This is a 100-problem objective long-context cybersecurity reasoning subset built from the full flAWS CloudTrail corpus. Each row contains an objective query prompt, a large CloudTrail JSONL context, and the JSON schema the answer must match. Gold answers are deterministic query results over the serialized slice and are not included in this public export. Families apigateway_restapi_event_profile: 10… See the full description on the dataset page: https://huggingface.co/datasets/achinta3/cybersec-jsonschemabench-cloudtrail-objective-hard-v3.tabularquestion-answeringn<1K0 likes8 downloads5mo agoHugging Face22achinta3 /cybersec-jsonschemabench-cloudtrail-natural-hard-v4 CybersecJSONSchemaBench CloudTrail Natural Hard v4 This is a 100-problem natural-prompt long-context cybersecurity reasoning subset built from the full flAWS CloudTrail corpus. Each row contains a natural analyst-style question, a large CloudTrail JSONL context, and the JSON schema the answer must match. Gold answers are deterministic hidden-oracle results over the serialized slice and are not included in this public export. Families actor_recon_to_change: 20… See the full description on the dataset page: https://huggingface.co/datasets/achinta3/cybersec-jsonschemabench-cloudtrail-natural-hard-v4.tabularquestion-answeringn<1K0 likes5 downloads5mo agoHugging Face23Harvin1111 /MMLU-Pro-json MMLU-Pro json This is a reupload of MMLU-Pro in json format. Please, refer to the original dataset for details. tabularquestion-answering10K<n<100K0 likes4 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.