CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01opencompass /SWEBench-Pro-Verified SWE-Bench Pro Verified: Anti-hacking & Task refinement SWE-Bench Pro has emerged as a standard benchmark for evaluating software engineering agents on challenging repository-level tasks. However, our analysis work show that its evaluation is undermined by two sources of unreliability: reward hacking, enabled by leakage of gold solutions or hidden evaluation information, and task quality issues, including misleading problem statements and improperly scoped tests. These issues can… See the full description on the dataset page: https://huggingface.co/datasets/opencompass/SWEBench-Pro-Verified.tabularn<1K3 likes1.6k downloads16d agoHugging Face02YefanZhou98 /LLMVerify-Verifier LLMVerify-Verifier Verification results dataset for the paper "Variation in Verification: Understanding Verification Dynamics in Large Language Models", accepted at ICLR 2026 (arXiv:2509.17995). This dataset contains the binary verdicts and chain-of-thought verification reasoning produced by 15 verifier models judging candidate solutions from 15 generator models across three task domains. It supports systematic analysis of how problem difficulty, generator capability, and verifier… See the full description on the dataset page: https://huggingface.co/datasets/YefanZhou98/LLMVerify-Verifier.tabular1M<n<10M1 likes232 downloads5mo agoHugging Face03fatihdx /tr-rss-haber-akisi-verisi TR-RSS Haber Akışı Verisi TL;DR — Bu veri seti, Türkiye odaklı haber/RSS akışlarından toplanan kayıtları; mükerrerlik, spam, reklam, amaç dışı kategori, yurtdışı odak ve editoryal çerçeve yoğunluğu açısından katmanlı kalite kontrolden geçirerek erken sinyal üretimine uygun hâle getirir. Doğrulama kararı / verdict üretmez; ClaimReview ve dezenformasyon araştırmaları için upstream izleme ve kaynak önceliklendirme katmanı olarak tasarlanmıştır. Ölçek: 307.800 öğe incelendi →… See the full description on the dataset page: https://huggingface.co/datasets/fatihdx/tr-rss-haber-akisi-verisi.tabulartext-classification10K<n<100K0 likes204 downloads3mo agoHugging Face04daaain /swebench-verified-deepseek-v4-flash-failure-analysis SWE-bench Verified runs & failure analysis — DeepSeek-V4-flash (local) × mini-swe-agent Per-instance analysis of SWE-bench Verified runs of a locally-served DeepSeek-V4-flash model driven by mini-swe-agent, graded with the official SWE-bench harness. Each instance carries the full agent trajectory, a readable transcript, the submitted patch, the harness test output, deterministic metrics, and a hand-verified qualitative root-cause diagnosis. Current numbers (resolve rates… See the full description on the dataset page: https://huggingface.co/datasets/daaain/swebench-verified-deepseek-v4-flash-failure-analysis.tabulartext-generationn<1K0 likes160 downloads3mo agoHugging Face05vinod-anbalagan /chart-reasoning-verified chart-reasoning-verified Chart reasoning examples generated from an explicit latent representation. The data, the question and the answer are computed before the chart is drawn, so the image is a rendering of known ground truth rather than the source of it. No model was asked to label anything. Each row carries both a rendered chart and a text serialisation of the same chart, so the set is usable for vision-language training and for text-only language model training without… See the full description on the dataset page: https://huggingface.co/datasets/vinod-anbalagan/chart-reasoning-verified.imagevisual-question-answering1K<n<10K0 likes152 downloads8d agoHugging Face06ulamai /verified-research-reasoning-trajectories Verified Research Reasoning Trajectories for RLVR This repository is the public sample and schema repository for Ulam's research-level mathematical reasoning trajectories for reinforcement learning with verifiable rewards (RLVR), process supervision, judge training, proof criticism, and private evaluations. Ulam Verified Research Reasoning Trajectories are proof-process data for RLVR. Each record contains a normalized research problem, a golden or partial-golden proof graph… See the full description on the dataset page: https://huggingface.co/datasets/ulamai/verified-research-reasoning-trajectories.documenttext-generationn<1K3 likes148 downloads2mo agoHugging Face07ulamai /verified-math-olympiad-trajectories Verified Math Olympiad Reasoning Trajectories for RLVR This repository is the public sample and schema repository for Ulam's math olympiad reasoning trajectories for reinforcement learning with verifiable rewards (RLVR), answer-verifier evaluation, process-supervision candidates, judge training, proof criticism, and private evaluations. The goal is not merely to provide final-answer math examples. Each record is a structured olympiad reasoning object containing a normalized problem… See the full description on the dataset page: https://huggingface.co/datasets/ulamai/verified-math-olympiad-trajectories.documentn<1K2 likes122 downloads4mo agoHugging Face08cfahlgren1 /verizon-codex-trace Verizon Codex Trace A sanitized Codex session trace of GPT-5.5/Codex working through a Verizon billing and trade-in support flow. The trace includes the original JSONL session structure, assistant/user turns, tool calls, shell output, selected Computer Use browser screenshots, and the Verizon live-chat/bill context needed to understand the agent's work. What's included One saved Codex session JSONL under sessions/ Assistant messages, user messages, tool calls… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/verizon-codex-trace.tabularn<1K1 likes90 downloads5mo agoHugging Face09likaixin /APPS-verified Introduction This dataset contains verified solutions from the APPS dataset's training set. Solutions that fail to pass all the test cases are removed. Problems with no correct solution are also removed. The solutions were executed on Intel E5-2620 v3 CPUs with the execution timeout set to 10 seconds. Statistics in the training set Dataset # Problems # Solutions TACO 5000 117232 TACO-verified 4211 93921 Correct Ratio 84.22% 80.12% tabularquestion-answering1K<n<10K5 likes89 downloads2y agoHugging Face10jang1563 /sci-agent-verification-cascade Scientific Agent Verification Cascade Public evaluation fixtures and verified aggregate results for testing whether scientific claims keep their source, meaning, uncertainty, and verification requirements as they move between AI agents. This dataset accompanies the Scientific Agent Verification Cascade codebase. Version 0.2.0 contains synthetic evaluation data and aggregate-only results. It contains no raw hosted-model response, private holdout identifier, source-record… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/sci-agent-verification-cascade.tabulartext-generationn<1K0 likes84 downloads16d agoHugging Face11eewer /swerebench-traces-raw-source-verification-enhanced-20260617 SWE-rebench Raw Source Verification Enhanced 20260617 This is a private raw source dataset for building refined mini-swe-agent SFT datasets. It is intentionally not tokenized and intentionally preserves source data plus metadata for downstream filtering, masking, weighting, and audit. Do not treat every row as a clean endpoint solve. Download The full dataset directory is uploaded as a single compressed archive: hf download… See the full description on the dataset page: https://huggingface.co/datasets/eewer/swerebench-traces-raw-source-verification-enhanced-20260617.tabulartext-generationn<1K0 likes74 downloads3mo agoHugging Face12Veri-Code /ReForm-DafnyComp-Benchmark Re:Form Datasets This repository contains the datasets used in the paper Re:Form -- Reducing Human Priors in Scalable Formal Software Verification with RL in LLMs: A Preliminary Study on Dafny. The Re:Form project introduces a framework for code to specification generation using large language models, based on Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL). This work systematically explores ways to reduce human priors in scalable formal software verification by… See the full description on the dataset page: https://huggingface.co/datasets/Veri-Code/ReForm-DafnyComp-Benchmark.tabulartext-generationn<1K1 likes67 downloads5mo agoHugging Face13beneficial-ai-foundation /vericoding Vericoding A benchmark for vericoding: formally verified program synthesis Sergiu Bursuc, Theodore Ehrenborg, Shaowei Lin, Lacramioara Astefanoaei, Ionel Emilian Chiosa, Jure Kukovec, Alok Singh, Oliver Butterley, Adem Bizid, Quinn Dougherty, Miranda Zhao, Max Tan, Max Tegmark We present and test the largest benchmark for vericoding, LLM-generation of formally verified code from formal specifications - in contrast to vibe coding, which generates potentially buggy code from a… See the full description on the dataset page: https://huggingface.co/datasets/beneficial-ai-foundation/vericoding.tabulartext-generation10K<n<100K0 likes62 downloads11mo agoHugging Face14deepinquiry /verified-facts-sample-100 DeepInquiry Verified Facts (Sample-100) A 90-fact sample from the DeepInquiry verified-facts corpus. Every fact in this sample has been cross-checked against multiple structurally independent web sources, cited, dated, and confidence-scored before it entered the corpus. This is a preview sample. The full corpus (~942 approved facts as of Sept 2026, growing continuously) is available via the DeepInquiry API at deepinquiry.ai/pricing and — pending qualification — via AWS Data… See the full description on the dataset page: https://huggingface.co/datasets/deepinquiry/verified-facts-sample-100.tabularquestion-answeringn<1K0 likes62 downloads25d agoHugging Face15tarsur385 /swe-verified-gemini3-flash-trajectories SWE-bench Verified — Gemini-3-flash agent trajectories (graded, 3 samples/instance) Agent trajectories from gemini-3-flash-preview (high reasoning, temperature 0.8) run with the OpenHands agent on SWE-bench Verified, in Modal sandboxes. For each of 100 instances we sampled multiple trajectories and graded them with the SWE-bench harness; this dataset holds the 3 graded samples per instance = 296 trajectories, 198 resolved (67%). pass@1 ≈ 66/100, pass@3 (oracle) = 73/100.… See the full description on the dataset page: https://huggingface.co/datasets/tarsur385/swe-verified-gemini3-flash-trajectories.tabularothern<1K1 likes56 downloads3mo agoHugging Face16nickh007 /hw-verify hw-verify — a hardware-security verification dataset with controls Every positive example ships beside a deliberately broken counterpart, so a model or tool is graded against controls instead of against itself. Try the checker that generated this data: 🔒 hw-verify Space — paste Verilog, get a verdict, in your browser, no install. Install pip install datasets 30-second quickstart from datasets import load_dataset rtl =… See the full description on the dataset page: https://huggingface.co/datasets/nickh007/hw-verify.tabulartext-classificationn<1K0 likes52 downloads2mo agoHugging Face17parsaidp /swe-bench-verified-raw-traces-qwen3-coder SWE-bench Verified raw mini-SWE-agent traces Raw mini-SWE-agent trajectories from 20250802_mini-v1.0.0_qwen3-coder-480b-a35b-instruct for SWE-bench Verified. The raw/easy split uses exactly the 194 instance IDs from parsaidp/SWE-bench_Verified_easy. That public dataset contains SWE-bench Verified questions and Kimi-generated answers; this dataset uses only its instance IDs. The trajectory contents here are local mini-SWE-agent/Qwen traces. Files data/full.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/parsaidp/swe-bench-verified-raw-traces-qwen3-coder.tabular1K<n<10K0 likes49 downloads4mo agoHugging Face18nickh007 /hw-verify-paths hw-verify-paths ▶ Try the checker in your browser · Docs & overview Dependency graphs and witness paths for constant-time RTL analysis — the reasoning, not just the label. The companion dataset records what each design is: CONSTANT_TIME or LEAKY. This one records why. For every fixture it carries the full signal dependency graph, and for every leaky one the concrete chains of signals that carry a secret to the observation. Why witness paths and not just verdicts… See the full description on the dataset page: https://huggingface.co/datasets/nickh007/hw-verify-paths.tabulartext-classificationn<1K0 likes48 downloads2mo agoHugging Face19ctokx /regexgym-verified-traces RegexGym-Verified-Traces Reasoning traces for writing regexes from examples. Each record shows a task (some strings that should match, some that shouldn't), the teacher's chain-of-thought, and the regex it landed on. Every trace here actually solved the task's hidden holdout — the regex was run against examples the teacher never saw, and only exact solves were kept. The ground-truth regexes aren't in the released records; the model has to earn its answer. What's in… See the full description on the dataset page: https://huggingface.co/datasets/ctokx/regexgym-verified-traces.tabulartext-generationn<1K0 likes43 downloads2mo agoHugging Face20mishface123 /numeric-claim-verifier Numeric Claim Verifier (Science) — Adaption AutoScientist Programmatically verified prompt/completion pairs for scientific and statistical numeric claim verification. Labels correct — claim matches ground-truth tables wrong_direction — trend/sign reversed wrong_magnitude — right direction, wrong size (25–70% offset) unverifiable — no matching source row (real entity + absent metric) Sources Our World in Data CO₂ / Energy WHO GHO life expectancy… See the full description on the dataset page: https://huggingface.co/datasets/mishface123/numeric-claim-verifier.tabulartext-classificationn<1K0 likes39 downloads2mo agoHugging Face21nagygabor /Z3-Verified-Reasoning-Graphs Z3-Verified Constraint Reasoning Dataset 5k Baseline · Production-Ready · Zero Label Noise The Problem This Solves Most synthetic reasoning datasets only show the "happy path". Real reasoning requires knowing when to backtrack. Open-source LLMs hallucinate on constraint satisfaction problems because they are trained on fluent-sounding but logically inconsistent traces. This dataset is different: ❌ No LLM-generated reasoning — zero hallucinations, zero label noise ✅… See the full description on the dataset page: https://huggingface.co/datasets/nagygabor/Z3-Verified-Reasoning-Graphs.tabulartext-generation1K<n<10K1 likes37 downloads6mo agoHugging Face22nikitamounier /swe-bench-verified-raw-traces-qwen3-coder SWE-bench Verified raw mini-SWE-agent traces Raw mini-SWE-agent trajectories from 20250802_mini-v1.0.0_qwen3-coder-480b-a35b-instruct for SWE-bench Verified. The raw/easy split uses exactly the 194 instance IDs from parsaidp/SWE-bench_Verified_easy. That public dataset contains SWE-bench Verified questions and Kimi-generated answers; this dataset uses only its instance IDs. The trajectory contents here are local mini-SWE-agent/Qwen traces. Files data/full.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/nikitamounier/swe-bench-verified-raw-traces-qwen3-coder.tabular1K<n<10K1 likes36 downloads4mo agoHugging Face23dsb117 /brainblast-verified-footgun-corpus Brainblast — Verified SDK Footgun Corpus (free sample) The only code-training data that ships with a machine-checkable proof. Each record is a real insecure→fixed code footgun with a replayable RED→GREEN receipt: a deterministic checker fails the insecure version and passes the fixed one. You don't trust the labels — you replay the proof. This repo is a free 40-record sample (receipt-only tier). The full corpus is 4,183 proven records across 154 SDKs and 9 vulnerability classes… See the full description on the dataset page: https://huggingface.co/datasets/dsb117/brainblast-verified-footgun-corpus.tabulartext-generationn<1K0 likes36 downloads2mo agoHugging Face24Eve39570 /verified-analytics-tasks Verified Analytics Tasks 150+ mainly small analytics and data-engineering tasks. The point of the set is the answer key: every task ships its own automated checker, and every gold answer was run through that checker and scored a clean 1.0 before the task was allowed in. So the labels are more like "here's the checker, score it yourself" instead of "just trust me bro." I wanted to create a synthetic dataset inspired by this paper: Autodata: An agentic data scientist to create… See the full description on the dataset page: https://huggingface.co/datasets/Eve39570/verified-analytics-tasks.tabulartext-generationn<1K0 likes28 downloads3mo agoHugging Face25LLMTeamAkiyama /clean_cot_verification_340k元データ: https://huggingface.co/datasets/Zigeng/CoT-Verification-340k 使用したコード: https://github.com/LLMTeamAkiyama/0-data_prepare/tree/master/src/CoT-Verification-340k データ件数: 140,980 平均トークン数: 602 最大トークン数: 2,040 合計トークン数: 84,894,510 ファイル形式: JSONL ファイル分割数: 2 合計ファイルサイズ: 256.3 MB 加工内容: データセットIDの付与: データフレームのインデックスに1を加算して、base_datasets_idとして新しいID列を付与しました。 response列のフィルタリング: response列が「Yes,」で始まる行のみを保持し、それ以外の行を除外しました。 prompt列の文字長によるフィルタリング: prompt列の文字列の長さが80… See the full description on the dataset page: https://huggingface.co/datasets/LLMTeamAkiyama/clean_cot_verification_340k.tabularquestion-answering100K<n<1M0 likes26 downloads1y agoHugging Face26maxRyeery /VeriSoftBench VeriSoftBench VeriSoftBench is a benchmark for evaluating neural theorem provers on software verification tasks in Lean 4. The dataset contains 500 theorem-proving tasks drawn from 23 real-world Lean 4 repositories spanning compiler verification, type system formalization, applied verification (zero-knowledge proofs, smart contracts), semantic frameworks, and more. 📄 Paper (arXiv): https://arxiv.org/html/2602.18307v1💻 Full benchmark + pipeline + setup:… See the full description on the dataset page: https://huggingface.co/datasets/maxRyeery/VeriSoftBench.tabularn<1K0 likes25 downloads7mo agoHugging Face27daviddata1 /bullet-swebench-verified Bullet on SWE-bench Verified — 479/500 = 95.8% Results for the Bullet coding agent on all 500 instances of SWE-bench Verified, graded by the official swebench.harness.run_evaluation scorer. Every instance was attempted and graded; there are no empty patches. 479 / 500 resolved = 95.8% Run with the Bullet harness on gpt-5.6-sol at high reasoning effort, one attempt per instance. By repository repository resolved django 223/231 96.5% sympy 73/75 97.3%… See the full description on the dataset page: https://huggingface.co/datasets/daviddata1/bullet-swebench-verified.tabulartext-generationn<1K0 likes25 downloads2mo agoHugging Face28YSenseAI /verifimind-peas-eval VerifiMind-PEAS Evaluation Dataset DOI: 10.5281/zenodo.21276884 · License: MIT · Version: 1.0 A human-annotated evaluation dataset for measuring the performance of the VerifiMind-PEAS multi-agent epistemic verification system. Ground-truth labels were assigned by a single domain-expert annotator whose final verdicts and confidence were human judgments; disclosed LLM comprehension assistance was used during annotation (see Annotation Protocol — transparency is a design commitment… See the full description on the dataset page: https://huggingface.co/datasets/YSenseAI/verifimind-peas-eval.tabulartext-classificationn<1K0 likes17 downloads3mo agoHugging Face29dispatchAI /all-verified-benchmarks All Verified Benchmarks Real CPU benchmark data for ALL 31 working dispatchAI models. Zero broken models. Zero partial models. All verified. 🚀 dispatchAI tabularn<1K0 likes16 downloads3mo agoHugging Face30Pittro /verifiable-ai-provenance-bench Verifiable AI Provenance Bench (TTTPS) 25 real timestamp-provenance receipts generated on 2026-08-04 by calling the live KPP (Kenosian Protocol Platform) provenance API (POST /v1/anchor, POST /v1/verify), which implements the TTTPS (Time-Token Tamper-evident Provenance Seal) scheme. Each row is one real API round trip: a content_hash was submitted to /v1/anchor, the returned receipt_id was then submitted to /v1/verify, and both raw responses are recorded. This dataset was built… See the full description on the dataset page: https://huggingface.co/datasets/Pittro/verifiable-ai-provenance-bench.tabularothern<1K1 likes15 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.