CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01shhu2001 /SciCode-Verified SciCode-Verified SciCode-Verified is the corrected, human-verified release of the SciCode scientific-code-generation benchmark. A problem-by-problem audit identified 263 defects in the 65-problem SciCode test split and corrected every confirmable defect. The released evaluation set contains 64 main problems and 287 scored subproblems; one original problem is excluded because its specification does not determine a unique, verifiable answer. Paper: SciCode-Verified: How Benchmark… See the full description on the dataset page: https://huggingface.co/datasets/shhu2001/SciCode-Verified.texttext-generationn<1K1 likes1.7k downloads2mo agoHugging Face02likaixin /TACO-verified Introduction This dataset contains verified solutions from the TACO dataset's training set. Solutions that fail to pass all the test cases are removed. Problems with no correct solution are also removed. The solutions were executed on Intel E5-2620 v3 CPUs with the execution timeout set to 10 seconds. Statistics in the training set Dataset # Problems # Solutions TACO 25443 1468722 TACO-verified 12898 1043251 Correct Ratio 50.69 % 71.03 %… See the full description on the dataset page: https://huggingface.co/datasets/likaixin/TACO-verified.textquestion-answering10K<n<100K20 likes1.4k downloads1y agoHugging Face03harithoppil /terminal-bench-2-verified Terminal-Bench 2.0 Verified: Instruction & Environment Fix Version 中文版本 We conducted a comprehensive review of the entire Terminal-Bench 2.0 dataset and identified various issues. Both GLM-5 and Step 3.5-Flash were evaluated using this verified version. This modified version addresses environment and instruction issues we discovered in Terminal-Bench 2.0. It includes two types of fixes: Environment Fixes: Updated Dockerfiles and instructions to support Claude Code Agent runtime… See the full description on the dataset page: https://huggingface.co/datasets/harithoppil/terminal-bench-2-verified.documenttext-generationn<1K2 likes1.3k downloads5mo agoHugging Face04FreedomIntelligence /medical-o1-verifiable-problem Introduction This dataset features open-ended medical problems designed to improve LLMs' medical reasoning. Each entry includes a open-ended question and a ground-truth answer based on challenging medical exams. The verifiable answers enable checking LLM outputs, refining their reasoning processes. For details, see our paper and GitHub repository. Citation If you find our data useful, please consider citing our work! @misc{chen2024huatuogpto1medicalcomplexreasoning… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/medical-o1-verifiable-problem.textquestion-answering10K<n<100K124 likes776 downloads2y agoHugging Face05salimayed /verified-defi-datasets Verified Solana Sealevel & Anchor Program Optimization Fine-Tuning Corpus Dataset Description High-density, verified AI fine-tuning dataset in ALPACA format. Domain: Solana Sealevel & Anchor Program Optimization Verified Records: 3 Estimated Tokens: 339 Quality QA Score: 99.0% Monetization Status: Direct Zero-Gas Web3 & HuggingFace Distribution texttext-generationn<1K0 likes520 downloads14d agoHugging Face06pankajmathur /nemotron-nano-30b-miniswe-swebench-verified Nemotron Nano 30B + mini-swe-agent SWE-bench Verified Trajectories Agent trajectories from running NVIDIA Nemotron 3 Nano 30B A3B (MoE, 8B active params) on SWE-bench Verified using mini-swe-agent. ⚠️ Incomplete Run This benchmark was terminated early due to poor performance. The model struggled with the agentic coding task. Model Information Attribute Value Model NVIDIA Nemotron 3 Nano 30B A3B Architecture MoE (30B total, 8B active) Serving vLLM… See the full description on the dataset page: https://huggingface.co/datasets/pankajmathur/nemotron-nano-30b-miniswe-swebench-verified.texttext-generationn<1K0 likes357 downloads9mo agoHugging Face07Lego-X /Lego-RL-SWE-Bench-Verified Lego-RL-SWE-Bench-Verified The 500 SWE-bench Verified instances as ready-to-run harbor RL environments — the exact evaluation set behind every SWE-bench Verified number in LEGO-RL, packaged the same way as the training set Lego-X/Lego-RL-2699 so one trainer reads both. Two parallel views of the same 500 instances: View Path What it is Official SWE-bench records swebench_verified_official_500/ The upstream princeton-nlp/SWE-bench_Verified rows, verbatim Harbor RL… See the full description on the dataset page: https://huggingface.co/datasets/Lego-X/Lego-RL-SWE-Bench-Verified.texttext-generationn<1K0 likes255 downloads1mo agoHugging Face08Fujitsu-FRE /MAPS_Verified Dataset Card for Multilingual Benchmark for Global Agent Performance and Security This is the first Multilingual Agentic AI Benchmark for evaluating agentic AI systems across different languages and diverse tasks. Benchmark enables systematic analysis of how agents perform under multilingual conditions. This dataset contains 550 instances for GAIA, 660 instances for ASB, 737 instances for Maths, and 1100 instances for SWE. Each task was translated into 10 target languages resulting… See the full description on the dataset page: https://huggingface.co/datasets/Fujitsu-FRE/MAPS_Verified.texttext-generation1K<n<10K3 likes233 downloads8mo agoHugging Face09Mungus451 /verified_wiki_historian_the_beatles_anthology_dataset_active Verified-Wiki-Historian: The Beatles Anthology Verified-Wiki-Historian (The Beatles Anthology) is a refined, citation-grounded instruction dataset for Beatles-specific historical question answering, summarization, and supervised fine-tuning. This dataset is a cleaned and rebuilt refinement of: Mungus451/verified_wiki_historian_the_beatles_anthology_dataset_active The current release contains 4,000 instruction records focused on Beatles history, recording sessions, release… See the full description on the dataset page: https://huggingface.co/datasets/Mungus451/verified_wiki_historian_the_beatles_anthology_dataset_active.textquestion-answering1K<n<10K1 likes224 downloads3mo agoHugging Face10Doc2Feat-bench /Doc2Feat-bench_Verified Dataset Summary NoCode-bench Verified is subset of NoCode-bench, a dataset that tests systems’ no-code feature addition ability automatically. Languages The text of the dataset is primarily English, but we make no effort to filter or otherwise clean based on language type. Dataset Structure An example of a SWE-bench datum is as follows: repo: (str) - The repository owner/name identifier from GitHub. instance_id: (str) - A formatted instance… See the full description on the dataset page: https://huggingface.co/datasets/Doc2Feat-bench/Doc2Feat-bench_Verified.texttext-generationn<1K1 likes221 downloads1y agoHugging Face11HayleyZhou1113 /VeriTime VeriTime: Time Series Reasoning via Process-Verifiable Thinking Data Synthesis and Scheduling for Tailored LLM Reasoning This is the dataset associated with our paper: Time Series Reasoning via Process-Verifiable Thinking Data Synthesis and Scheduling for Tailored LLM Reasoning Jiahui Zhou, Dan Li, Boxin Li, Xiao Zhang, Erli Meng, Lin Li, Zhuomin Chen, Jian Lou, See-Kiong Ng ICML 2026 &nbsp;|&nbsp; Paper &nbsp; Dataset Construction Pipeline: TSRgen TSRgen is an… See the full description on the dataset page: https://huggingface.co/datasets/HayleyZhou1113/VeriTime.texttime-series-forecasting1K<n<10K0 likes217 downloads18d agoHugging Face12vericava /sft-tool-calling-structured-output-v1 vericava/sft-tool-calling-structured-output-v1 Dataset to train (SFT) 3-20B LLMs for tool calling and structured outputs/classifications. Includes contents in English as well as some Japanese. texttext-classification100K<n<1M2 likes158 downloads8mo agoHugging Face13ulamai /verified-research-reasoning-trajectories Verified Research Reasoning Trajectories for RLVR This repository is the public sample and schema repository for Ulam's research-level mathematical reasoning trajectories for reinforcement learning with verifiable rewards (RLVR), process supervision, judge training, proof criticism, and private evaluations. Ulam Verified Research Reasoning Trajectories are proof-process data for RLVR. Each record contains a normalized research problem, a golden or partial-golden proof graph… See the full description on the dataset page: https://huggingface.co/datasets/ulamai/verified-research-reasoning-trajectories.documenttext-generationn<1K3 likes156 downloads2mo agoHugging Face14daaain /swebench-verified-deepseek-v4-flash-failure-analysis SWE-bench Verified runs & failure analysis — DeepSeek-V4-flash (local) × mini-swe-agent Per-instance analysis of SWE-bench Verified runs of a locally-served DeepSeek-V4-flash model driven by mini-swe-agent, graded with the official SWE-bench harness. Each instance carries the full agent trajectory, a readable transcript, the submitted patch, the harness test output, deterministic metrics, and a hand-verified qualitative root-cause diagnosis. Current numbers (resolve rates… See the full description on the dataset page: https://huggingface.co/datasets/daaain/swebench-verified-deepseek-v4-flash-failure-analysis.tabulartext-generationn<1K0 likes153 downloads3mo agoHugging Face15namakoo /idfu-verified-code IDFU Code Negative Dataset — Free Preview A curated dataset of Python code samples that failed execution-based validation, designed for training reward models, DPO rejected-side pairs, and error-detection classifiers. Free 100-sample preview; paid full versions available separately. What's inside this preview 100 unique Python samples, all AST-validated 19 CS domains represented (MCMC, FFT, distributed consensus, ZKP, formal methods, HFT microstructure, and more)… See the full description on the dataset page: https://huggingface.co/datasets/namakoo/idfu-verified-code.texttext-generationn<1K1 likes124 downloads5mo agoHugging Face16protogonos /verified-tool-use-dataset Verified tool-use trajectories for LLM agents This was a time-boxed experiment by an autonomous agent (Protogonos), now concluded. Nothing here is offered for sale or for hire, and no payment is accepted. Multi-turn function-calling conversations for training and evaluating tool-using agents — 48 trajectories across 16 domains, with every tool call checked against its tool's JSON-Schema. The free sample in this repo is a real slice of the full set: the viewer above renders it… See the full description on the dataset page: https://huggingface.co/datasets/protogonos/verified-tool-use-dataset.texttext-generationn<1K1 likes120 downloads28d agoHugging Face17F-A-I-L /kodcode-verified-python-235k KodCode-Verified Python — 234,555 execution-verified Python SFT rows One row per problem. Every assistant turn is code that passed its own unit tests when actually run — real pytest against KodCode-V1's tests, in a pinned interpreter, in a sandboxed subprocess. No LLM judge, no heuristic filter, no model-generated answers. Unlike the v3 release this supersedes, the corpus is deduplicated, decontaminated against HumanEval/MBPP, and stripped of rows whose tests cannot constrain… See the full description on the dataset page: https://huggingface.co/datasets/F-A-I-L/kodcode-verified-python-235k.texttext-generation100K<n<1M0 likes115 downloads10d agoHugging Face18giggiovpg /android-kotlin-compose-compiler-verified Qwandroid — Compiler-Verified Modern Android (Kotlin + Jetpack Compose) Dataset 5,777 SFT examples + 150 held-out eval + 8,027 DPO preference pairs. Every SFT row was actually compiled — not LLM-approved, not heuristically filtered. A subset was verified behaviorally by running JUnit tests. Built to fine-tune small models into focused Android specialists rather than general-purpose coders. Why this exists Android code in pretraining corpora is largely stale —… See the full description on the dataset page: https://huggingface.co/datasets/giggiovpg/android-kotlin-compose-compiler-verified.texttext-generation10K<n<100K1 likes99 downloads1mo agoHugging Face19likaixin /APPS-verified Introduction This dataset contains verified solutions from the APPS dataset's training set. Solutions that fail to pass all the test cases are removed. Problems with no correct solution are also removed. The solutions were executed on Intel E5-2620 v3 CPUs with the execution timeout set to 10 seconds. Statistics in the training set Dataset # Problems # Solutions TACO 5000 117232 TACO-verified 4211 93921 Correct Ratio 84.22% 80.12% tabularquestion-answering1K<n<10K5 likes89 downloads2y agoHugging Face20Nellyw888 /VeriReason-RTL-Coder_7b_reasoning_tb_simple Verireason-RTL-Coder_7b_reasoning_tb_simple For implementation details, visit our GitHub repository: VeriReason and our page Check out our paper: VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation Update Log 2025.05.17: Initial release of Nellyw888/Verireason-RTL-Coder_7b_reasoning_tb_simple Project Description This study introduces VeriReason, a novel approach utilizing reinforcement learning with… See the full description on the dataset page: https://huggingface.co/datasets/Nellyw888/VeriReason-RTL-Coder_7b_reasoning_tb_simple.texttext-generationn<1K0 likes88 downloads1y agoHugging Face21jang1563 /sci-agent-verification-cascade Scientific Agent Verification Cascade Public evaluation fixtures and verified aggregate results for testing whether scientific claims keep their source, meaning, uncertainty, and verification requirements as they move between AI agents. This dataset accompanies the Scientific Agent Verification Cascade codebase. Version 0.2.0 contains synthetic evaluation data and aggregate-only results. It contains no raw hosted-model response, private holdout identifier, source-record… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/sci-agent-verification-cascade.tabulartext-generationn<1K0 likes87 downloads16d agoHugging Face22Nellyw888 /VeriReason-RTL-Coder_7b_reasoning_tb Verireason-RTL-Coder_7b_reasoning_tb For implementation details, visit our GitHub repository: VeriReason and our page Check out our paper: VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation Update Log 2025.05.17: Initial release of Nellyw888/Verireason-RTL-Coder_7b_reasoning_tb Project Description This study introduces VeriReason, a novel approach utilizing reinforcement learning with testbench feedback to… See the full description on the dataset page: https://huggingface.co/datasets/Nellyw888/VeriReason-RTL-Coder_7b_reasoning_tb.texttext-generation1K<n<10K4 likes83 downloads1y agoHugging Face23Veri-Code /ReForm-Python2Dafny-Dataset Re:Form Datasets This repository contains the datasets associated with the paper "Re:Form -- Reducing Human Priors in Scalable Formal Software Verification with RL in LLMs: A Preliminary Study on Dafny". The paper introduces a framework that leverages Reinforcement Learning (RL) within Large Language Models (LLMs) to reduce reliance on human-annotated priors for formal software verification. The datasets provided are integral for training and evaluating models in formal language… See the full description on the dataset page: https://huggingface.co/datasets/Veri-Code/ReForm-Python2Dafny-Dataset.texttext-generation10K<n<100K2 likes80 downloads5mo agoHugging Face24eewer /swerebench-traces-raw-source-verification-enhanced-20260617 SWE-rebench Raw Source Verification Enhanced 20260617 This is a private raw source dataset for building refined mini-swe-agent SFT datasets. It is intentionally not tokenized and intentionally preserves source data plus metadata for downstream filtering, masking, weighting, and audit. Do not treat every row as a clean endpoint solve. Download The full dataset directory is uploaded as a single compressed archive: hf download… See the full description on the dataset page: https://huggingface.co/datasets/eewer/swerebench-traces-raw-source-verification-enhanced-20260617.tabulartext-generationn<1K0 likes77 downloads3mo agoHugging Face25TheJackBright /verisci-verified-science-math-code VeriSci Verified Science Math Code Verifier-grounded dataset for the Adaption AutoScientist Challenge Part 2, targeting the Science category with secondary Math and Code coverage. Summary VeriSci trains models to solve scientific computations, finite-difference PDE updates, numerical ODE steps, unit-checked mechanics, thermodynamics, circuits, chemistry stoichiometry, molarity, unit conversion, vector decomposition, two-point linear modeling, small Python… See the full description on the dataset page: https://huggingface.co/datasets/TheJackBright/verisci-verified-science-math-code.texttext-generation1K<n<10K0 likes72 downloads2mo agoHugging Face26theepicflyer /openrtlset-permissive-verified openrtlset-permissive-verified A permissive-filtered, machine-verified subset of ESCAD/OpenRTLSet. Every record in this dataset compiles. Each completion was elaborated and linted with Verilator (--lint-only -Wall, warnings fatal) and admitted only on a clean run. Nothing here is graded by an LLM judge. Contents 13,625 records drawn from 2,174 distinct upstream repositories Task: specification + pinned module interface -> complete Verilog module Upstream… See the full description on the dataset page: https://huggingface.co/datasets/theepicflyer/openrtlset-permissive-verified.texttext-generation10K<n<100K0 likes70 downloads2mo agoHugging Face27Veri-Code /ReForm-DafnyComp-Benchmark Re:Form Datasets This repository contains the datasets used in the paper Re:Form -- Reducing Human Priors in Scalable Formal Software Verification with RL in LLMs: A Preliminary Study on Dafny. The Re:Form project introduces a framework for code to specification generation using large language models, based on Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL). This work systematically explores ways to reduce human priors in scalable formal software verification by… See the full description on the dataset page: https://huggingface.co/datasets/Veri-Code/ReForm-DafnyComp-Benchmark.tabulartext-generationn<1K1 likes67 downloads5mo agoHugging Face28AbiralArch /hardware-verilogeval-v2 hardware-verilogeval-v2 VerilogEval v2 - 471 Verilog evaluation problems Dataset Overview This dataset is part of a comprehensive collection of hardware design datasets for training and evaluating LLMs on Verilog/SystemVerilog code generation and hardware design tasks. Files verilog_eval_problems.json: 471 VerilogEval v2 problems Usage from datasets import load_dataset # Load the dataset dataset = load_dataset('AbiralArch/hardware-verilogeval-v2')… See the full description on the dataset page: https://huggingface.co/datasets/AbiralArch/hardware-verilogeval-v2.texttext-generationn<1K0 likes66 downloads1y agoHugging Face29LDJnr /Verified-Camel This is the Official Verified Camel dataset. Just over 100 verified examples, and many more coming soon! Comprised of over 100 highly filtered and curated examples from specific portions of CamelAI stem datasets. These examples are verified to be true by experts in the specific related field, with atleast a bachelors degree in the subject. Roughly 30-40% of the originally curated data from CamelAI was found to have atleast minor errors and/or incoherent questions(as determined… See the full description on the dataset page: https://huggingface.co/datasets/LDJnr/Verified-Camel.textquestion-answeringn<1K43 likes60 downloads2y agoHugging Face30beneficial-ai-foundation /vericoding Vericoding A benchmark for vericoding: formally verified program synthesis Sergiu Bursuc, Theodore Ehrenborg, Shaowei Lin, Lacramioara Astefanoaei, Ionel Emilian Chiosa, Jure Kukovec, Alok Singh, Oliver Butterley, Adem Bizid, Quinn Dougherty, Miranda Zhao, Max Tan, Max Tegmark We present and test the largest benchmark for vericoding, LLM-generation of formally verified code from formal specifications - in contrast to vibe coding, which generates potentially buggy code from a… See the full description on the dataset page: https://huggingface.co/datasets/beneficial-ai-foundation/vericoding.tabulartext-generation10K<n<100K0 likes59 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.