CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01abhinavpola /tau2-bench-verified-airline tau2-bench-verified — airline domain (mirror) Mirror of the airline domain from amazon-agi/tau2-bench-verified (MIT License), pinned at commit 864350a8971a8f8ee9e7b8472e2edc380a806b0c. Re-hosted for the OpenRouter native TypeScript benchmark harness so it can fetch the verified airline tasks + environment DB at runtime. Contents tasks/test.jsonl — 50 verified airline tasks. Each row has a single task_json string column holding one verbatim tau2 v2 task object (id… See the full description on the dataset page: https://huggingface.co/datasets/abhinavpola/tau2-bench-verified-airline.textn<1K0 likes4.3k downloads3mo agoHugging Face02razzant /ouroboros-osworld-verified-opus5 Ouroboros on OSWorld-Verified: 90.69%, the highest result reported to date Status: Self-reported result over all 361 tasks. The official per-task scores, prompts, manifests and feasibility records are public here, together with every acting task record that the run produced. Start here Result 90.69% (327.39 / 361) Model anthropic/claude-opus-5 Method Screenshot only, one rollout, 100 policy turns Exact evidence f52ebf2 and evidence.json… See the full description on the dataset page: https://huggingface.co/datasets/razzant/ouroboros-osworld-verified-opus5.textothern<1K1 likes2.1k downloads2mo agoHugging Face03shhu2001 /SciCode-Verified SciCode-Verified SciCode-Verified is the corrected, human-verified release of the SciCode scientific-code-generation benchmark. A problem-by-problem audit identified 263 defects in the 65-problem SciCode test split and corrected every confirmable defect. The released evaluation set contains 64 main problems and 287 scored subproblems; one original problem is excluded because its specification does not determine a unique, verifiable answer. Paper: SciCode-Verified: How Benchmark… See the full description on the dataset page: https://huggingface.co/datasets/shhu2001/SciCode-Verified.texttext-generationn<1K1 likes1.7k downloads2mo agoHugging Face04razzant /ouroboros-osworld-verified-sonnet46 Ouroboros on OSWorld-Verified: best published result on Claude Sonnet 4.6 Status: Self-reported result over all 361 task packages. Prompts, manifests, outcomes and feasibility records are public for every task. The official evaluator produced 360 score files; one unscored task is counted as zero. Start here Result 83.27% (300.59 / 361) Model anthropic/claude-sonnet-4.6 Method Screenshot only, one rollout, 100 policy turns Exact evidence… See the full description on the dataset page: https://huggingface.co/datasets/razzant/ouroboros-osworld-verified-sonnet46.textothern<1K0 likes1.6k downloads2mo agoHugging Face05opencompass /SWEBench-Pro-Verified SWE-Bench Pro Verified: Anti-hacking & Task refinement SWE-Bench Pro has emerged as a standard benchmark for evaluating software engineering agents on challenging repository-level tasks. However, our analysis work show that its evaluation is undermined by two sources of unreliability: reward hacking, enabled by leakage of gold solutions or hidden evaluation information, and task quality issues, including misleading problem statements and improperly scoped tests. These issues can… See the full description on the dataset page: https://huggingface.co/datasets/opencompass/SWEBench-Pro-Verified.tabularn<1K3 likes1.6k downloads16d agoHugging Face06likaixin /TACO-verified Introduction This dataset contains verified solutions from the TACO dataset's training set. Solutions that fail to pass all the test cases are removed. Problems with no correct solution are also removed. The solutions were executed on Intel E5-2620 v3 CPUs with the execution timeout set to 10 seconds. Statistics in the training set Dataset # Problems # Solutions TACO 25443 1468722 TACO-verified 12898 1043251 Correct Ratio 50.69 % 71.03 %… See the full description on the dataset page: https://huggingface.co/datasets/likaixin/TACO-verified.textquestion-answering10K<n<100K20 likes1.4k downloads1y agoHugging Face07harithoppil /terminal-bench-2-verified Terminal-Bench 2.0 Verified: Instruction & Environment Fix Version 中文版本 We conducted a comprehensive review of the entire Terminal-Bench 2.0 dataset and identified various issues. Both GLM-5 and Step 3.5-Flash were evaluated using this verified version. This modified version addresses environment and instruction issues we discovered in Terminal-Bench 2.0. It includes two types of fixes: Environment Fixes: Updated Dockerfiles and instructions to support Claude Code Agent runtime… See the full description on the dataset page: https://huggingface.co/datasets/harithoppil/terminal-bench-2-verified.documenttext-generationn<1K2 likes1.2k downloads5mo agoHugging Face082077AIDataFoundation /VeriWebVeriWeb: Verifiable Long-Chain Web Benchmark for Agentic Information-Seeking [!NOTE] This project was originally named VeriGUI. As our initial data collection focused on web-based tasks that primarily involve information-seeking rather than GUI interaction, we now define this part as the standalone VeriWeb benchmark, while desktop and other GUI-oriented scenarios will be released as a separate benchmark (in progress). We apologize for any resulting confusion. Overview… See the full description on the dataset page: https://huggingface.co/datasets/2077AIDataFoundation/VeriWeb.textn<1K27 likes1.1k downloads8mo agoHugging Face09FreedomIntelligence /medical-o1-verifiable-problem Introduction This dataset features open-ended medical problems designed to improve LLMs' medical reasoning. Each entry includes a open-ended question and a ground-truth answer based on challenging medical exams. The verifiable answers enable checking LLM outputs, refining their reasoning processes. For details, see our paper and GitHub repository. Citation If you find our data useful, please consider citing our work! @misc{chen2024huatuogpto1medicalcomplexreasoning… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/medical-o1-verifiable-problem.textquestion-answering10K<n<100K124 likes776 downloads2y agoHugging Face10lmms-lab-eval /HLE-Verified HLE-Verified (HF-native) This dataset is a Hugging Face-native conversion of skylenage/HLE-Verified at revision becad9f339dfce27df0ebb38e55dabef12ca5735. Why this exists The source dataset stores nested verification fields with mixed runtime types (for example 0/1/"uncertain"), which breaks strict Arrow JSON parsing in datasets.load_dataset. This converted dataset normalizes those fields and publishes split-ready JSONL files for direct use in lmms_eval. Split… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-eval/HLE-Verified.text1K<n<10K0 likes522 downloads7mo agoHugging Face11salimayed /verified-defi-datasets Verified Solana Sealevel & Anchor Program Optimization Fine-Tuning Corpus Dataset Description High-density, verified AI fine-tuning dataset in ALPACA format. Domain: Solana Sealevel & Anchor Program Optimization Verified Records: 3 Estimated Tokens: 339 Quality QA Score: 99.0% Monetization Status: Direct Zero-Gas Web3 & HuggingFace Distribution texttext-generationn<1K0 likes516 downloads14d agoHugging Face12SZLHOLDINGS /k-verify-benchmark-v1 Part of the SZL Holdings governed estate — claims are designed to carry checkable receipts. Verification proves integrity & origin, never accuracy or performance. K-Verify Benchmark v1 Author: Yachay / SZL Holdings · Version: 1.0.0 · Items: 100 K-Verify measures whether an AI's claimed factual answer is verifiable via a receipt chain — not just whether it is correct. It is the first benchmark we know of that scores provenance and honest refusal as… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/k-verify-benchmark-v1.textquestion-answeringn<1K0 likes380 downloads26d agoHugging Face13pankajmathur /nemotron-nano-30b-miniswe-swebench-verified Nemotron Nano 30B + mini-swe-agent SWE-bench Verified Trajectories Agent trajectories from running NVIDIA Nemotron 3 Nano 30B A3B (MoE, 8B active params) on SWE-bench Verified using mini-swe-agent. ⚠️ Incomplete Run This benchmark was terminated early due to poor performance. The model struggled with the agentic coding task. Model Information Attribute Value Model NVIDIA Nemotron 3 Nano 30B A3B Architecture MoE (30B total, 8B active) Serving vLLM… See the full description on the dataset page: https://huggingface.co/datasets/pankajmathur/nemotron-nano-30b-miniswe-swebench-verified.texttext-generationn<1K0 likes368 downloads9mo agoHugging Face14rasinmuhammed /verified-sql-rewards Verified SQL Rewards A text-to-SQL corpus where every reward carries a machine-checkable proof that it is correct. Questions, all independently verified 109,306 Databases 1,400 across 7 schema families Tables / data rows 4,400 / ~19.6 million Unique (question, answer) pairs 102,764 Candidates refused and published 12,150 Verification pass rate 90.00% Trivial baseline (always answer 0) 1.83% Each item is a natural-language question, a gold SQL query… See the full description on the dataset page: https://huggingface.co/datasets/rasinmuhammed/verified-sql-rewards.texttable-question-answering100K<n<1M0 likes291 downloads22d agoHugging Face15csoai /gspc-verify GSPC Verify This dataset carries inputs and pointers for verifying published GSPC evidence. It is a reader surface, not a mill, ranking engine, or certificate. The live board remains the authority: GET https://councilof.ai/api/gspc. Lid: 23 axes measured · 14 model fleets · 3 public leader scores · 9 fact runs · TIE is TIE · not a certificate. Do not freeze a MEASURED/UNMEASURED table here. Re-read the live endpoints; fetch failure is UNCHECKABLE, never zero. The verifier… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-verify.textothern<1K0 likes260 downloads4d agoHugging Face16Mungus451 /verified_wiki_historian_the_beatles_anthology_dataset_active Verified-Wiki-Historian: The Beatles Anthology Verified-Wiki-Historian (The Beatles Anthology) is a refined, citation-grounded instruction dataset for Beatles-specific historical question answering, summarization, and supervised fine-tuning. This dataset is a cleaned and rebuilt refinement of: Mungus451/verified_wiki_historian_the_beatles_anthology_dataset_active The current release contains 4,000 instruction records focused on Beatles history, recording sessions, release… See the full description on the dataset page: https://huggingface.co/datasets/Mungus451/verified_wiki_historian_the_beatles_anthology_dataset_active.textquestion-answering1K<n<10K1 likes247 downloads3mo agoHugging Face17Lego-X /Lego-RL-SWE-Bench-Verified Lego-RL-SWE-Bench-Verified The 500 SWE-bench Verified instances as ready-to-run harbor RL environments — the exact evaluation set behind every SWE-bench Verified number in LEGO-RL, packaged the same way as the training set Lego-X/Lego-RL-2699 so one trainer reads both. Two parallel views of the same 500 instances: View Path What it is Official SWE-bench records swebench_verified_official_500/ The upstream princeton-nlp/SWE-bench_Verified rows, verbatim Harbor RL… See the full description on the dataset page: https://huggingface.co/datasets/Lego-X/Lego-RL-SWE-Bench-Verified.texttext-generationn<1K0 likes235 downloads1mo agoHugging Face18Fujitsu-FRE /MAPS_Verified Dataset Card for Multilingual Benchmark for Global Agent Performance and Security This is the first Multilingual Agentic AI Benchmark for evaluating agentic AI systems across different languages and diverse tasks. Benchmark enables systematic analysis of how agents perform under multilingual conditions. This dataset contains 550 instances for GAIA, 660 instances for ASB, 737 instances for Maths, and 1100 instances for SWE. Each task was translated into 10 target languages resulting… See the full description on the dataset page: https://huggingface.co/datasets/Fujitsu-FRE/MAPS_Verified.texttext-generation1K<n<10K3 likes233 downloads8mo agoHugging Face19YefanZhou98 /LLMVerify-Verifier LLMVerify-Verifier Verification results dataset for the paper "Variation in Verification: Understanding Verification Dynamics in Large Language Models", accepted at ICLR 2026 (arXiv:2509.17995). This dataset contains the binary verdicts and chain-of-thought verification reasoning produced by 15 verifier models judging candidate solutions from 15 generator models across three task domains. It supports systematic analysis of how problem difficulty, generator capability, and verifier… See the full description on the dataset page: https://huggingface.co/datasets/YefanZhou98/LLMVerify-Verifier.tabular1M<n<10M1 likes232 downloads5mo agoHugging Face20Doc2Feat-bench /Doc2Feat-bench_Verified Dataset Summary NoCode-bench Verified is subset of NoCode-bench, a dataset that tests systems’ no-code feature addition ability automatically. Languages The text of the dataset is primarily English, but we make no effort to filter or otherwise clean based on language type. Dataset Structure An example of a SWE-bench datum is as follows: repo: (str) - The repository owner/name identifier from GitHub. instance_id: (str) - A formatted instance… See the full description on the dataset page: https://huggingface.co/datasets/Doc2Feat-bench/Doc2Feat-bench_Verified.texttext-generationn<1K1 likes221 downloads1y agoHugging Face21HayleyZhou1113 /VeriTime VeriTime: Time Series Reasoning via Process-Verifiable Thinking Data Synthesis and Scheduling for Tailored LLM Reasoning This is the dataset associated with our paper: Time Series Reasoning via Process-Verifiable Thinking Data Synthesis and Scheduling for Tailored LLM Reasoning Jiahui Zhou, Dan Li, Boxin Li, Xiao Zhang, Erli Meng, Lin Li, Zhuomin Chen, Jian Lou, See-Kiong Ng ICML 2026 &nbsp;|&nbsp; Paper &nbsp; Dataset Construction Pipeline: TSRgen TSRgen is an… See the full description on the dataset page: https://huggingface.co/datasets/HayleyZhou1113/VeriTime.texttime-series-forecasting1K<n<10K0 likes219 downloads18d agoHugging Face22ZJU-REAL /VerifyBench VerifyBench: Benchmarking Reference-based Reward Systems for Large Language Models Yuchen Yan1,2,*, Jin Jiang2,3, Zhenbang Ren1,4, Yijun Li1, Xudong Cai1, Yang Liu2, Xin Xu5, Mengdi Zhang2, Jian Shao1,†, Yongliang Shen1,†, Jun Xiao1, Yueting Zhuang1 1Zhejiang University 2Meituan Group 3Peking university 4University of Electronic Science and Technology of China 5The Hong Kong University of Science and Technology ICLR 2026… See the full description on the dataset page: https://huggingface.co/datasets/ZJU-REAL/VerifyBench.texttext-ranking1K<n<10K17 likes209 downloads7mo agoHugging Face23fatihdx /tr-rss-haber-akisi-verisi TR-RSS Haber Akışı Verisi TL;DR — Bu veri seti, Türkiye odaklı haber/RSS akışlarından toplanan kayıtları; mükerrerlik, spam, reklam, amaç dışı kategori, yurtdışı odak ve editoryal çerçeve yoğunluğu açısından katmanlı kalite kontrolden geçirerek erken sinyal üretimine uygun hâle getirir. Doğrulama kararı / verdict üretmez; ClaimReview ve dezenformasyon araştırmaları için upstream izleme ve kaynak önceliklendirme katmanı olarak tasarlanmıştır. Ölçek: 307.800 öğe incelendi →… See the full description on the dataset page: https://huggingface.co/datasets/fatihdx/tr-rss-haber-akisi-verisi.tabulartext-classification10K<n<100K0 likes204 downloads3mo agoHugging Face24dislove /evidence-backed-authority-verification Evidence-Backed Authority Verification for Autonomous Agents Measuring and Governing Root-Equivalent Execution Paths A verifier that was asked whether an autonomous agent could reach root on its host, could not prove that it couldn't, and said so. This repository is the paper, the verifier, and every artifact the paper's numbers are computed from. Verdict BLOCKED_ROOT_EQUIVALENCE_DOCKER — exclusivity not proven Paper 39 pages, 17,302 words, 40 references —… See the full description on the dataset page: https://huggingface.co/datasets/dislove/evidence-backed-authority-verification.textn<1K0 likes185 downloads2mo agoHugging Face25SWE-Gym /MoatlessTools-Agent-Verifier-Train-Datatext1K<n<10K0 likes178 downloads2y agoHugging Face26vericava /sft-tool-calling-structured-output-v1 vericava/sft-tool-calling-structured-output-v1 Dataset to train (SFT) 3-20B LLMs for tool calling and structured outputs/classifications. Includes contents in English as well as some Japanese. texttext-classification100K<n<1M2 likes171 downloads8mo agoHugging Face27daaain /swebench-verified-deepseek-v4-flash-failure-analysis SWE-bench Verified runs & failure analysis — DeepSeek-V4-flash (local) × mini-swe-agent Per-instance analysis of SWE-bench Verified runs of a locally-served DeepSeek-V4-flash model driven by mini-swe-agent, graded with the official SWE-bench harness. Each instance carries the full agent trajectory, a readable transcript, the submitted patch, the harness test output, deterministic metrics, and a hand-verified qualitative root-cause diagnosis. Current numbers (resolve rates… See the full description on the dataset page: https://huggingface.co/datasets/daaain/swebench-verified-deepseek-v4-flash-failure-analysis.tabulartext-generationn<1K0 likes160 downloads3mo agoHugging Face28vinod-anbalagan /chart-reasoning-verified chart-reasoning-verified Chart reasoning examples generated from an explicit latent representation. The data, the question and the answer are computed before the chart is drawn, so the image is a rendering of known ground truth rather than the source of it. No model was asked to label anything. Each row carries both a rendered chart and a text serialisation of the same chart, so the set is usable for vision-language training and for text-only language model training without… See the full description on the dataset page: https://huggingface.co/datasets/vinod-anbalagan/chart-reasoning-verified.imagevisual-question-answering1K<n<10K0 likes152 downloads8d agoHugging Face29ulamai /verified-research-reasoning-trajectories Verified Research Reasoning Trajectories for RLVR This repository is the public sample and schema repository for Ulam's research-level mathematical reasoning trajectories for reinforcement learning with verifiable rewards (RLVR), process supervision, judge training, proof criticism, and private evaluations. Ulam Verified Research Reasoning Trajectories are proof-process data for RLVR. Each record contains a normalized research problem, a golden or partial-golden proof graph… See the full description on the dataset page: https://huggingface.co/datasets/ulamai/verified-research-reasoning-trajectories.documenttext-generationn<1K3 likes148 downloads2mo agoHugging Face30yale-nlp /physics-verified PHYSICS-Verified PHYSICS-Verified is a benchmark of 1,109 PhD-qualifying-exam physics problems with 2,803 scored answers, covering six core areas of physics. Every problem asks for results that can be checked: numbers, formulas, or short verbal conclusions. Each answer has been checked against its reference solution. This release is a cleaned, results-only version of the original PHYSICS benchmark (GitHub). Problems that required a proof, explanation, or drawing were removed, as… See the full description on the dataset page: https://huggingface.co/datasets/yale-nlp/physics-verified.textquestion-answering1K<n<10K0 likes136 downloads5d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.