CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01abhinavpola /tau2-bench-verified-airline tau2-bench-verified — airline domain (mirror) Mirror of the airline domain from amazon-agi/tau2-bench-verified (MIT License), pinned at commit 864350a8971a8f8ee9e7b8472e2edc380a806b0c. Re-hosted for the OpenRouter native TypeScript benchmark harness so it can fetch the verified airline tasks + environment DB at runtime. Contents tasks/test.jsonl — 50 verified airline tasks. Each row has a single task_json string column holding one verbatim tau2 v2 task object (id… See the full description on the dataset page: https://huggingface.co/datasets/abhinavpola/tau2-bench-verified-airline.textn<1K0 likes4.7k downloads3mo agoHugging Face02razzant /ouroboros-osworld-verified-opus5 Ouroboros on OSWorld-Verified: 90.69%, the highest result reported to date Status: Self-reported result over all 361 tasks. The official per-task scores, prompts, manifests and feasibility records are public here, together with every acting task record that the run produced. Start here Result 90.69% (327.39 / 361) Model anthropic/claude-opus-5 Method Screenshot only, one rollout, 100 policy turns Exact evidence f52ebf2 and evidence.json… See the full description on the dataset page: https://huggingface.co/datasets/razzant/ouroboros-osworld-verified-opus5.textothern<1K1 likes2.6k downloads1mo agoHugging Face03razzant /ouroboros-osworld-verified-sonnet46 Ouroboros on OSWorld-Verified: best published result on Claude Sonnet 4.6 Status: Self-reported result over all 361 task packages. Prompts, manifests, outcomes and feasibility records are public for every task. The official evaluator produced 360 score files; one unscored task is counted as zero. Start here Result 83.27% (300.59 / 361) Model anthropic/claude-sonnet-4.6 Method Screenshot only, one rollout, 100 policy turns Exact evidence… See the full description on the dataset page: https://huggingface.co/datasets/razzant/ouroboros-osworld-verified-sonnet46.textothern<1K0 likes2.2k downloads1mo agoHugging Face04shhu2001 /SciCode-Verified SciCode-Verified SciCode-Verified is the corrected, human-verified release of the SciCode scientific-code-generation benchmark. A problem-by-problem audit identified 263 defects in the 65-problem SciCode test split and corrected every confirmable defect. The released evaluation set contains 64 main problems and 287 scored subproblems; one original problem is excluded because its specification does not determine a unique, verifiable answer. Paper: SciCode-Verified: How Benchmark… See the full description on the dataset page: https://huggingface.co/datasets/shhu2001/SciCode-Verified.texttext-generationn<1K1 likes1.8k downloads2mo agoHugging Face05likaixin /TACO-verified Introduction This dataset contains verified solutions from the TACO dataset's training set. Solutions that fail to pass all the test cases are removed. Problems with no correct solution are also removed. The solutions were executed on Intel E5-2620 v3 CPUs with the execution timeout set to 10 seconds. Statistics in the training set Dataset # Problems # Solutions TACO 25443 1468722 TACO-verified 12898 1043251 Correct Ratio 50.69 % 71.03 %… See the full description on the dataset page: https://huggingface.co/datasets/likaixin/TACO-verified.textquestion-answering10K<n<100K20 likes1.5k downloads1y agoHugging Face06harithoppil /terminal-bench-2-verified Terminal-Bench 2.0 Verified: Instruction & Environment Fix Version 中文版本 We conducted a comprehensive review of the entire Terminal-Bench 2.0 dataset and identified various issues. Both GLM-5 and Step 3.5-Flash were evaluated using this verified version. This modified version addresses environment and instruction issues we discovered in Terminal-Bench 2.0. It includes two types of fixes: Environment Fixes: Updated Dockerfiles and instructions to support Claude Code Agent runtime… See the full description on the dataset page: https://huggingface.co/datasets/harithoppil/terminal-bench-2-verified.documenttext-generationn<1K2 likes1.5k downloads5mo agoHugging Face07opencompass /SWEBench-Pro-Verified SWE-Bench Pro Verified: Anti-hacking & Task refinement SWE-Bench Pro has emerged as a standard benchmark for evaluating software engineering agents on challenging repository-level tasks. However, our analysis work show that its evaluation is undermined by two sources of unreliability: reward hacking, enabled by leakage of gold solutions or hidden evaluation information, and task quality issues, including misleading problem statements and improperly scoped tests. These issues can… See the full description on the dataset page: https://huggingface.co/datasets/opencompass/SWEBench-Pro-Verified.tabularn<1K2 likes1.5k downloads13d agoHugging Face08salimayed /verified-defi-datasets Verified Solana Sealevel & Anchor Program Optimization Fine-Tuning Corpus Dataset Description High-density, verified AI fine-tuning dataset in ALPACA format. Domain: Solana Sealevel & Anchor Program Optimization Verified Records: 3 Estimated Tokens: 339 Quality QA Score: 99.0% Monetization Status: Direct Zero-Gas Web3 & HuggingFace Distribution texttext-generationn<1K0 likes640 downloads10d agoHugging Face09lmms-lab-eval /HLE-Verified HLE-Verified (HF-native) This dataset is a Hugging Face-native conversion of skylenage/HLE-Verified at revision becad9f339dfce27df0ebb38e55dabef12ca5735. Why this exists The source dataset stores nested verification fields with mixed runtime types (for example 0/1/"uncertain"), which breaks strict Arrow JSON parsing in datasets.load_dataset. This converted dataset normalizes those fields and publishes split-ready JSONL files for direct use in lmms_eval. Split… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-eval/HLE-Verified.text1K<n<10K0 likes422 downloads7mo agoHugging Face10Doc2Feat-bench /Doc2Feat-bench_Verified Dataset Summary NoCode-bench Verified is subset of NoCode-bench, a dataset that tests systems’ no-code feature addition ability automatically. Languages The text of the dataset is primarily English, but we make no effort to filter or otherwise clean based on language type. Dataset Structure An example of a SWE-bench datum is as follows: repo: (str) - The repository owner/name identifier from GitHub. instance_id: (str) - A formatted instance… See the full description on the dataset page: https://huggingface.co/datasets/Doc2Feat-bench/Doc2Feat-bench_Verified.texttext-generationn<1K1 likes352 downloads1y agoHugging Face11rasinmuhammed /verified-sql-rewards Verified SQL Rewards A text-to-SQL corpus where every reward carries a machine-checkable proof that it is correct. Questions, all independently verified 109,306 Databases 1,400 across 7 schema families Tables / data rows 4,400 / ~19.6 million Unique (question, answer) pairs 102,764 Candidates refused and published 12,150 Verification pass rate 90.00% Trivial baseline (always answer 0) 1.83% Each item is a natural-language question, a gold SQL query… See the full description on the dataset page: https://huggingface.co/datasets/rasinmuhammed/verified-sql-rewards.texttable-question-answering100K<n<1M0 likes290 downloads18d agoHugging Face12pankajmathur /nemotron-nano-30b-miniswe-swebench-verified Nemotron Nano 30B + mini-swe-agent SWE-bench Verified Trajectories Agent trajectories from running NVIDIA Nemotron 3 Nano 30B A3B (MoE, 8B active params) on SWE-bench Verified using mini-swe-agent. ⚠️ Incomplete Run This benchmark was terminated early due to poor performance. The model struggled with the agentic coding task. Model Information Attribute Value Model NVIDIA Nemotron 3 Nano 30B A3B Architecture MoE (30B total, 8B active) Serving vLLM… See the full description on the dataset page: https://huggingface.co/datasets/pankajmathur/nemotron-nano-30b-miniswe-swebench-verified.texttext-generationn<1K0 likes279 downloads9mo agoHugging Face13Lego-X /Lego-RL-SWE-Bench-Verified Lego-RL-SWE-Bench-Verified The 500 SWE-bench Verified instances as ready-to-run harbor RL environments — the exact evaluation set behind every SWE-bench Verified number in LEGO-RL, packaged the same way as the training set Lego-X/Lego-RL-2699 so one trainer reads both. Two parallel views of the same 500 instances: View Path What it is Official SWE-bench records swebench_verified_official_500/ The upstream princeton-nlp/SWE-bench_Verified rows, verbatim Harbor RL… See the full description on the dataset page: https://huggingface.co/datasets/Lego-X/Lego-RL-SWE-Bench-Verified.texttext-generationn<1K0 likes264 downloads28d agoHugging Face14Fujitsu-FRE /MAPS_Verified Dataset Card for Multilingual Benchmark for Global Agent Performance and Security This is the first Multilingual Agentic AI Benchmark for evaluating agentic AI systems across different languages and diverse tasks. Benchmark enables systematic analysis of how agents perform under multilingual conditions. This dataset contains 550 instances for GAIA, 660 instances for ASB, 737 instances for Maths, and 1100 instances for SWE. Each task was translated into 10 target languages resulting… See the full description on the dataset page: https://huggingface.co/datasets/Fujitsu-FRE/MAPS_Verified.texttext-generation1K<n<10K3 likes253 downloads8mo agoHugging Face15tsinghua-sigs-robot-lab /VeriLoop-Structural-Repair-Verified VLR-StructuralRepair v1.0.0 — non-regressive repair of real semantic defects Evidence-convergent supervision for function-level semantic repair under a hidden set of protected obligations. A candidate is positive only when it preserves every already-satisfied obligation and strictly repairs at least one. Aggregate improvement that breaks a protected obligation is a negative, however far the total failure count drops. The previous generation of this dataset… See the full description on the dataset page: https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-Structural-Repair-Verified.tabulartext-generation10K<n<100K0 likes252 downloads27d agoHugging Face16ulamai /verified-math-olympiad-trajectories Verified Math Olympiad Reasoning Trajectories for RLVR This repository is the public sample and schema repository for Ulam's math olympiad reasoning trajectories for reinforcement learning with verifiable rewards (RLVR), answer-verifier evaluation, process-supervision candidates, judge training, proof criticism, and private evaluations. The goal is not merely to provide final-answer math examples. Each record is a structured olympiad reasoning object containing a normalized problem… See the full description on the dataset page: https://huggingface.co/datasets/ulamai/verified-math-olympiad-trajectories.documentn<1K2 likes206 downloads4mo agoHugging Face17tsinghua-sigs-robot-lab /VeriLoop-Governed-Recurrence-Verified VLR-Recurrence-Verified VLR-Recurrence-Verified is a synthetic-data construction release for studying evidence-convergent program repair. It operationalizes a protected partial order: a candidate is positive only when it preserves every already-satisfied obligation and strictly improves at least one unresolved obligation. Scale Split Tasks Families Transitions Balanced pairs Certified finals Train 3,500 28 12,250 49,000 3,500 Validation 750 10 2,623… See the full description on the dataset page: https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-Governed-Recurrence-Verified.tabulartext-generation100K<n<1M1 likes178 downloads1mo agoHugging Face18ulamai /verified-research-reasoning-trajectories Verified Research Reasoning Trajectories for RLVR This repository is the public sample and schema repository for Ulam's research-level mathematical reasoning trajectories for reinforcement learning with verifiable rewards (RLVR), process supervision, judge training, proof criticism, and private evaluations. Ulam Verified Research Reasoning Trajectories are proof-process data for RLVR. Each record contains a normalized research problem, a golden or partial-golden proof graph… See the full description on the dataset page: https://huggingface.co/datasets/ulamai/verified-research-reasoning-trajectories.documenttext-generationn<1K3 likes173 downloads2mo agoHugging Face19Mungus451 /verified_wiki_historian_the_beatles_anthology_dataset_active Verified-Wiki-Historian: The Beatles Anthology Verified-Wiki-Historian (The Beatles Anthology) is a refined, citation-grounded instruction dataset for Beatles-specific historical question answering, summarization, and supervised fine-tuning. This dataset is a cleaned and rebuilt refinement of: Mungus451/verified_wiki_historian_the_beatles_anthology_dataset_active The current release contains 4,000 instruction records focused on Beatles history, recording sessions, release… See the full description on the dataset page: https://huggingface.co/datasets/Mungus451/verified_wiki_historian_the_beatles_anthology_dataset_active.textquestion-answering1K<n<10K1 likes158 downloads3mo agoHugging Face20namakoo /idfu-verified-code IDFU Code Negative Dataset — Free Preview A curated dataset of Python code samples that failed execution-based validation, designed for training reward models, DPO rejected-side pairs, and error-detection classifiers. Free 100-sample preview; paid full versions available separately. What's inside this preview 100 unique Python samples, all AST-validated 19 CS domains represented (MCMC, FFT, distributed consensus, ZKP, formal methods, HFT microstructure, and more)… See the full description on the dataset page: https://huggingface.co/datasets/namakoo/idfu-verified-code.texttext-generationn<1K1 likes136 downloads5mo agoHugging Face21daaain /swebench-verified-deepseek-v4-flash-failure-analysis SWE-bench Verified runs & failure analysis — DeepSeek-V4-flash (local) × mini-swe-agent Per-instance analysis of SWE-bench Verified runs of a locally-served DeepSeek-V4-flash model driven by mini-swe-agent, graded with the official SWE-bench harness. Each instance carries the full agent trajectory, a readable transcript, the submitted patch, the harness test output, deterministic metrics, and a hand-verified qualitative root-cause diagnosis. Current numbers (resolve rates… See the full description on the dataset page: https://huggingface.co/datasets/daaain/swebench-verified-deepseek-v4-flash-failure-analysis.tabulartext-generationn<1K0 likes124 downloads3mo agoHugging Face22dots-studio /IMO-AnswerBench-Verified IMO AnswerBench Verified IMO AnswerBench Verified is a human-expert-verified derivative of OpenEvals/IMO-AnswerBench, originally curated by the Google DeepMind Superhuman Reasoning team. Every record in the 400-problem benchmark was reviewed individually. The review identified and corrected 13 records while preserving the benchmark's balanced coverage of four major mathematical areas. Dataset summary Total records: 400 Verification method: record-by-record human… See the full description on the dataset page: https://huggingface.co/datasets/dots-studio/IMO-AnswerBench-Verified.textquestion-answeringn<1K1 likes117 downloads1mo agoHugging Face23protogonos /verified-tool-use-dataset Verified tool-use trajectories for LLM agents This was a time-boxed experiment by an autonomous agent (Protogonos), now concluded. Nothing here is offered for sale or for hire, and no payment is accepted. Multi-turn function-calling conversations for training and evaluating tool-using agents — 48 trajectories across 16 domains, with every tool call checked against its tool's JSON-Schema. The free sample in this repo is a real slice of the full set: the viewer above renders it… See the full description on the dataset page: https://huggingface.co/datasets/protogonos/verified-tool-use-dataset.texttext-generationn<1K1 likes117 downloads24d agoHugging Face24databounty-io /regex-pattern-generation-with-verified-match-sets-cmskdvm7 Regex Pattern Generation with Verified Match Sets A dataset of regex pattern generation with verified match sets examples for training and evaluation. Good items are unambiguous and verifiable across difficulty levels; skip synthetic-looking or low-effort cases. About This dataset was produced by the DataBounty community and published here as part of an open, karma-only program. Accepted items: 1000 Language: Regex Framework: Community License: CC-BY-4.0… See the full description on the dataset page: https://huggingface.co/datasets/databounty-io/regex-pattern-generation-with-verified-match-sets-cmskdvm7.text1K<n<10K0 likes110 downloads13d agoHugging Face25vinod-anbalagan /chart-reasoning-verified chart-reasoning-verified Chart reasoning examples generated from an explicit latent representation. The data, the question and the answer are computed before the chart is drawn, so the image is a rendering of known ground truth rather than the source of it. No model was asked to label anything. Each row carries both a rendered chart and a text serialisation of the same chart, so the set is usable for vision-language training and for text-only language model training without… See the full description on the dataset page: https://huggingface.co/datasets/vinod-anbalagan/chart-reasoning-verified.imagevisual-question-answering1K<n<10K0 likes109 downloads4d agoHugging Face26giggiovpg /android-kotlin-compose-compiler-verified Qwandroid — Compiler-Verified Modern Android (Kotlin + Jetpack Compose) Dataset 5,777 SFT examples + 150 held-out eval + 8,027 DPO preference pairs. Every SFT row was actually compiled — not LLM-approved, not heuristically filtered. A subset was verified behaviorally by running JUnit tests. Built to fine-tune small models into focused Android specialists rather than general-purpose coders. Why this exists Android code in pretraining corpora is largely stale —… See the full description on the dataset page: https://huggingface.co/datasets/giggiovpg/android-kotlin-compose-compiler-verified.texttext-generation10K<n<100K1 likes104 downloads1mo agoHugging Face27yale-nlp /physics-verified PHYSICS-Verified PHYSICS-Verified is a benchmark of 1,109 PhD-qualifying-exam physics problems with 2,803 scored answers, covering six core areas of physics. Every problem asks for results that can be checked: numbers, formulas, or short verbal conclusions. Each answer has been checked against its reference solution. This release is a cleaned, results-only version of the original PHYSICS benchmark (GitHub). Problems that required a proof, explanation, or drawing were removed, as… See the full description on the dataset page: https://huggingface.co/datasets/yale-nlp/physics-verified.textquestion-answering1K<n<10K0 likes103 downloads2d agoHugging Face28F-A-I-L /kodcode-verified-python-235k KodCode-Verified Python — 234,555 execution-verified Python SFT rows One row per problem. Every assistant turn is code that passed its own unit tests when actually run — real pytest against KodCode-V1's tests, in a pinned interpreter, in a sandboxed subprocess. No LLM judge, no heuristic filter, no model-generated answers. Unlike the v3 release this supersedes, the corpus is deduplicated, decontaminated against HumanEval/MBPP, and stripped of rows whose tests cannot constrain… See the full description on the dataset page: https://huggingface.co/datasets/F-A-I-L/kodcode-verified-python-235k.texttext-generation100K<n<1M0 likes94 downloads7d agoHugging Face29cy-330 /UPBench-Error-verified-v2 UPBench-Error-verified-v2 LCZZZZ/UPBench-Error → generation_error 子集,经两轮人工核验后保留的 1,219 条样本。每条样本视觉上看不出明显的低级生成缺陷。 筛选过程 步骤 剩余 原始 generation_error 样本 5,761 剔除 is_gui=true(GUI-World / egoproactive 屏幕录制) 3,390 第一轮:逐条过目 error_clip.mp4 good 1,420 / bad 1,970 第二轮:对第一轮 good 再过一遍 good 1,219 / bad 201 第二轮刷掉了第一轮 14.2% 的样本,最终保留率 1,219 / 3,390 = 36.0%。 内容 metadata/manifest-verified.jsonl 1,219 条,原 manifest 全部 25 个字段逐字保留,… See the full description on the dataset page: https://huggingface.co/datasets/cy-330/UPBench-Error-verified-v2.textvideo-classification1K<n<10K0 likes88 downloads26d agoHugging Face30likaixin /APPS-verified Introduction This dataset contains verified solutions from the APPS dataset's training set. Solutions that fail to pass all the test cases are removed. Problems with no correct solution are also removed. The solutions were executed on Intel E5-2620 v3 CPUs with the execution timeout set to 10 seconds. Statistics in the training set Dataset # Problems # Solutions TACO 5000 117232 TACO-verified 4211 93921 Correct Ratio 84.22% 80.12% tabularquestion-answering1K<n<10K5 likes87 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.