CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01PrimeIntellect /Multi-SWE-RL-Verified Multi-SWE-RL-Verified Gold-patch-validated subset of PrimeIntellect/Multi-SWE-RL-Reupload (ByteDance's Multi-SWE-RL): 2,232 / 4,703 rows across C, Go, Java, JavaScript, Rust, and TypeScript that produce a clean reward signal end-to-end. Default dataset of the multiswe_v1 taskset. Changes vs upstream Starting from the 4,703-row re-upload: C++ dropped wholesale — 0/449 rows passed gold-patch validation in pass 1; the images are broken for scoring, not merely… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/Multi-SWE-RL-Verified.tabulartext-generation1K<n<10K4 likes5.6k downloads3mo agoHugging Face02naderalfares /epoch_ai_swebench_verified Epoch AI SWE-bench Verified Traces Complete public trace archives and an analysis-ready Parquet conversion of Epoch AI's SWE-bench Verified evaluations. Contents 34 published evaluation runs covering 16,456 traces (484 SWE-bench instances per run). data/: loadable Parquet data, one exact trace per row. original/: the byte-identical .eval archives published by Epoch AI. run_metadata/: non-sample files from each .eval archive (header.json, summaries, reductions… See the full description on the dataset page: https://huggingface.co/datasets/naderalfares/epoch_ai_swebench_verified.tabulartext-generation10K<n<100K1 likes4.2k downloads1mo agoHugging Face03PrimeIntellect /R2E-Gym-Subset-Verified R2E-Gym-Subset-Verified Gold-patch-validated subset of R2E-Gym/R2E-Gym-Subset (paper). The train split contains 4,522 / 4,578 rows (98.78%) verified scoreable end-to-end: apply the gold patch, run the upstream /testbed/run_tests.sh baked into the row's image, check the parsed outcomes against expected_output_json. Changes vs upstream Validation-only subset — our passes, run in fresh sandboxes per row: one full pass at concurrency 200, then a 10× retry pass over… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/R2E-Gym-Subset-Verified.tabulartext-generation1K<n<10K1 likes3.2k downloads3mo agoHugging Face04AmazonScience /SWE-PolyBench_Verified SWE-PolyBench SWE-PolyBench is a multi language repo level software engineering benchmark. Currently it includes 4 languages: Python, Java, Javascript, and Typescript. The number of instances in the verified split is: Javascript: 100 Typescript: 100 Python: 113 Java: 69 Datasets There are total three datasets available under SWE-PolyBench. AmazonScience/SWE-PolyBench is the full dataset, AmazonScience/SWE-PolyBench_500 is the stratified sampled dataset with 500 instances… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/SWE-PolyBench_Verified.tabularn<1K5 likes3.2k downloads10mo agoHugging Face05nsk7153 /MedCalc-Bench-Verified Updates Updates to MedCalc-Bench Verified will be made on this page going forward. Here is the github link for our repository: https://github.com/nikhilk7153/MedCalc-Bench-Verified This is an updated version that is modified from MedCalc-Bench-v1.2. While we have audited MedCalc-Bench Verified on mulitple occasions, should there by any corrections or enhancements, we will update with a new release and specify any changes. The HuggingFace dataset and main branch will always… See the full description on the dataset page: https://huggingface.co/datasets/nsk7153/MedCalc-Bench-Verified.tabularquestion-answering10K<n<100K7 likes1.9k downloads18d agoHugging Face06opencompass /SWEBench-Pro-Verified SWE-Bench Pro Verified: Anti-hacking & Task refinement SWE-Bench Pro has emerged as a standard benchmark for evaluating software engineering agents on challenging repository-level tasks. However, our analysis work show that its evaluation is undermined by two sources of unreliability: reward hacking, enabled by leakage of gold solutions or hidden evaluation information, and task quality issues, including misleading problem statements and improperly scoped tests. These issues can… See the full description on the dataset page: https://huggingface.co/datasets/opencompass/SWEBench-Pro-Verified.tabularn<1K3 likes1.6k downloads16d agoHugging Face073it /bitaudit_verification_dataset_v2tabular1K<n<10K0 likes1.5k downloads3y agoHugging Face08castcheck /temperature-verification CastCheck — daily station-level verification of public weather forecasts Independent, automated verification of raw 2 m temperature forecasts from operational NWP (ECMWF IFS HRES, NCEP GFS) and AI models (ECMWF AIFS Single; NOAA/CIRA operational runs of GraphCast, Pangu-Weather, FourCastNet v2 and Aurora from both GFS and IFS initial conditions) at 23 U.S. first-order stations — 22 major airports plus New York Central Park. The headline metric is the instantaneous 2 m… See the full description on the dataset page: https://huggingface.co/datasets/castcheck/temperature-verification.tabular10M<n<100M0 likes836 downloads20h agoHugging Face09r2e-edits /deepswe-verifier-2582-v1tabular1K<n<10K0 likes508 downloads1y agoHugging Face10syvai /danish-asr-verified danish-asr-verified ALL rows of syvai/danish-asr-unified transcribed by the syv-transcribe ensemble (hviske-v5.3 + hviske-v5, confidence-weighted ROVER), each annotated with: verified — True when the ensemble independently reproduced the reference exactly (compared after lowercasing, punctuation-strip, whitespace-collapse). Two independent witnesses agree => near-certain label. wer_teacher_vs_ref / cer_teacher_vs_ref — word/character error rate between normalized teacher output… See the full description on the dataset page: https://huggingface.co/datasets/syvai/danish-asr-verified.tabularautomatic-speech-recognition1M<n<10M0 likes453 downloads2mo agoHugging Face113it /bitaudit_verification_dataset Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/3it/bitaudit_verification_dataset.tabularn<1K0 likes437 downloads3y agoHugging Face12HayatoHongoEveryonesAI /qa_verify_tir_5.9M_new_unfiltered_v1"HayatoHongoEveryonesAI/qa_verify_tir_2.9M_new_v1", "HayatoHongoEveryonesAI/qa_verify_1m_tir_3", "HayatoHongoEveryonesAI/qa_verify_1m_tir_4", "HayatoHongoEveryonesAI/qa_verify_1m_tir_5", tabular1M<n<10M0 likes398 downloads8mo agoHugging Face13AmineHA /WebArena-Verified WebArena-Verified Dataset description WebArena-Verified is a curated benchmark dataset of web tasks designed for reproducible evaluation of web agents across multiple realistic websites. Sources GitHub repository: webarena-verified Original WebArena benchmark: webarena.dev Splits full: 812 rows hard: 258 rows Tasks per site Counts below are task counts grouped by category. Tasks with more than one site are grouped under multi-category… See the full description on the dataset page: https://huggingface.co/datasets/AmineHA/WebArena-Verified.tabular1K<n<10K2 likes394 downloads8mo agoHugging Face14togethercomputer /CoderForge-Preview-32B-SWE-Bench-Verified-Evaluation-trajectoriestabularn<1K13 likes380 downloads8mo agoHugging Face15HayatoHongoEveryonesAI /qa_verify_cot_new_6M_unfiltered_v7dataset_names = [ "HayatoHongoEveryonesAI/qa_verify_1m_cot_1", "HayatoHongoEveryonesAI/qa_verify_1m_cot_2", "HayatoHongoEveryonesAI/qa_verify_1m_cot_3", "HayatoHongoEveryonesAI/qa_verify_1m_cot_4", "HayatoHongoEveryonesAI/qa_verify_1m_cot_5", "HayatoHongoEveryonesAI/qa_verify_2m_cot_2", "HayatoHongoEveryonesAI/qa_verify_2m_cot_3", ] https://colab.research.google.com/drive/1272DRwGt02zokQiHHOl4HpoKezdyw59O?usp=sharing tabular1M<n<10M0 likes360 downloads8mo agoHugging Face16mlfoundations-dev /pdf_science_questions_verified_r1_traces__2_24_25 Dataset card for pdf_science_questions_verified_r1_traces__2_24_25 This dataset was made with Curator. Dataset details A sample from the dataset: { "url": "https://www.ttcho.com/_files/ugd/988b76_01ceeff230b24cbbb0125b2bfa3f3475.pdf", "filename": "988b76_01ceeff230b24cbbb0125b2bfa3f3475.pdf", "success": true, "page_count": 37, "page_number": 1, "question_choices_solutions": "QUESTION: What is the identity of X in the reaction 14N + 1n \u2192… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/pdf_science_questions_verified_r1_traces__2_24_25.tabular1K<n<10K0 likes344 downloads2y agoHugging Face17HayatoHongoEveryonesAI /qa_verify_cot_new_5.1M_v7HayatoHongoEveryonesAI/qa_verify_cot_new_6M_unfiltered_v7 https://colab.research.google.com/drive/1tjJ14xLa0UZ0slYqnRuQR8ngk1sPyY8v?usp=sharing tabular1M<n<10M0 likes309 downloads8mo agoHugging Face18r2e-edits /deepswe-verifier-merged-with-regression-with-filenamestabular1K<n<10K0 likes259 downloads1y agoHugging Face19GenData-Research /scientific-verification Scientific Verification Benchmark: NMC Cathodes Dataset summary The benchmark contains 50 scientific claims about NMC (lithium nickel manganese cobalt oxide) battery cathodes. Each claim is answered by Claude Opus 5, GPT 5.6 Luna and Gemini 3.1 Pro using a set of 20 open-access papers, producing 150 scored answers. The accompanying reference set contains 1,991 experiment-grounded measurements curated from 227 open-access papers, with experimental conditions and… See the full description on the dataset page: https://huggingface.co/datasets/GenData-Research/scientific-verification.tabularquestion-answering1K<n<10K0 likes250 downloads9d agoHugging Face20tasksource /nli-veridicality-transitivity@inproceedings{yanaka-etal-2021-exploring, title = "Exploring Transitivity in Neural {NLI} Models through Veridicality", author = "Yanaka, Hitomi and Mineshima, Koji and Inui, Kentaro", booktitle = "Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume", year = "2021", pages = "920--934", } tabulartext-classification100K<n<1M1 likes241 downloads4y agoHugging Face21YefanZhou98 /LLMVerify-Verifier LLMVerify-Verifier Verification results dataset for the paper "Variation in Verification: Understanding Verification Dynamics in Large Language Models", accepted at ICLR 2026 (arXiv:2509.17995). This dataset contains the binary verdicts and chain-of-thought verification reasoning produced by 15 verifier models judging candidate solutions from 15 generator models across three task domains. It supports systematic analysis of how problem difficulty, generator capability, and verifier… See the full description on the dataset page: https://huggingface.co/datasets/YefanZhou98/LLMVerify-Verifier.tabular1M<n<10M1 likes232 downloads5mo agoHugging Face22fatihdx /tr-rss-haber-akisi-verisi TR-RSS Haber Akışı Verisi TL;DR — Bu veri seti, Türkiye odaklı haber/RSS akışlarından toplanan kayıtları; mükerrerlik, spam, reklam, amaç dışı kategori, yurtdışı odak ve editoryal çerçeve yoğunluğu açısından katmanlı kalite kontrolden geçirerek erken sinyal üretimine uygun hâle getirir. Doğrulama kararı / verdict üretmez; ClaimReview ve dezenformasyon araştırmaları için upstream izleme ve kaynak önceliklendirme katmanı olarak tasarlanmıştır. Ölçek: 307.800 öğe incelendi →… See the full description on the dataset page: https://huggingface.co/datasets/fatihdx/tr-rss-haber-akisi-verisi.tabulartext-classification10K<n<100K0 likes204 downloads3mo agoHugging Face23r2e-edits /deepswe-verifier-2582-v2tabular1K<n<10K0 likes201 downloads1y agoHugging Face24huyouare /SWE-bench_Verified_With_Annotationstabularn<1K1 likes167 downloads2y agoHugging Face25r2e-edits /deepswe-verifier-merged-with-regressiontabular1K<n<10K0 likes164 downloads1y agoHugging Face26daaain /swebench-verified-deepseek-v4-flash-failure-analysis SWE-bench Verified runs & failure analysis — DeepSeek-V4-flash (local) × mini-swe-agent Per-instance analysis of SWE-bench Verified runs of a locally-served DeepSeek-V4-flash model driven by mini-swe-agent, graded with the official SWE-bench harness. Each instance carries the full agent trajectory, a readable transcript, the submitted patch, the harness test output, deterministic metrics, and a hand-verified qualitative root-cause diagnosis. Current numbers (resolve rates… See the full description on the dataset page: https://huggingface.co/datasets/daaain/swebench-verified-deepseek-v4-flash-failure-analysis.tabulartext-generationn<1K0 likes160 downloads3mo agoHugging Face27XumengWen /AIME24-25_CoT_Verification Dataset for ICLR 2026 Paper: Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs 📌 Dataset Summary This dataset contains the rollouts (reasoning traces) and verification results used in our ICLR 2026 paper. The data allows for the analysis of how Reinforcement Learning with Verifiable Rewards (RLVR) incentivizes the correct reasoning of Large Language Models (LLMs) on challenging mathematics benchmarks. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/XumengWen/AIME24-25_CoT_Verification.tabulartext-generation100K<n<1M1 likes159 downloads8mo agoHugging Face28JetBrains-Research /agent-trajectories-swe-bench-test-minus-verified Agent Trajectories: SWE-bench Test \ Verified — Mixed Teachers (gpt-5.2 / gpt-5-mini) Summary Full multi-turn agent trajectories collected from the SWE-bench Test minus Verified split (i.e., SWE-bench Test instances that are not part of SWE-bench Verified). Intended for SFT of agent models on coding tasks. Data Collection Each trajectory was produced by a GT-aware lookahead agent that, at every turn: Sampled a candidate response from both gpt-5.2 and… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/agent-trajectories-swe-bench-test-minus-verified.tabulartext-generation1K<n<10K0 likes159 downloads6mo agoHugging Face29snorkelai /Tau2-Bench-Verified-Airline-With-Code-Agents Dataset Card for a Code Agent Version of Tau Bench 2 Airline Dataset Summary This dataset includes sample traces and associated metadata from multi-turn interactions between an code agent and AI assistant, along with the original verion of the tasks with more bespoke tools. The dataset is based on a verified version of the Airline environment from Sierra.ai's Tau^2 Bench with the verified version from Amazon AGI group here. You can find an earlier version of the dataset… See the full description on the dataset page: https://huggingface.co/datasets/snorkelai/Tau2-Bench-Verified-Airline-With-Code-Agents.tabularn<1K3 likes155 downloads7mo agoHugging Face30vinod-anbalagan /chart-reasoning-verified chart-reasoning-verified Chart reasoning examples generated from an explicit latent representation. The data, the question and the answer are computed before the chart is drawn, so the image is a rendering of known ground truth rather than the source of it. No model was asked to label anything. Each row carries both a rendered chart and a text serialisation of the same chart, so the set is usable for vision-language training and for text-only language model training without… See the full description on the dataset page: https://huggingface.co/datasets/vinod-anbalagan/chart-reasoning-verified.imagevisual-question-answering1K<n<10K0 likes152 downloads8d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.