CoolFace
20 results

Veri

princeton-nlp /SWE-bench_VerifiedDataset Summary SWE-bench Verified is a subset of 500 samples from the SWE-bench test set, which have been human-validated for quality. SWE-bench is a dataset that tests systems’ ability to solve GitHub issues automatically. See this post for more details on the human-validation process. The dataset collects 500 test Issue-Pull Request pairs from popular Python repositories. Evaluation is performed by unit test verification using post-PR behavior as the reference solution. The original… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/SWE-bench_Verified.textn<1K387 likes254k downloads2y agoHugging FaceSWE-bench /SWE-bench_VerifiedDataset Summary SWE-bench Verified is a subset of 500 samples from the SWE-bench test set, which have been human-validated for quality. SWE-bench is a dataset that tests systems’ ability to solve GitHub issues automatically. See this post for more details on the human-validation process. The dataset collects 500 test Issue-Pull Request pairs from popular Python repositories. Evaluation is performed by unit test verification using post-PR behavior as the reference solution. The original… See the full description on the dataset page: https://huggingface.co/datasets/SWE-bench/SWE-bench_Verified.textn<1K165 likes125k downloads1mo agoHugging Faceskylenage-ai /HLE-Verified HLE-Verified A Systematic Verification and Structured Revision of Humanity’s Last Exam Overview Humanity’s Last Exam (HLE) is a high-difficulty, multi-domain benchmark designed to evaluate advanced reasoning capabilities across diverse scientific and technical domains. Following its public release, members of the open-source community raised concerns regarding the reliability of certain items. Community discussions and informal replication attempts suggested that some… See the full description on the dataset page: https://huggingface.co/datasets/skylenage-ai/HLE-Verified.document1K<n<10K19 likes48k downloads7mo agoHugging Facezai-org /terminal-bench-2-verified Terminal-Bench 2.0 Verified: Instruction & Environment Fix Version 中文版本 2026.08.18 Update: Following the release of Terminal-Bench 2.1, we conducted another round of review on several tasks' grading and instructions, validating each fix end-to-end inside the actual task images. This round fixes 6 tasks in two categories: Grading/Test Fixes (dna-insert, make-doom-for-mips, filter-js-from-html, install-windows-3.11) — the test logic itself misjudged valid submissions or let… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/terminal-bench-2-verified.82 likes25k downloads1mo agoHugging FacehbXNov /distill_r1_qwen_math_1.5b_128_solns_math_verifications0 likes11k downloads2y agoHugging Facelmms-lab /HLE-Verified HLE-Verified (HF-native JSONL) This dataset is a lightweight, evaluation-ready reformatting of the HLE-Verified benchmark created by the Skylenage Team. Original work: Weiqi Zhai et al., "HLE-Verified: A Systematic Verification and Structured Revision of Humanity's Last Exam" (arXiv:2602.13964) Original dataset: skylenage/HLE-Verified Original repository: SKYLENAGE-AI/HLE-Verified Source & Snapshot Converted from skylenage/HLE-Verified snapshot… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab/HLE-Verified.5 likes9.3k downloads7mo agoHugging Face