CoolFace
20 results

verified

princeton-nlp /SWE-bench_VerifiedDataset Summary SWE-bench Verified is a subset of 500 samples from the SWE-bench test set, which have been human-validated for quality. SWE-bench is a dataset that tests systems’ ability to solve GitHub issues automatically. See this post for more details on the human-validation process. The dataset collects 500 test Issue-Pull Request pairs from popular Python repositories. Evaluation is performed by unit test verification using post-PR behavior as the reference solution. The original… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/SWE-bench_Verified.textn<1K387 likes275k downloads2y agoHugging FaceSWE-bench /SWE-bench_VerifiedDataset Summary SWE-bench Verified is a subset of 500 samples from the SWE-bench test set, which have been human-validated for quality. SWE-bench is a dataset that tests systems’ ability to solve GitHub issues automatically. See this post for more details on the human-validation process. The dataset collects 500 test Issue-Pull Request pairs from popular Python repositories. Evaluation is performed by unit test verification using post-PR behavior as the reference solution. The original… See the full description on the dataset page: https://huggingface.co/datasets/SWE-bench/SWE-bench_Verified.textn<1K162 likes125k downloads1mo agoHugging Faceskylenage-ai /HLE-Verified HLE-Verified A Systematic Verification and Structured Revision of Humanity’s Last Exam Overview Humanity’s Last Exam (HLE) is a high-difficulty, multi-domain benchmark designed to evaluate advanced reasoning capabilities across diverse scientific and technical domains. Following its public release, members of the open-source community raised concerns regarding the reliability of certain items. Community discussions and informal replication attempts suggested that some… See the full description on the dataset page: https://huggingface.co/datasets/skylenage-ai/HLE-Verified.document1K<n<10K19 likes45k downloads7mo agoHugging Facezai-org /terminal-bench-2-verified Terminal-Bench 2.0 Verified: Instruction & Environment Fix Version 中文版本 2026.08.18 Update: Following the release of Terminal-Bench 2.1, we conducted another round of review on several tasks' grading and instructions, validating each fix end-to-end inside the actual task images. This round fixes 6 tasks in two categories: Grading/Test Fixes (dna-insert, make-doom-for-mips, filter-js-from-html, install-windows-3.11) — the test logic itself misjudged valid submissions or let… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/terminal-bench-2-verified.82 likes27k downloads29d agoHugging Facelmms-lab /HLE-Verified HLE-Verified (HF-native JSONL) This dataset is a lightweight, evaluation-ready reformatting of the HLE-Verified benchmark created by the Skylenage Team. Original work: Weiqi Zhai et al., "HLE-Verified: A Systematic Verification and Structured Revision of Humanity's Last Exam" (arXiv:2602.13964) Original dataset: skylenage/HLE-Verified Original repository: SKYLENAGE-AI/HLE-Verified Source & Snapshot Converted from skylenage/HLE-Verified snapshot… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab/HLE-Verified.5 likes8.3k downloads7mo agoHugging Facexlangai /ubuntu_osworld_verified_trajs OSWorld-Verified Model Trajectories This repository contains trajectory results from various AI models evaluated on the OSWorld benchmark - a comprehensive evaluation environment for multimodal agents in real computer environments. Dataset Overview This dataset includes evaluation trajectories and results from multiple state-of-the-art models tested on OSWorld tasks. File Structure Each zip file contains complete evaluation trajectories including:… See the full description on the dataset page: https://huggingface.co/datasets/xlangai/ubuntu_osworld_verified_trajs.100K<n<1M22 likes5.9k downloads2mo agoHugging Face