verified
DeepSeek-V4-Flash-0731-StrixHalo-Verified-GGUFDCFT-OpenThinker-Verified-Llama-3_3-70B-bs-1024-i1-GGUFDCFT-Stratos-Verified-114k-Llama-3_3-70B-bs-256-i1-GGUFDCFT-Stratos-verified-114k-Llama-3_1-8B-i1-GGUFDCFT-Stratos-Verified-114k-qwen-2.5-72B-i1-GGUFDCFT-Stratos-Verified-114k-7B-4gpus-GGUFCORE-Qwen3-1.7B-MATH-verified-GGUFDCFT-Stratos-Verified-114k-7B-4gpus-i1-GGUF
Datasets
All datasets matching “verified”SWE-bench_VerifiedDataset Summary
SWE-bench Verified is a subset of 500 samples from the SWE-bench test set, which have been human-validated for quality. SWE-bench is a dataset that tests systems’ ability to solve GitHub issues automatically. See this post for more details on the human-validation process.
The dataset collects 500 test Issue-Pull Request pairs from popular Python repositories. Evaluation is performed by unit test verification using post-PR behavior as the reference solution.
The original… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/SWE-bench_Verified.SWE-bench_VerifiedDataset Summary
SWE-bench Verified is a subset of 500 samples from the SWE-bench test set, which have been human-validated for quality. SWE-bench is a dataset that tests systems’ ability to solve GitHub issues automatically. See this post for more details on the human-validation process.
The dataset collects 500 test Issue-Pull Request pairs from popular Python repositories. Evaluation is performed by unit test verification using post-PR behavior as the reference solution.
The original… See the full description on the dataset page: https://huggingface.co/datasets/SWE-bench/SWE-bench_Verified.HLE-Verified
HLE-Verified
A Systematic Verification and Structured Revision of Humanity’s Last Exam
Overview
Humanity’s Last Exam (HLE) is a high-difficulty, multi-domain benchmark designed to evaluate advanced reasoning capabilities across diverse scientific and technical domains.
Following its public release, members of the open-source community raised concerns regarding the reliability of certain items. Community discussions and informal replication attempts suggested that some… See the full description on the dataset page: https://huggingface.co/datasets/skylenage-ai/HLE-Verified.terminal-bench-2-verified
Terminal-Bench 2.0 Verified: Instruction & Environment Fix Version
中文版本
2026.08.18 Update: Following the release of Terminal-Bench 2.1, we conducted another round of review on several tasks' grading and instructions, validating each fix end-to-end inside the actual task images. This round fixes 6 tasks in two categories: Grading/Test Fixes (dna-insert, make-doom-for-mips, filter-js-from-html, install-windows-3.11) — the test logic itself misjudged valid submissions or let… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/terminal-bench-2-verified.HLE-Verified
HLE-Verified (HF-native JSONL)
This dataset is a lightweight, evaluation-ready reformatting of the HLE-Verified benchmark created by the Skylenage Team.
Original work: Weiqi Zhai et al., "HLE-Verified: A Systematic Verification and Structured Revision of Humanity's Last Exam" (arXiv:2602.13964)
Original dataset: skylenage/HLE-Verified
Original repository: SKYLENAGE-AI/HLE-Verified
Source & Snapshot
Converted from skylenage/HLE-Verified snapshot… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab/HLE-Verified.ubuntu_osworld_verified_trajs
OSWorld-Verified Model Trajectories
This repository contains trajectory results from various AI models evaluated on the OSWorld benchmark - a comprehensive evaluation environment for multimodal agents in real computer environments.
Dataset Overview
This dataset includes evaluation trajectories and results from multiple state-of-the-art models tested on OSWorld tasks.
File Structure
Each zip file contains complete evaluation trajectories including:… See the full description on the dataset page: https://huggingface.co/datasets/xlangai/ubuntu_osworld_verified_trajs.
