datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
distill_qwen_7b_aime_verifications_7b_ft_verifierHLE-Verifications
HLE with Gemini 3 Pro
This dataset contains 649 multiple-choice and exact-match questions from the Humanity's Last Exam (HLE) benchmark with 50 candidate responses generated by Gemini 3 Pro for each problem. Each response has been evaluated for correctness using a mixture of Qwen3-Next-80B-A3B-instruct and Python code to parse different answer formats, and scored by multiple LLM judges according to a 0-5 rubric.
Dataset Structure
Split: Single split named "data"
Number… See the full description on the dataset page: https://huggingface.co/datasets/FUSE-verifiers/HLE-Verifications.distill_qwen_7b_math_verifications_7b_ft_verifierlean-verifier-formalizations
Lean Verifier Formalizations
A dataset of Lean 4 theorem-proving tasks for evaluating agentic coding harnesses. Each row pairs a formal task_statement (with the reference proof body removed) against a real Lean 4 repository, plus the informal_excerpt/informal_source_text describing what the theorem claims, permitted_axioms for the verifier, and provenance fields (repo_url, repo_commit_sha, license) tracing back to the source project.
Sources
Every row is pulled… See the full description on the dataset page: https://huggingface.co/datasets/Pradheep1647/lean-verifier-formalizations.deepswe-verifier-2582-v1OpenHands-Verifier-TrajectoriesR2EGym-Verifier-TrajectoriesVA-EXP-0109llm-verifier-freelancer-qwen3.5-122b-131k-opencode-traces
Agent trace dataset
Decoding the literal token IDs
The prompt_token_ids / completion_token_ids / logprobs columns are the
verbatim tokens the serving engine emitted, stored PER AGENT STEP as a
list-of-lists (one inner list per turn). To turn them back into text you MUST
use the exact tokenizer the model was served with — a generic same-family
tokenizer will decode word tokens to garbage.
Served model / tokenizer source: Qwen/Qwen3.5-122B-A10B-FP8
from transformers… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/llm-verifier-freelancer-qwen3.5-122b-131k-opencode-traces.rl__24GPU_shaped__inferredbugs-sandboxes-verifier__exp_tas_optimal_comb__40-0deepswe-verifier-merged-with-regression-with-filenamesLLMVerify-Verifier
LLMVerify-Verifier
Verification results dataset for the paper "Variation in Verification: Understanding Verification Dynamics in Large Language Models", accepted at ICLR 2026 (arXiv:2509.17995).
This dataset contains the binary verdicts and chain-of-thought verification reasoning produced by 15 verifier models judging candidate solutions from 15 generator models across three task domains. It supports systematic analysis of how problem difficulty, generator capability, and verifier… See the full description on the dataset page: https://huggingface.co/datasets/YefanZhou98/LLMVerify-Verifier.deepswe-verifier-2582-v2swebench_verified_random_100_folders_a3_rl_DCAgent_inferredbugs_sandboxes_verifierb984e9c9a3-rl-DCAgent_inferredbugs-sandboxes-verifierdev_set_v2_a3_rl_DCAgent_inferredbugs_sandboxes_verifier_55_8B_20260827_092854MoatlessTools-Agent-Verifier-Train-Dataatlas-16-verifier-permission-prompt-ablation
ATLAS report 16: does the orchestrator's "cannot solve" clause suppress candidate verification?
Complete raw products of the ATLAS rl-training report 16 experiment
(GitHub issue #36). Two system-prompt arms of the same model over the
same 78 fixed states, greedy decoding, one shared vLLM server.
What the experiment did
The ATLAS orchestrator's frozen system prompt contains the clause
You cannot solve the problem yourself; you decide when to explore
further and… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-16-verifier-permission-prompt-ablation.deepswe-verifier-merged-with-regressiona3-rl-DCAgent_llm-verifier-freelancerterminal_bench_2_a3_rl_DCAgent_inferredbugs_sandboxes_verifier_55_8B_20260525_080234R2EGym-VerifierTrajectories-PatchOnlyswebench_verified_a3_rl_DCAgent_inferredbugs_sandboxes_verifier_55_8B_20260528_003627gaia_127_a3_rl_DCAgent_inferredbugs_sandboxes_verifier_55_8B_20260526_230125verifier-debias-v2
verifier-debias-v2 — de-biased SFT data for a generative verifier
This dataset was presented in the paper One Token to Fool LLM-as-a-Judge.
GitHub repository: yulaizhao/Master-RM
SFT data to train a generative verifier (GenRM, arXiv:2408.15240)
from Qwen/Qwen2.5-7B-Instruct. Each row is a conversational example
(messages = system + user + gold assistant) plus a verdict (PASS/FAIL) and
the underlying 1–5 score. The assistant target is a critique ending in
Verdict: PASS / Verdict:… See the full description on the dataset page: https://huggingface.co/datasets/narcolepticchicken/verifier-debias-v2.inferredbugs-sandboxes-verifier-qwen3.5-122b-131k-opencode-traces
Agent trace dataset
Decoding the literal token IDs
The prompt_token_ids / completion_token_ids / logprobs columns are the
verbatim tokens the serving engine emitted, stored PER AGENT STEP as a
list-of-lists (one inner list per turn). To turn them back into text you MUST
use the exact tokenizer the model was served with — a generic same-family
tokenizer will decode word tokens to garbage.
Served model / tokenizer source: Qwen/Qwen3.5-122B-A10B-FP8
from transformers… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/inferredbugs-sandboxes-verifier-qwen3.5-122b-131k-opencode-traces.aider_polyglot_a3_rl_DCAgent_inferredbugs_sandboxes_verifier_55_8B_20260526_230125captcha-ocr-verifiervietnamese-acoustic-boundary-verifier-data
Vietnamese Acoustic Boundary & Speaker Purity Dataset (Gemini 3.8 Flash Distilled)
This dataset contains 202 curated Vietnamese audio samples with fine-grained acoustic boundary annotations distilled from Google Gemini 3.8 Flash (thinkingLevel="LOW").
It is specifically designed to train and evaluate multimodal models (e.g., Gemma 4 E4B Audio) on acoustic quality control for speech synthesis and speaker diarization pipelines.
Dataset Structure
Each sample is… See the full description on the dataset page: https://huggingface.co/datasets/tungnguyenlam/vietnamese-acoustic-boundary-verifier-data.deepswe-verifier-merged-with-regression-with-filenames-clean
