datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lean-verifier-formalizations
Lean Verifier Formalizations
A dataset of Lean 4 theorem-proving tasks for evaluating agentic coding harnesses. Each row pairs a formal task_statement (with the reference proof body removed) against a real Lean 4 repository, plus the informal_excerpt/informal_source_text describing what the theorem claims, permitted_axioms for the verifier, and provenance fields (repo_url, repo_commit_sha, license) tracing back to the source project.
Sources
Every row is pulled… See the full description on the dataset page: https://huggingface.co/datasets/Pradheep1647/lean-verifier-formalizations.deepswe-verifier-2582-v1OpenHands-Verifier-TrajectoriesR2EGym-Verifier-Trajectoriesllm-verifier-freelancer-qwen3.5-122b-131k-opencode-traces
Agent trace dataset
Decoding the literal token IDs
The prompt_token_ids / completion_token_ids / logprobs columns are the
verbatim tokens the serving engine emitted, stored PER AGENT STEP as a
list-of-lists (one inner list per turn). To turn them back into text you MUST
use the exact tokenizer the model was served with — a generic same-family
tokenizer will decode word tokens to garbage.
Served model / tokenizer source: Qwen/Qwen3.5-122B-A10B-FP8
from transformers… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/llm-verifier-freelancer-qwen3.5-122b-131k-opencode-traces.rl__24GPU_shaped__inferredbugs-sandboxes-verifier__exp_tas_optimal_comb__40-0deepswe-verifier-merged-with-regression-with-filenamesdeepswe-verifier-2582-v2swebench_verified_random_100_folders_a3_rl_DCAgent_inferredbugs_sandboxes_verifierb984e9c9a3-rl-DCAgent_inferredbugs-sandboxes-verifierdev_set_v2_a3_rl_DCAgent_inferredbugs_sandboxes_verifier_55_8B_20260827_092854deepswe-verifier-merged-with-regressionverifier-debias-v2
verifier-debias-v2 — de-biased SFT data for a generative verifier
This dataset was presented in the paper One Token to Fool LLM-as-a-Judge.
GitHub repository: yulaizhao/Master-RM
SFT data to train a generative verifier (GenRM, arXiv:2408.15240)
from Qwen/Qwen2.5-7B-Instruct. Each row is a conversational example
(messages = system + user + gold assistant) plus a verdict (PASS/FAIL) and
the underlying 1–5 score. The assistant target is a critique ending in
Verdict: PASS / Verdict:… See the full description on the dataset page: https://huggingface.co/datasets/narcolepticchicken/verifier-debias-v2.a3-rl-DCAgent_llm-verifier-freelancerR2EGym-VerifierTrajectories-PatchOnlyinferredbugs-sandboxes-verifier-qwen3.5-122b-131k-opencode-traces
Agent trace dataset
Decoding the literal token IDs
The prompt_token_ids / completion_token_ids / logprobs columns are the
verbatim tokens the serving engine emitted, stored PER AGENT STEP as a
list-of-lists (one inner list per turn). To turn them back into text you MUST
use the exact tokenizer the model was served with — a generic same-family
tokenizer will decode word tokens to garbage.
Served model / tokenizer source: Qwen/Qwen3.5-122B-A10B-FP8
from transformers… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/inferredbugs-sandboxes-verifier-qwen3.5-122b-131k-opencode-traces.terminal_bench_2_a3_rl_DCAgent_inferredbugs_sandboxes_verifier_55_8B_20260525_080234swebench_verified_a3_rl_DCAgent_inferredbugs_sandboxes_verifier_55_8B_20260528_003627gaia_127_a3_rl_DCAgent_inferredbugs_sandboxes_verifier_55_8B_20260526_230125aider_polyglot_a3_rl_DCAgent_inferredbugs_sandboxes_verifier_55_8B_20260526_230125deepswe-verifier-merged-with-regression-with-filenames-cleanvietnamese-acoustic-boundary-verifier-data
Vietnamese Acoustic Boundary & Speaker Purity Dataset (Gemini 3.8 Flash Distilled)
This dataset contains 202 curated Vietnamese audio samples with fine-grained acoustic boundary annotations distilled from Google Gemini 3.8 Flash (thinkingLevel="LOW").
It is specifically designed to train and evaluate multimodal models (e.g., Gemma 4 E4B Audio) on acoustic quality control for speech synthesis and speaker diarization pipelines.
Dataset Structure
Each sample is… See the full description on the dataset page: https://huggingface.co/datasets/tungnguyenlam/vietnamese-acoustic-boundary-verifier-data.financeagent_terminal_a3_rl_DCAgent_inferredbugs_sandboxes_verifier_55_8B_20260cf484ec3dev_set_v2_a3_rl_DCAgent_inferredbugs_sandboxes_verifier_55_8B_20260525_080234medagentbench_a3_rl_DCAgent_inferredbugs_sandboxes_verifier_55_8B_20260526_230706dev_set_v2_rl__24GPU_shaped__inferredbugs_sandboxes_verifier__exp_tas_optimal_c937091c1deepswe-verifier-3582-exitreason-agent-priority-v1terminal_bench_2_rl__24GPU_shaped__inferredbugs_sandboxes_verifier__exp_tas_opt2ea25390deepswe-swebv-eval-n16-verifier-v1gpqa-llama-3-8b-verifier
