datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sci-agent-verification-cascade
Scientific Agent Verification Cascade
Public evaluation fixtures and verified aggregate results for testing whether
scientific claims keep their source, meaning, uncertainty, and verification
requirements as they move between AI agents.
This dataset accompanies the
Scientific Agent Verification Cascade
codebase. Version 0.2.0
contains synthetic evaluation data and aggregate-only results. It contains no
raw hosted-model response, private holdout identifier,
source-record… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/sci-agent-verification-cascade.swerebench-traces-raw-source-verification-enhanced-20260617
SWE-rebench Raw Source Verification Enhanced 20260617
This is a private raw source dataset for building refined mini-swe-agent SFT datasets. It is intentionally not tokenized and intentionally preserves source data plus metadata for downstream filtering, masking, weighting, and audit. Do not treat every row as a clean endpoint solve.
Download
The full dataset directory is uploaded as a single compressed archive:
hf download… See the full description on the dataset page: https://huggingface.co/datasets/eewer/swerebench-traces-raw-source-verification-enhanced-20260617.keural-v2-self-verification
Self-Verification Dataset
Status: generation in progress. 4,477 / 50,000 target rows (~9%), growing. Being generated in parallel across multiple environments/models — see Generation models mix below. This card describes the file as of this snapshot; row count and model mix will change on re-upload.
File: 05_self_verification_generated.jsonl (one JSON object per line).
What this is
Synthetic "wrong draft → self-critique → corrected solution" traces, for training a… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/keural-v2-self-verification.successprm-1K-verification-cots
SuccessPRM 1K Verification CoTs
This dataset contains 1,000 math verification chain-of-thought examples in a success/fail step-verification format.
Split
train: 1000 rows
Columns
problem
prefix
cot
raw_cot
prefix_steps
gt_step_labels
prefix_label
step_success_labels
step_success_ratings
format_valid
Generation summary
Source dataset: launch/thinkprm-1K-verification-cots
Source split: train
Generator model: Qwen/QwQ-32B-Preview
Prompt version:… See the full description on the dataset page: https://huggingface.co/datasets/christinakopi/successprm-1K-verification-cots.
