datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sci-agent-verification-cascade
Scientific Agent Verification Cascade
Public evaluation fixtures and verified aggregate results for testing whether
scientific claims keep their source, meaning, uncertainty, and verification
requirements as they move between AI agents.
This dataset accompanies the
Scientific Agent Verification Cascade
codebase. Version 0.2.0
contains synthetic evaluation data and aggregate-only results. It contains no
raw hosted-model response, private holdout identifier,
source-record… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/sci-agent-verification-cascade.swerebench-traces-raw-source-verification-enhanced-20260617
SWE-rebench Raw Source Verification Enhanced 20260617
This is a private raw source dataset for building refined mini-swe-agent SFT datasets. It is intentionally not tokenized and intentionally preserves source data plus metadata for downstream filtering, masking, weighting, and audit. Do not treat every row as a clean endpoint solve.
Download
The full dataset directory is uploaded as a single compressed archive:
hf download… See the full description on the dataset page: https://huggingface.co/datasets/eewer/swerebench-traces-raw-source-verification-enhanced-20260617.clean_cot_verification_340k元データ: https://huggingface.co/datasets/Zigeng/CoT-Verification-340k
使用したコード: https://github.com/LLMTeamAkiyama/0-data_prepare/tree/master/src/CoT-Verification-340k
データ件数: 140,980
平均トークン数: 602
最大トークン数: 2,040
合計トークン数: 84,894,510
ファイル形式: JSONL
ファイル分割数: 2
合計ファイルサイズ: 256.3 MB
加工内容:
データセットIDの付与: データフレームのインデックスに1を加算して、base_datasets_idとして新しいID列を付与しました。
response列のフィルタリング: response列が「Yes,」で始まる行のみを保持し、それ以外の行を除外しました。
prompt列の文字長によるフィルタリング: prompt列の文字列の長さが80… See the full description on the dataset page: https://huggingface.co/datasets/LLMTeamAkiyama/clean_cot_verification_340k.math7500_train_verifications_llama3.1-8b_gt_soln_in_contextgsm8k-answer-verificationllama3.1-8b_train_correct_verifications_gt_soln_in_context_consistenct_verificationsllama-3.1-8B-correct-verifications-with-gt-solutionsllama-3.1-8b-instruct-math_train-correct-verificationsveritas-claim-verification-dataset
VERITAS Claim Verification Dataset
775,294 labeled verification calls (770,000 synthetic + 5,294 adversarial)Multi-model consensus verdicts from the VERITAS Oracle across 8 domains.
Paper
VERITAS: Independence-Weighted Multi-Model Consensus for AI Output Verification
Adversarial Robustness Findings (NEW)
5,294 adversarial variants generated via deterministic transforms:
Transform
N
Accuracy
negation
2,379
99.7%
jurisdiction
114
100.0%
quantifier… See the full description on the dataset page: https://huggingface.co/datasets/VeritasOmega/veritas-claim-verification-dataset.
