datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bitaudit_verification_dataset_v2temperature-verification
CastCheck — daily station-level verification of public weather forecasts
Independent, automated verification of raw 2 m temperature forecasts from operational NWP
(ECMWF IFS HRES, NCEP GFS) and AI models (ECMWF AIFS Single; NOAA/CIRA operational runs of GraphCast,
Pangu-Weather, FourCastNet v2 and Aurora from both GFS and IFS initial conditions) at 23
U.S. first-order stations — 22 major airports plus New York Central Park. The headline metric is the instantaneous 2 m… See the full description on the dataset page: https://huggingface.co/datasets/castcheck/temperature-verification.bitaudit_verification_dataset
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/3it/bitaudit_verification_dataset.scientific-verification
Scientific Verification Benchmark: NMC Cathodes
Dataset summary
The benchmark contains 50 scientific claims about NMC (lithium nickel manganese cobalt oxide) battery cathodes. Each claim is answered by Claude Opus 5, GPT 5.6 Luna and Gemini 3.1 Pro using a set of 20 open-access papers, producing 150 scored answers. The accompanying reference set contains 1,991 experiment-grounded measurements curated from 227 open-access papers, with experimental conditions and… See the full description on the dataset page: https://huggingface.co/datasets/GenData-Research/scientific-verification.AIME24-25_CoT_Verification
Dataset for ICLR 2026 Paper: Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
📌 Dataset Summary
This dataset contains the rollouts (reasoning traces) and verification results used in our ICLR 2026 paper. The data allows for the analysis of how Reinforcement Learning with Verifiable Rewards (RLVR) incentivizes the correct reasoning of Large Language Models (LLMs) on challenging mathematics benchmarks.
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/XumengWen/AIME24-25_CoT_Verification.sci-agent-verification-cascade
Scientific Agent Verification Cascade
Public evaluation fixtures and verified aggregate results for testing whether
scientific claims keep their source, meaning, uncertainty, and verification
requirements as they move between AI agents.
This dataset accompanies the
Scientific Agent Verification Cascade
codebase. Version 0.2.0
contains synthetic evaluation data and aggregate-only results. It contains no
raw hosted-model response, private holdout identifier,
source-record… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/sci-agent-verification-cascade.swerebench-traces-raw-source-verification-enhanced-20260617
SWE-rebench Raw Source Verification Enhanced 20260617
This is a private raw source dataset for building refined mini-swe-agent SFT datasets. It is intentionally not tokenized and intentionally preserves source data plus metadata for downstream filtering, masking, weighting, and audit. Do not treat every row as a clean endpoint solve.
Download
The full dataset directory is uploaded as a single compressed archive:
hf download… See the full description on the dataset page: https://huggingface.co/datasets/eewer/swerebench-traces-raw-source-verification-enhanced-20260617.clean_cot_verification_340k元データ: https://huggingface.co/datasets/Zigeng/CoT-Verification-340k
使用したコード: https://github.com/LLMTeamAkiyama/0-data_prepare/tree/master/src/CoT-Verification-340k
データ件数: 140,980
平均トークン数: 602
最大トークン数: 2,040
合計トークン数: 84,894,510
ファイル形式: JSONL
ファイル分割数: 2
合計ファイルサイズ: 256.3 MB
加工内容:
データセットIDの付与: データフレームのインデックスに1を加算して、base_datasets_idとして新しいID列を付与しました。
response列のフィルタリング: response列が「Yes,」で始まる行のみを保持し、それ以外の行を除外しました。
prompt列の文字長によるフィルタリング: prompt列の文字列の長さが80… See the full description on the dataset page: https://huggingface.co/datasets/LLMTeamAkiyama/clean_cot_verification_340k.PersonaSignal-LeakageCheck-Verification-Orientation-claude-sonnet-4-5-20250929PersonaSignal-LeakageCheck-Verification-Orientation-gpt-4oPersonaSignal-LeakageCheck-Verification-Orientation-Meta-Llama-3.1-8B-Instruct-Turboclaim_verification_training_set
Dataset Card for "20k_claims_train"
More Information needed
PersonaSignal-LeakageCheck-Verification-Orientation-gpt-4o-miniPersonaSignal-LeakageCheck-Verification-Orientation-DPO-TinkerSpeculative-Verification2026-06-23_verification_past_spectrogramThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/jogarulfop/2026-06-23_verification_past_spectrogram.retrieval_verification_roberta
Dataset Card for "retrieval_verification_roberta"
More Information needed
retrieval_verification_bm25_squeezebert_v2
Dataset Card for "retrieval_verification_bm25_squeezebert_v2"
More Information needed
datasets_for_magnetic_MTP_NatSR2024_verification
Cite this dataset Kotykhov, A. S., Gubaev, K., Hodapp, M., Tantardini, C., Shapeev, A. V., and Novikov, I. S. datasets for magnetic MTP NatSR2024 verification. ColabFit, 2024. https://doi.org/10.60732/acd42be9
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_wu6xd9i8cf7i_0
Visit the ColabFit Exchange to search additional datasets by… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/datasets_for_magnetic_MTP_NatSR2024_verification.retrieval_verification_bm25_distilbert
Dataset Card for "retrieval_verification_bm25_distilbert"
More Information needed
retrieval_verification_bm25_bert
Dataset Card for "retrieval_verification_bm25_bert"
More Information needed
retrieval_verification_bm25_roberta
Dataset Card for "retrieval_verification_bm25_roberta"
More Information needed
retrieval_verification_squeezebert
Dataset Card for "retrieval_verification_squeezebert"
More Information needed
retrieval_verification_distilbert
Dataset Card for "retrieval_verification_distilbert"
More Information needed
retrieval_verification_squeezebert_v2
Dataset Card for "retrieval_verification_squeezebert_v2"
More Information needed
math7500_train_verifications_llama3.1-8b_gt_soln_in_contextFeAl-mMTP-Verification
Cite this dataset Kotykhov, A. S., Gubaev, K., Hodapp, M., Tantardini, C., Shapeev, A. V., and Novikov, I. S. FeAl-mMTP-Verification. ColabFit, 2023. https://doi.org/None
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_5aea6ukansx8_0
Visit the ColabFit Exchange to search additional datasets by author, description, element content and… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/FeAl-mMTP-Verification.verification-bandwidth-derived-results
Verification Bandwidth Under Correlated Evaluators — derived results
Derived numerical results for the paper Verification Bandwidth Under Correlated Evaluators: What an Effective-Sample-Size Statistic Measures in an Acceptance Cascade.
Paper concept DOI: 10.5281/zenodo.21891435
Paper version DOI (v1.0.0): 10.5281/zenodo.21891436
Code: github.com/spectralbranding/orgschema-papers/tree/main/verification-bandwidth/code
This dataset's DOI: 10.57967/hf/9953
What this is… See the full description on the dataset page: https://huggingface.co/datasets/spectralbranding/verification-bandwidth-derived-results.verification_eval_verification_resultsSpeculative-Verification-Online
