verus
Datasets
All datasets matching “verus”Verus_Training_Data
Readme
This is our Verus training data for SAFE and VeruSyn.
SAFE's SFT data:
sft_safe_25k.json: 25K samples (11K proof generation and 14K debugging)
SAFE's SFT data:
sft_part1_6.9M.json: 6.9M samples (5.7M proof generation and 1.2M debugging)
sft_part2_4557.json: 4.6K long CoT samples (1.4K proof generation and 3.2K debugging)
Raw trajectories of VeruSyn's Part 2:
algorithmic_trajectory_9040.jsonl: 'success': 7396, 'error': 271, 'timeout': 1373… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/Verus_Training_Data.verusyn
VeruSyn — Anonymous SFT Data
Supervised fine-tuning data for training language models to write Verus proofs
(loop invariants, assertions, proof fn bodies) for Rust code, and to debug
existing Verus proofs that fail to verify.
This is the anonymous mirror used for the NeurIPS 2026 review process and
contains only the two SFT splits referenced in the paper.
Files
File
Records
Size
Purpose
sft_part1_6.9M.json
6,896,180
~15.7 GB
Large-scale synthesized SFT… See the full description on the dataset page: https://huggingface.co/datasets/anon-neurips26/verusyn.
