neurips-2026-avs-bench/formal-anytime-valid-stats
Formal-AVS: A Lean Benchmark for Anytime-Valid Confidence-Sequence Theorem Proving 60 Lean 4 theorem targets on anytime-valid confidence sequences across four families (Howard-Ramdas, betting, Whitehouse vector, asymptotic CLT). Benchmark Structure 60 targets grouped into tiers T0-T3 (pre-evaluation) and categories T4-T5 (empirical) 7 drafters evaluated across single-shot, agentic, and unbounded modes 14 Aristotle sessions (unbounded refinement)… See the full description on the dataset page: https://huggingface.co/datasets/neurips-2026-avs-bench/formal-anytime-valid-stats.
fix: add prov:wasDerivedFrom + prov:wasGeneratedBy (exact spec field names)
fix: RAI as markdown sections + update 48→60 + add Opus
fix: structured cr:RAI block with sourceDatasets + provenanceActivities
fix: add rai:dataSource + rai:dataCollectionProcess (spec-compliant field names)
fix: add missing RAI fields (sourceDatasets + provenanceActivities)
fix: real theorem names for all 60 targets + T0-T5 tiers + remove identity leaks
add: Croissant file with RAI metadata (NeurIPS 2026 requirement)
fix: remove identity leak (closed_by_asabi_local → closed_locally)
Add aristotle_history.jsonl (14 sessions)
Update dataset card with full headline results table
Add dataset card
Initial upload: 60-target benchmark with drafter + agentic results
initial commit
