datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
souslab-us-restaurant-menus
Souslab — US Restaurant Menus
A structured sample of the Souslab US restaurant menu dataset: real restaurants, real menu items, real prices — normalized into a clean schema you can train on or analyze directly.
This sample is published openly under CC-BY-NC-4.0 for research and non-commercial evaluation. The full dataset — 449,000+ US restaurants and 44.3M+ menu items, refreshed continuously with chain-level aggregation — is available via the Souslab API under commercial… See the full description on the dataset page: https://huggingface.co/datasets/AnyStackLabsdev/souslab-us-restaurant-menus.formal-anytime-valid-stats
Formal-AVS: A Lean Benchmark for Anytime-Valid Confidence-Sequence Theorem Proving
60 Lean 4 theorem targets on anytime-valid confidence sequences across four families (Howard-Ramdas, betting, Whitehouse vector, asymptotic CLT).
Benchmark Structure
60 targets grouped into tiers T0-T3 (pre-evaluation) and categories T4-T5 (empirical)
7 drafters evaluated across single-shot, agentic, and unbounded modes
14 Aristotle sessions (unbounded refinement)
Headline Results… See the full description on the dataset page: https://huggingface.co/datasets/neurips-2026-avs-bench/formal-anytime-valid-stats.anycrap
ANYCRAP: Absurdist AI-Generated Product Catalog
A dataset of 120,000+ AI-generated absurdist and creative product concepts, collected over ~1 year from the ANYCRAP catalog platform. Products range from gentle wordplay to surreal absurdism, with multilingual names, AI-generated descriptions and images, engagement data, manual category labels, and AI-powered quality scores.
What's in it
Config
Size
Description
full
126K
All products — names, descriptions… See the full description on the dataset page: https://huggingface.co/datasets/kafked/anycrap.formal-anytime-valid-stats
Formal-AVS: A Lean Benchmark for Anytime-Valid Confidence-Sequence Theorem Proving
60 Lean 4 theorem targets on anytime-valid confidence sequences across four families (Howard-Ramdas, betting, Whitehouse vector, asymptotic CLT).
Benchmark Structure
60 targets grouped into tiers T0-T3 (pre-evaluation) and categories T4-T5 (empirical)
48-target evaluated slate (headline drafter sweeps)
14 Aristotle sessions (unbounded refinement)
Headline Results (pass@5 on… See the full description on the dataset page: https://huggingface.co/datasets/athanor-ai/formal-anytime-valid-stats.AnythingLLM
