datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Align-Anything-Instruction-100K-zh
Dataset Card for Align-Anything-Instruction-100K-zh
[🏠 Homepage]
[🤗 Instruction-Dataset-100K(en)]
[🤗 Instruction-Dataset-100K(zh)]
[🤗 Align-Anything Datasets]
Instruction-Dataset-100K(zh)
Highlights
Data sources:
Firefly (47.8%),
COIG (2.9%),
and our meticulously constructed QA pairs (49.3%).
100K QA pairs (zh): 104,550 meticulously crafted instructions, selected and polished from various Chinese datasets… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/Align-Anything-Instruction-100K-zh.Align-Anything-Instruction-100K
Dataset Card for Align-Anything-Instruction-100K
[🏠 Homepage]
[🤗 Instruction-Dataset-100K(en)]
[🤗 Instruction-Dataset-100K(zh)]
[🤗 Align-Anything Datasets]
Highlights
Data sources:
PKU-SafeRLHF QA ,
DialogSum,
Empathetic,
Instruction-Wild,
and Alpaca.
100K QA pairs: By leveraging GPT-4 to annotate meticulously refined instructions, we obtain 105,333 QA pairs.… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/Align-Anything-Instruction-100K.formal-anytime-valid-stats
Formal-AVS: A Lean Benchmark for Anytime-Valid Confidence-Sequence Theorem Proving
60 Lean 4 theorem targets on anytime-valid confidence sequences across four families (Howard-Ramdas, betting, Whitehouse vector, asymptotic CLT).
Benchmark Structure
60 targets grouped into tiers T0-T3 (pre-evaluation) and categories T4-T5 (empirical)
7 drafters evaluated across single-shot, agentic, and unbounded modes
14 Aristotle sessions (unbounded refinement)
Headline Results… See the full description on the dataset page: https://huggingface.co/datasets/neurips-2026-avs-bench/formal-anytime-valid-stats.anycrap
ANYCRAP: Absurdist AI-Generated Product Catalog
A dataset of 120,000+ AI-generated absurdist and creative product concepts, collected over ~1 year from the ANYCRAP catalog platform. Products range from gentle wordplay to surreal absurdism, with multilingual names, AI-generated descriptions and images, engagement data, manual category labels, and AI-powered quality scores.
What's in it
Config
Size
Description
full
126K
All products — names, descriptions… See the full description on the dataset page: https://huggingface.co/datasets/kafked/anycrap.formal-anytime-valid-stats
Formal-AVS: A Lean Benchmark for Anytime-Valid Confidence-Sequence Theorem Proving
60 Lean 4 theorem targets on anytime-valid confidence sequences across four families (Howard-Ramdas, betting, Whitehouse vector, asymptotic CLT).
Benchmark Structure
60 targets grouped into tiers T0-T3 (pre-evaluation) and categories T4-T5 (empirical)
48-target evaluated slate (headline drafter sweeps)
14 Aristotle sessions (unbounded refinement)
Headline Results (pass@5 on… See the full description on the dataset page: https://huggingface.co/datasets/athanor-ai/formal-anytime-valid-stats.
