datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sai-mash
Multilingual Audits: Structured & Harmonized
MASH is a dataset of Supreme Audit Institution (SAI) reports harmonized to a common language and format. SAIs publish their work in national languages and with varying structures, making cross-country analysis difficult. MASH resolves this by processing each report through a standardized pipeline that produces English summaries, structured metadata, and controlled-vocabulary tags — enabling researchers, auditors, and developers to… See the full description on the dataset page: https://huggingface.co/datasets/Riksrevisjonen/sai-mash.unsolved-math-clean
🧠 Unsolved Math — Clean
8,626 curated open research problems in mathematics and CS — including 122 Millennium Prize Problems — deduplicated, schema-flattened, and packaged as proper parquet configs with an eval-only benchmark view.
A reasoning frontier dataset: every problem here is actually unsolved or partially solved — ideal for honest capability probing instead of contaminated benchmarks.
Clean derivative of ulamai/UnsolvedMath (8,785 problems). License unchanged:… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/unsolved-math-clean.PulmoBench
PulmoBench: A Structured Clinical Reasoning Benchmark for Pulmonology
Overview
PulmoBench is a structured benchmark designed to evaluate Large Language Models (LLMs) in clinical risk stratification, safety-aware pulmonary reasoning, and zero-shot diagnostic generalization.
The benchmark features two evaluation tracks:
Core Benchmark (train, validation, test): Synthesized natural-language clinical vignettes with structured risk tiers, escalation flags, and… See the full description on the dataset page: https://huggingface.co/datasets/saibhossain/PulmoBench.
