datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Zephyrus
ZephyrusBench
ZephyrusBench is a weather-science benchmark released with the paper Zephyrus: An Agentic Framework for Weather Science. It contains 2,230 question-answer pairs across 49 tasks spanning geospatial reasoning, temporal reasoning, forecasting, simulation, climatology, and scientific question answering.Accepted at the International Conference on Learning Representations, 2026.
Paper and Resources
Paper: arXiv
Poster: ICLR 2026 Poster
Code: Rose-STL-Lab/Zephyrus… See the full description on the dataset page: https://huggingface.co/datasets/Rose-STL-Lab/Zephyrus.ClimaQA
ClimaQA: An Automated Evaluation Framework for Climate Question Answering Models (ICLR 2025)
Check the paper's webpage and GitHub for more info!
The ClimaQA benchmark is designed to evaluate Large Language Models (LLMs) on climate science question-answering tasks by ensuring scientific rigor and complexity. It is built from graduate-level climate science textbooks, which provide a reliable foundation for generating questions with precise terminology and complex scientific theories.… See the full description on the dataset page: https://huggingface.co/datasets/Rose-STL-Lab/ClimaQA.STL-MCQA-results
STL Prompting for Zero-Shot MCQA
Results of zero-shot Multiple-Choice Question Answering (MCQA) experiments for the paper:
Dang, Q. P., Tran-Truong, P. T., Vu, D. L., Nguyen, L. S. T., Vo, Q. T. N., & Quan, T. (2026).
Enhancing large language model performance for automatic zero-shot multiple-choice question answering via single-token logit prompting.
Computers and Education: Artificial Intelligence. DOI: 10.1016/j.caeai.2026.100578
Source code:… See the full description on the dataset page: https://huggingface.co/datasets/p-storm/STL-MCQA-results.
