datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
noidea-wisdom-v1wisdombench
WisdomBench and Wisdom Science Data
WisdomBench is a longitudinal benchmark for measuring whether an AI agent changes after repeated exposure to feedback and failure.
This dataset contains 3,600 scored evaluation events under the included conditions (3 models x 4 strategies x 20 tasks x 5 rounds x 3 seeds).
It also mirrors the Wisdom Science Research Portfolio release:
Zenodo record: https://zenodo.org/records/20027295
Portfolio DOI: 10.5281/zenodo.20027295
Portfolio folder:… See the full description on the dataset page: https://huggingface.co/datasets/MMJBDS/wisdombench.ADG-Qwen2.5-Alpaca-GPT4ADG-Qwen2.5-CoTADG-LLaMa3-Alpaca-GPT4ADG-LLaMa3-WizardLMADG-LLaMa3-CoTADG-Qwen2.5-WizardLM
