datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SciKnowEval
SciKnowEval
Evaluating Multi-level Scientific Knowledge of Large Language Models
Please refer to our repository and paper for more details.
博学之 ,审问之 ,慎思之 ,明辨之 ,笃行之。
—— 《礼记 · 中庸》 Doctrine of the Mean
The Scientific Knowledge Evaluation (SciKnowEval) benchmark for Large Language Models (LLMs) is inspired by the profound principles outlined in the “Doctrine of the Mean” from ancient Chinese philosophy. This benchmark is designed to assess LLMs based on their proficiency in… See the full description on the dataset page: https://huggingface.co/datasets/hicai-zju/SciKnowEval.SciKnowEval
SciKnowEval
Evaluating Multi-level Scientific Knowledge of Large Language Models
Please refer to our repository and paper for more details.
博学之 ,审问之 ,慎思之 ,明辨之 ,笃行之。
—— 《礼记 · 中庸》 Doctrine of the Mean
The Scientific Knowledge Evaluation (SciKnowEval) benchmark for Large Language Models (LLMs) is inspired by the profound principles outlined in the “Doctrine of the Mean” from ancient Chinese philosophy. This benchmark is designed to assess LLMs based on their proficiency in… See the full description on the dataset page: https://huggingface.co/datasets/Xh1Xxhg/SciKnowEval.sciknoweval-v2-hard-autogradable-512-2026-04-28
SciKnowEval v2 Hard Autogradable 512 - 2026-04-28
A 512-example sanity subset sampled from hicai-zju/SciKnowEval (v2, test) for Plan-CRL scientific reasoning evals.
Selection seed: 20260428.
Filtering and balancing:
excludes L1
keeps L2, L3, L4
keeps autogradable types: mcq-4-choices, mcq-2-choices, true_or_false, filling
requires answerKey or answer
balances domains at 128 examples each: Biology, Chemistry, Material, Physics
per domain: 32 L2, 48 L3, 48 L4
Useful fields for… See the full description on the dataset page: https://huggingface.co/datasets/1337xyz1337xyz/sciknoweval-v2-hard-autogradable-512-2026-04-28.
