datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SciKnowEval
SciKnowEval
Evaluating Multi-level Scientific Knowledge of Large Language Models
Please refer to our repository and paper for more details.
博学之 ,审问之 ,慎思之 ,明辨之 ,笃行之。
—— 《礼记 · 中庸》 Doctrine of the Mean
The Scientific Knowledge Evaluation (SciKnowEval) benchmark for Large Language Models (LLMs) is inspired by the profound principles outlined in the “Doctrine of the Mean” from ancient Chinese philosophy. This benchmark is designed to assess LLMs based on their proficiency in… See the full description on the dataset page: https://huggingface.co/datasets/hicai-zju/SciKnowEval.SciKnowEval
SciKnowEval
Evaluating Multi-level Scientific Knowledge of Large Language Models
Please refer to our repository and paper for more details.
博学之 ,审问之 ,慎思之 ,明辨之 ,笃行之。
—— 《礼记 · 中庸》 Doctrine of the Mean
The Scientific Knowledge Evaluation (SciKnowEval) benchmark for Large Language Models (LLMs) is inspired by the profound principles outlined in the “Doctrine of the Mean” from ancient Chinese philosophy. This benchmark is designed to assess LLMs based on their proficiency in… See the full description on the dataset page: https://huggingface.co/datasets/Xh1Xxhg/SciKnowEval.SciKnowEval
SciKnowEval — evaluation data (OpenCompass format)
Bud Ecosystem eval mirror (config SciKnowEval_gen). Source hicai-zju/SciKnowEval — license MIT, unchanged; all rights remain with the original authors.
sciknoweval
