sciknoweval
bs16-k10-lr5e-7-ema0.01-eopd0.8-qwen3-4b-think-sciknoweval_chem_sensitive20pct_nogapk10-lr5e-7-ema0-eopd0.8-sciknoweval_bio_sensitive100pct-pos_gap20pctk10-lr5e-7-ema0-eopd0.8-sciknoweval_chem_sensitive20pct-pos_gap20pctk10-lr5e-7-ema0-eopd0.8-sciknoweval_physics_sensitive100pct-pos_gap20pctbs16-k10-lr5e-7-ema0.01-eopd0.8-qwen3-4b-think-sciknoweval_material_bottom20_nogapbs16-k10-lr5e-7-ema0.01-eopd0.8-qwen3-4b-think-sciknoweval_material_sensitive20pct_nogap-maxstepk10-lr5e-7-ema0.01-eopd0.8-sciknoweval_material_sensitive20pct-pos_gap20pctbs16-k10-lr5e-7-ema0.01-eopd0.8-qwen3-4b-think-sciknoweval_material_sensitive20pct_nogap
SciKnowEval
SciKnowEval
Evaluating Multi-level Scientific Knowledge of Large Language Models
Please refer to our repository and paper for more details.
博学之 ,审问之 ,慎思之 ,明辨之 ,笃行之。
—— 《礼记 · 中庸》 Doctrine of the Mean
The Scientific Knowledge Evaluation (SciKnowEval) benchmark for Large Language Models (LLMs) is inspired by the profound principles outlined in the “Doctrine of the Mean” from ancient Chinese philosophy. This benchmark is designed to assess LLMs based on their proficiency in… See the full description on the dataset page: https://huggingface.co/datasets/hicai-zju/SciKnowEval.SciKnowEval_mcqa
Dataset Card for SciKnowEval_mcqa
Dataset Description
This dataset is a modified version of the original SciKnowEval dataset.
SciKnowEval is a comprehensive dataset designed to evaluate the scientific knowledge reasoning capabilities of Large Language Models (LLMs). It spans primarily across a few domains (Physics, Chemistry, Biology, Materials).
Modifications in this Version
In this release, we have curated this dataset to focus only on MCQA questions… See the full description on the dataset page: https://huggingface.co/datasets/ShAIkespear/SciKnowEval_mcqa.sciknoweval_l3SciKnowEval
SciKnowEval
Evaluating Multi-level Scientific Knowledge of Large Language Models
Please refer to our repository and paper for more details.
博学之 ,审问之 ,慎思之 ,明辨之 ,笃行之。
—— 《礼记 · 中庸》 Doctrine of the Mean
The Scientific Knowledge Evaluation (SciKnowEval) benchmark for Large Language Models (LLMs) is inspired by the profound principles outlined in the “Doctrine of the Mean” from ancient Chinese philosophy. This benchmark is designed to assess LLMs based on their proficiency in… See the full description on the dataset page: https://huggingface.co/datasets/Xh1Xxhg/SciKnowEval.sciKnowEval-OEsciknoweval_l3
