CoolFace
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hicai-zju /SciKnowEval SciKnowEval Evaluating Multi-level Scientific Knowledge of Large Language Models Please refer to our repository and paper for more details. 博学之 ,审问之 ,慎思之 ,明辨之 ,笃行之。 —— 《礼记 · 中庸》 Doctrine of the Mean The Scientific Knowledge Evaluation (SciKnowEval) benchmark for Large Language Models (LLMs) is inspired by the profound principles outlined in the “Doctrine of the Mean” from ancient Chinese philosophy. This benchmark is designed to assess LLMs based on their proficiency in… See the full description on the dataset page: https://huggingface.co/datasets/hicai-zju/SciKnowEval.textquestion-answering10K<n<100K18 likes4.7k downloads1y agoHugging Face02ShAIkespear /SciKnowEval_mcqa Dataset Card for SciKnowEval_mcqa Dataset Description This dataset is a modified version of the original SciKnowEval dataset. SciKnowEval is a comprehensive dataset designed to evaluate the scientific knowledge reasoning capabilities of Large Language Models (LLMs). It spans primarily across a few domains (Physics, Chemistry, Biology, Materials). Modifications in this Version In this release, we have curated this dataset to focus only on MCQA questions… See the full description on the dataset page: https://huggingface.co/datasets/ShAIkespear/SciKnowEval_mcqa.textmultiple-choice10K<n<100K4 likes104 downloads10mo agoHugging Face03alvinming /sciknoweval_l3text1K<n<10K0 likes78 downloads7mo agoHugging Face04Xh1Xxhg /SciKnowEval SciKnowEval Evaluating Multi-level Scientific Knowledge of Large Language Models Please refer to our repository and paper for more details. 博学之 ,审问之 ,慎思之 ,明辨之 ,笃行之。 —— 《礼记 · 中庸》 Doctrine of the Mean The Scientific Knowledge Evaluation (SciKnowEval) benchmark for Large Language Models (LLMs) is inspired by the profound principles outlined in the “Doctrine of the Mean” from ancient Chinese philosophy. This benchmark is designed to assess LLMs based on their proficiency in… See the full description on the dataset page: https://huggingface.co/datasets/Xh1Xxhg/SciKnowEval.textquestion-answering10K<n<100K0 likes48 downloads5mo agoHugging Face05guanning-ai /sciknoweval_l3text1K<n<10K0 likes24 downloads7mo agoHugging Face06violetxi /sciKnowEval-OEtext1K<n<10K0 likes23 downloads4mo agoHugging Face07budecosystem /SciKnowEval SciKnowEval — evaluation data (OpenCompass format) Bud Ecosystem eval mirror (config SciKnowEval_gen). Source hicai-zju/SciKnowEval — license MIT, unchanged; all rights remain with the original authors. text10K<n<100K0 likes22 downloads2mo agoHugging Face081337xyz1337xyz /sciknoweval-v2-hard-autogradable-512-2026-04-28 SciKnowEval v2 Hard Autogradable 512 - 2026-04-28 A 512-example sanity subset sampled from hicai-zju/SciKnowEval (v2, test) for Plan-CRL scientific reasoning evals. Selection seed: 20260428. Filtering and balancing: excludes L1 keeps L2, L3, L4 keeps autogradable types: mcq-4-choices, mcq-2-choices, true_or_false, filling requires answerKey or answer balances domains at 128 examples each: Biology, Chemistry, Material, Physics per domain: 32 L2, 48 L3, 48 L4 Useful fields for… See the full description on the dataset page: https://huggingface.co/datasets/1337xyz1337xyz/sciknoweval-v2-hard-autogradable-512-2026-04-28.tabularquestion-answeringn<1K0 likes15 downloads5mo agoHugging Face09graf /olmo_sft_sciknoweval_testtext10K<n<100K0 likes12 downloads2mo agoHugging Face10Ba2han /Sciknoweval-mcqa_Turkish Translated with: Karga-DPO-v0.1 Broken translations were cleaned off. Beware of hallucinations though. text10K<n<100K0 likes8 downloads9mo agoHugging Face11Sujal0077 /sciknowevaltext1K<n<10K0 likes4 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.