CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mteb /scifact SciFact An MTEB dataset Massive Text Embedding Benchmark SciFact verifies scientific claims using evidence from the research literature containing scientific paper abstracts. Task category t2t Domains Academic, Medical, Written Reference https://github.com/allenai/scifact How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["SciFact"]) evaluator = mteb.MTEB(task)… See the full description on the dataset page: https://huggingface.co/datasets/mteb/scifact.texttext-retrieval1K<n<10K5 likes22k downloads1y agoHugging Face02mteb /scidocs SCIDOCS An MTEB dataset Massive Text Embedding Benchmark SciDocs, a new evaluation benchmark consisting of seven document-level tasks ranging from citation prediction, to document classification and recommendation. Task category t2t Domains Academic, Written, Non-fiction Reference https://allenai.org/data/scidocs How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/scidocs.texttext-retrieval10K<n<100K6 likes22k downloads7mo agoHugging Face03SciCode1 /SciCodeThis dataset was presented in SciCode: A Research Coding Benchmark Curated by Scientists. textquestion-answeringn<1K16 likes16k downloads2y agoHugging Face04Weyaxi /sci-datasets Mainly science focused but other datasets exist too! Einstein models are based on this repo. text100K<n<1M28 likes7k downloads2y agoHugging Face05xw27 /scibench SciBench SciBench is a novel benchmark for college-level scientific problems sourced from instructional textbooks. The benchmark is designed to evaluate the complex reasoning capabilities, strong domain knowledge, and advanced calculation skills of LLMs. Please refer to our paper or website for full description: SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models . Citation If you find our paper useful, please cite our… See the full description on the dataset page: https://huggingface.co/datasets/xw27/scibench.textn<1K27 likes6.7k downloads2y agoHugging Face06nvidia /Nemotron-SFT-Science-v2 Dataset Description: Nemotron-Science-v2 is a science reasoning dataset with synthetic (synthetic MCQ, RQA) and non-synthetic vendor problems and LLM-generated solutions. It comprises three domains (Physics, Biology, and Chemistry), two question formats (multiple-choice questions [MCQ] and open questions [OpenQ]), and three generation setups: chain-of-thought (CoT) reasoning without tools, Python tool usage, and search tools usage with the Tavily API. The solutions were… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-SFT-Science-v2.texttext-generation1M<n<10M16 likes6.4k downloads4mo agoHugging Face07Nbardy /science-theory-textbookstext10K<n<100K9 likes5.1k downloads3y agoHugging Face08hicai-zju /SciKnowEval SciKnowEval Evaluating Multi-level Scientific Knowledge of Large Language Models Please refer to our repository and paper for more details. 博学之 ,审问之 ,慎思之 ,明辨之 ,笃行之。 —— 《礼记 · 中庸》 Doctrine of the Mean The Scientific Knowledge Evaluation (SciKnowEval) benchmark for Large Language Models (LLMs) is inspired by the profound principles outlined in the “Doctrine of the Mean” from ancient Chinese philosophy. This benchmark is designed to assess LLMs based on their proficiency in… See the full description on the dataset page: https://huggingface.co/datasets/hicai-zju/SciKnowEval.textquestion-answering10K<n<100K18 likes4.7k downloads1y agoHugging Face09scimdr /SciMDR-Evalimagequestion-answeringn<1K1 likes3k downloads7mo agoHugging Face10nvidia /Nemotron-Science-v1 Dataset Description: Nemotron-Science-v1 is a synthetic science reasoning dataset with two subsets: an MCQA set that improves on the STEM portion of Nemotron-Post-Training-v1 using GPT-OSS-120B to generate GPQA-style questions and reasoning traces, and an RQA set of synthetic chemistry questions. This dataset is ready for commercial use. The Nemotron-Science-v1 dataset contains the following subsets: MCQA This subset is an improvement of the STEM subset in… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Science-v1.text100K<n<1M32 likes2.8k downloads9mo agoHugging Face11mteb /scidocs-rerankingtext1K<n<10K2 likes1.7k downloads4y agoHugging Face12shhu2001 /SciCode-Verified SciCode-Verified SciCode-Verified is the corrected, human-verified release of the SciCode scientific-code-generation benchmark. A problem-by-problem audit identified 263 defects in the 65-problem SciCode test split and corrected every confirmable defect. The released evaluation set contains 64 main problems and 287 scored subproblems; one original problem is excluded because its specification does not determine a unique, verifiable answer. Paper: SciCode-Verified: How Benchmark… See the full description on the dataset page: https://huggingface.co/datasets/shhu2001/SciCode-Verified.texttext-generationn<1K1 likes1.7k downloads2mo agoHugging Face13R2MED /Medical-Sciences 🔭 Overview R2MED: First Reasoning-Driven Medical Retrieval Benchmark R2MED is a high-quality, high-resolution synthetic information retrieval (IR) dataset designed for medical scenarios. It contains 876 queries with three retrieval tasks, five medical scenarios, and twelve body systems. Dataset #Q #D Avg. Pos Q-Len D-Len Biology 103 57359 3.6 115.2 83.6 Bioinformatics77 47473 2.9 273.8 150.5 Medical Sciences 88 34810 2.8 107.1 122.7 MedXpertQA-Exam 97… See the full description on the dataset page: https://huggingface.co/datasets/R2MED/Medical-Sciences.texttext-retrieval10K<n<100K0 likes1.3k downloads1y agoHugging Face14Nbardy /wild-science-theory-textbookstext10K<n<100K3 likes1.1k downloads3y agoHugging Face15Hugodonotexit /math-code-science-deepseek-r1-en R1 Dataset Collection Aggregated high-quality English prompts and model-generated responses from DeepSeek R1 and DeepSeek R1-0528. Dataset Summary The R1 Dataset Collection combines multiple public DeepSeek-generated instruction-response corpora into a single, cleaned, English-only JSONL file. Each example consists of a <|user|> prompt and a <|assistant|> response in one "text" field. This release includes: ~21,000 examples from the DeepSeek-R1-0528 Distilled Custom… See the full description on the dataset page: https://huggingface.co/datasets/Hugodonotexit/math-code-science-deepseek-r1-en.textquestion-answering1M<n<10M5 likes1.1k downloads1y agoHugging Face16LLaMAX /BenchMAX_Science Dataset Sources Paper: BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models Link: https://huggingface.co/papers/2502.07346 Repository: https://github.com/CONE-MT/BenchMAX Dataset Description BenchMAX_Science is a dataset of BenchMAX, sourcing from GPQA, which evaluates the natural science reasoning capability in multilingual scenarios. We extend the original English dataset to 16 non-English languages. The data is first translated by Google… See the full description on the dataset page: https://huggingface.co/datasets/LLaMAX/BenchMAX_Science.textquestion-answering1K<n<10K2 likes957 downloads2y agoHugging Face17ScienceOne-AI /S1-Omni-Corpus-10K S1-Omni-Corpus-10K An open-source scientific multimodal reasoning dataset subset for S1-Omni 🧬 Model Introduction S1-Omni is a unified scientific multimodal reasoning model for scientific understanding, prediction, and generation. It is developed by the ScienceOne AI team of the Chinese Academy of Sciences. S1-Omni addresses fragmented scientific AI capabilities with a shared backbone for cross-disciplinary, cross-modal, and cross-task understanding and reasoning… See the full description on the dataset page: https://huggingface.co/datasets/ScienceOne-AI/S1-Omni-Corpus-10K.image10K<n<100K1 likes920 downloads2mo agoHugging Face18oss-codes /Computer-Science-Conversational-Dataset-Indictext10K<n<100K0 likes783 downloads1y agoHugging Face19sci-m-wang /C4-Eval C4-Eval C4-Eval is the evaluation set for C4 Bench, a Chengyu-based benchmark for measuring whether multimodal language models can understand cross-concept creativity. The release contains the original images, the corresponding idiom answers, and the complete task-specific questions used for evaluation. 221 base items: 37 human-designed seed figures and 184 bridge-controlled synthetic figures. 1,105 evaluation instances: five task formulations for every base item. Language:… See the full description on the dataset page: https://huggingface.co/datasets/sci-m-wang/C4-Eval.imageimage-text-to-text1K<n<10K0 likes634 downloads1mo agoHugging Face20ScienceOne-AI /SciGenEdit-10K SciGenEdit-10K An Open Dataset for Scientific Image Generation and Editing English | 简体中文 📖 Introduction SciGenEdit-10K is a public subset released with the S1-Omni-Image project. It is designed for research on scientific image generation, scientific image editing, and multi-turn scientific image generation and editing. S1-Omni-Image is a unified multimodal model developed by the ScienceOne team at the Chinese Academy of Sciences for scientific… See the full description on the dataset page: https://huggingface.co/datasets/ScienceOne-AI/SciGenEdit-10K.imagetext-to-image10K<n<100K3 likes606 downloads3mo agoHugging Face21ScienceOne-AI /S1-DeepResearch-15k S1-DeepResearch-15k Dataset Overview The S1-DeepResearch dataset is a curated collection of approximately 15k samples designed to improve deep research capabilities of large language models. The dataset includes two types of tasks: Verifiable tasks (labeled as "Closed-ended Multi-hop Resolution") Open-ended tasks (labeled as "Open-ended Exploration") Dataset Composition The dataset is organized into five core capability dimensions: Long-chain complex… See the full description on the dataset page: https://huggingface.co/datasets/ScienceOne-AI/S1-DeepResearch-15k.text10K<n<100K12 likes590 downloads5mo agoHugging Face22nvidia /Nemotron-RL-Science-v1 Dataset Description: Nemotron-RL-Science-v1 is a reinforcement learning (RL) dataset for science reasoning. Each example provides a problem, a reference answer, and a verifiable RL environment configuration (the agent prompt, the agent/verifier reference, and the answer-extraction template) so that a policy model can be trained with verifiable rewards. It covers three domains (Physics, Biology, and Chemistry), the open-question (OpenQ) format, and two generation setups:… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Science-v1.texttext-generation100K<n<1M13 likes534 downloads4mo agoHugging Face23placeholderlabs /Nemotron-SFT-Science-v2-Sharded Nemotron-SFT-Science-v2-Sharded Byte-preserving sequential 100 MB JSONL shards of selected files from nvidia/Nemotron-SFT-Science-v2. All credit for data generation and upstream curation belongs to the source authors. See the upstream dataset card for attribution, source descriptions and license terms. Included files: vendor.jsonl, so.jsonl, rqa.jsonl, syn_mcq.jsonl. No filtering, shuffling, normalization, tokenization or truncation was performed. Complete records and all… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/Nemotron-SFT-Science-v2-Sharded.text1M<n<10M0 likes514 downloads17d agoHugging Face24zd21 /SciInstructtext10K<n<100K7 likes459 downloads2y agoHugging Face25deep-principle /science_chemistrytextn<1K2 likes445 downloads12h agoHugging Face26alabnii /sciclaimeval-shared-task SciClaimEval Shared Task: All information is available at sciclaimeval.github.io Evaluation scripts & examples: github.com/SciClaimEval/sciclaimeval-shared-task More Information: paper Version Info Please use the latest version, v1.1. Changes from v1.0 to v1.1 Compared with v1.0, v1.1 includes the following changes. Removed Samples The following 20 samples have been removed: val_tab_1594 val_tab_0067… See the full description on the dataset page: https://huggingface.co/datasets/alabnii/sciclaimeval-shared-task.imagetext-classification1K<n<10K4 likes441 downloads1mo agoHugging Face27simpleG2023 /chinese-materials-science-open-intelligence 🔬 Chinese Materials Science & Metallurgy Open Intelligence Dataset Curated open intelligence dataset providing English research briefs, authoritative DOIs, executive summaries, and high-resolution micrographs of breakthrough Chinese scientific research in Materials Science, Metallurgy, Advanced Alloys, and Mining Engineering. [!IMPORTANT] Data Completeness & Research Authenticity Notice: Included in this Hugging Face Open Dataset: English structured abstracts, core… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-materials-science-open-intelligence.tabulartext-retrieval1K<n<10K0 likes391 downloads2h agoHugging Face28Zilinghan /scicode Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Zilinghan/scicode.textquestion-answeringn<1K1 likes366 downloads2y agoHugging Face29groundmore /scivideobench SciVideoBench 📄 Paper | 🌐 Project Page | 💻 Code SciVideoBench is the first comprehensive benchmark for scientific video reasoning, covering disciplines in Physics, Chemistry, Biology, and Medicine. It provides challenging multiple-choice QA pairs grounded in real scientific videos. 🔬 Overview Scientific experiments present unique challenges for video-language models (VLMs): precise perception of visual details, integration of multimodal signals (video, audio… See the full description on the dataset page: https://huggingface.co/datasets/groundmore/scivideobench.textvideo-text-to-text1K<n<10K5 likes337 downloads1y agoHugging Face30seonjeongh /science_reasoning science_reasoning Mistral-7B의 과학 지식·추론 능력 향상을 위해 6개 공개 과학 객관식 QA 데이터셋을 통일 포맷으로 변환하고, ARC-Challenge test와의 오염을 제거한 데이터셋입니다. 원본 데이터셋 allenai/sciq allenai/openbookqa (main) allenai/qasc allenai/quartz allenai/ai2_arc (ARC-Easy / ARC-Challenge) nguyen-brat/worldtree 전처리 포맷 통일: 각 데이터셋의 서로 다른 스키마를 unique_id, orig_id, source, question, choices, answer, support 필드로 변환. support는 근거 문단/문장으로, 데이터셋별 원본 필드(support/fact/para/cot)에서 구성하거나 없으면 빈 문자열.… See the full description on the dataset page: https://huggingface.co/datasets/seonjeongh/science_reasoning.textmultiple-choice10K<n<100K0 likes333 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.