CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nvidia /Nemotron-SFT-Science-v2 Dataset Description: Nemotron-Science-v2 is a science reasoning dataset with synthetic (synthetic MCQ, RQA) and non-synthetic vendor problems and LLM-generated solutions. It comprises three domains (Physics, Biology, and Chemistry), two question formats (multiple-choice questions [MCQ] and open questions [OpenQ]), and three generation setups: chain-of-thought (CoT) reasoning without tools, Python tool usage, and search tools usage with the Tavily API. The solutions were… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-SFT-Science-v2.texttext-generation1M<n<10M16 likes6.4k downloads4mo agoHugging Face02TalentZHOU /hle_material_science HLE Material Science: A Specialized Benchmark for Materials Science A Materials Science Subset of Humanity's Last Exam (HLE) Overview HLE Material Science is a carefully curated materials science subset derived from the Humanity's Last Exam (HLE) dataset, containing 106 high-quality expert-level questions covering 25+ materials science subfields, with 97% of questions rated as high confidence. This dataset is designed to evaluate large language models'… See the full description on the dataset page: https://huggingface.co/datasets/TalentZHOU/hle_material_science.textquestion-answeringn<1K1 likes781 downloads8mo agoHugging Face03dvilasuero /natural-science-reasoning Natural Sciences Reasoning: the "smolest" reasoning dataset A smol-scale open dataset for reasoning tasks using Hugging Face Inference Endpoints. While intentionally limited in scale, this resource prioritizes: Reproducible pipeline for reasoning tasks using a variety of models (Deepseek V3, Deepsek-R1, Llama70B-Instruct, etc.) Knowledge sharing for domains other than Math and Code reasoning In this repo, you can find: The prompts and the pipeline (see the config file). The… See the full description on the dataset page: https://huggingface.co/datasets/dvilasuero/natural-science-reasoning.texttext-generationn<1K40 likes702 downloads2y agoHugging Face04islamlab /islamic-sciences islamlab — The Islamic Sciences Corpus The Islamic sciences other than Qur'an and hadith, as their authors wrote them: 4,022 works by scholars who died between the 0st and the 14th Hijri century, cut along their own chapter and biographical-entry boundaries into 1,864,389 units (3.41 billion characters of Arabic), each carrying the volume and page it sits on so a quotation can be cited rather than merely produced. Scope is Ahl al-Sunnah wa'l-Jamāʿah, and the gate is the author… See the full description on the dataset page: https://huggingface.co/datasets/islamlab/islamic-sciences.tabulartext-generation1M<n<10M2 likes648 downloads1mo agoHugging Face05nvidia /Nemotron-RL-Science-v1 Dataset Description: Nemotron-RL-Science-v1 is a reinforcement learning (RL) dataset for science reasoning. Each example provides a problem, a reference answer, and a verifiable RL environment configuration (the agent prompt, the agent/verifier reference, and the answer-extraction template) so that a policy model can be trained with verifiable rewards. It covers three domains (Physics, Biology, and Chemistry), the open-question (OpenQ) format, and two generation setups:… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Science-v1.texttext-generation100K<n<1M13 likes534 downloads4mo agoHugging Face06laion /llama-nemotron-science-reasoning-on-canonical-think-full Llama-Nemotron science reasoning — Delphi canonical-think (COMPLETE, no length filter) The complete reasoning:on science split of nvidia/Llama-Nemotron-Post-Training-Dataset, converted once into the canonical Delphi chat-template thinking format. 708,920 rows. Unlike the cold-start warmup slice open-athena/llama-nemotron-science-reasoning-on-le3000tok-100k (and its -canonical-think variant), this build applies no length cap and no subsample — every long-CoT science example is… See the full description on the dataset page: https://huggingface.co/datasets/laion/llama-nemotron-science-reasoning-on-canonical-think-full.texttext-generation100K<n<1M0 likes379 downloads19d agoHugging Face07marin-community /openthoughts4-science-26041-prompts-qwen3-30b-a3B-thinking-2507-n8-flattened-logprobs-k16 OpenThoughts-4 Science SDG: Qwen3-30B-A3B-Thinking-2507 (n=8, top-16 logprobs) Synthetic generations from Qwen/Qwen3-30B-A3B-Thinking-2507 on the Marin OpenThoughts-4 science SDG prompt set. Each prompt is sampled n=8 times, and for every generated token the dataset stores the chosen-token log probability plus the top-16 log probabilities over the vocabulary, enabling distillation, KL-style fine-tuning, reranking, and uncertainty analysis. Generation setup Field… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-science-26041-prompts-qwen3-30b-a3B-thinking-2507-n8-flattened-logprobs-k16.tabulartext-generation100K<n<1M0 likes324 downloads5mo agoHugging Face08169Pi /Science-QnA Science-QnA The Science-QnA is a large-scale, high-quality science-focused dataset (~5.63M rows) curated using synthetic data generation through distillation techniques and select open-source resources. Designed to train and evaluate reasoning-capable models in science domains with emphasis on conceptual understanding, numerical problem-solving, and exam-style Q&A patterns across Physics, Chemistry, Biology, and Mathematics. Summary • Domain: Science, Physics… See the full description on the dataset page: https://huggingface.co/datasets/169Pi/Science-QnA.texttext-generation1M<n<10M3 likes314 downloads7mo agoHugging Face09stonelight /hle_material_science HLE Material Science: A Specialized Benchmark for Materials Science A Materials Science Subset of Humanity's Last Exam (HLE) Overview HLE Material Science is a carefully curated materials science subset derived from the Humanity's Last Exam (HLE) dataset, containing 106 high-quality expert-level questions covering 25+ materials science subfields, with 97% of questions rated as high confidence. This dataset is designed to evaluate large language models'… See the full description on the dataset page: https://huggingface.co/datasets/stonelight/hle_material_science.textquestion-answeringn<1K0 likes269 downloads3mo agoHugging Face10marin-community /openthoughts4-science-26041-prompts-qwen3-32b-n8-flattened-logprobs-k16 OpenThoughts-4 Science SDG: Qwen3-32B (n=8, top-16 logprobs) Synthetic generations from Qwen/Qwen3-32B on the Marin OpenThoughts-4 science SDG prompt set. Each prompt is sampled n=8 times, and for every generated token the dataset stores the chosen-token log probability plus the top-16 log probabilities over the vocabulary, enabling distillation, KL-style fine-tuning, reranking, and uncertainty analysis. Generation setup Field Value Generator model… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-science-26041-prompts-qwen3-32b-n8-flattened-logprobs-k16.tabulartext-generation100K<n<1M0 likes239 downloads5mo agoHugging Face11Royal-lobster /10001-Science-Facts 10,001 Science Facts 10,000+ obscure, surprising, and verifiable science facts The kind that make you go "wait, really?" 🔗 GitHub Repository • 📁 Download by Category 🤔 What is this? A curated dataset of 10,003 science facts across 32 categories — from quantum physics to parasites to the history of food. Every fact is: Sourced — from Wikipedia, Wikidata, academic sources Verifiable — no LLM hallucinations Surprising — passes the "dinner party test"… See the full description on the dataset page: https://huggingface.co/datasets/Royal-lobster/10001-Science-Facts.texttext-generation10K<n<100K1 likes238 downloads8mo agoHugging Face12stindardlogic /science-qa-sft-100k Science QA SFT (100K) 100,000 science Q&A examples with step-by-step explanations for SFT fine-tuning. Covers physics, chemistry, biology, astronomy, and earth science at beginner through advanced difficulty. Motivation Models trained on general text often give superficially plausible but mechanistically wrong answers to science questions — stating the right conclusion without understanding the underlying reasoning. This dataset trains models to explain why an… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/science-qa-sft-100k.texttext-generation100K<n<1M0 likes234 downloads2mo agoHugging Face13Ik45 /data-science-en-id Data Science EN-ID Parallel Corpus (Scientific Domain) Dataset Description This dataset is a curated English-Indonesian (EN-ID) parallel corpus specifically designed for the Scientific and Data Science domains. It was developed to support the training of Machine Translation (NMT) models and Large Language Models (LLMs) to better handle technical terminology, academic structures, and formal scientific language. Primary Languages: English (EN) and Indonesian (ID) Domain:… See the full description on the dataset page: https://huggingface.co/datasets/Ik45/data-science-en-id.texttext-generation10M<n<100M0 likes190 downloads6mo agoHugging Face14ianncity /GLM-5.2-Science GLM-5.2 · Science-50000x 50,000x traces distilled from GLM-5.2 on High reasoning Physics · Chemistry · Biology Token Count: 160M Theres prompt overlap with my Kimi K2.5 dataset science subset, which I think those prompts are getting used in alot of places now You can use this dataset for any purpose and you dont need to credit me, preferably dont claim it as your own. hi - ianncity texttext-generation10K<n<100K19 likes176 downloads2mo agoHugging Face15Lots-of-LoRAs /task047_miscellaneous_answering_science_questions Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task047_miscellaneous_answering_science_questions Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task047_miscellaneous_answering_science_questions.texttext-generationn<1K0 likes164 downloads2y agoHugging Face16KellanF89 /Cannabis_Science_Data Cannabis Science Literature QA Dataset This dataset contains 161,170 high-quality question-answer pairs derived from over 400 peer-reviewed cannabis science research papers and textbooks. Created to advance AI research in cannabis science and medical applications, it provides a comprehensive resource for training language models on cannabis-related scientific knowledge. Dataset Details Dataset Description This dataset was systematically generated from a curated… See the full description on the dataset page: https://huggingface.co/datasets/KellanF89/Cannabis_Science_Data.question-answering100K<n<1M5 likes112 downloads1y agoHugging Face17tilikumotp /Global-Ocean-Science-Corpus 🌊 Global-Ocean-Science-Corpus (v2.0 Curated & Cleaned) A Highly Curated, Large-Scale Pre-Training & RAG Corpus for Deep Ocean Sciences, Marine Biology, and Oceanography Language Note: This dataset is a 100% English-language scientific corpus (language: "en") aggregating peer-reviewed literature, deep-sea exploration dossiers, and technical oceanographic reports from leading global marine institutes. Global-Ocean-Science-Corpus, derin okyanus bilimleri, deniz biyolojisi… See the full description on the dataset page: https://huggingface.co/datasets/tilikumotp/Global-Ocean-Science-Corpus.tabulartext-generation10K<n<100K0 likes84 downloads2d agoHugging Face18Lots-of-LoRAs /task701_mmmlu_answer_generation_high_school_computer_science Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task701_mmmlu_answer_generation_high_school_computer_science Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task701_mmmlu_answer_generation_high_school_computer_science.texttext-generationn<1K1 likes82 downloads2y agoHugging Face19islamlab /islamic-sciences-training islamlab — Islamic Sciences Training Sets Training data derived from islamlab/islamic-sciences: text for domain adaptation, retrieval pairs with hard negatives, and citation questions whose answers are read out of the corpus rather than written by a model. Nothing here is generated. Questions come from a fixed set of templates and every answer is a field already present in the corpus. That buys a narrow dataset in exchange for one that cannot teach a model a fact the sources do… See the full description on the dataset page: https://huggingface.co/datasets/islamlab/islamic-sciences-training.texttext-generation1M<n<10M0 likes82 downloads1mo agoHugging Face20percepteyeAI /10001-Science-Facts 10,001 Science Facts 10,000+ obscure, surprising, and verifiable science facts The kind that make you go "wait, really?" 🔗 GitHub Repository • 📁 Download by Category 🤔 What is this? A curated dataset of 10,003 science facts across 32 categories — from quantum physics to parasites to the history of food. Every fact is: Sourced — from Wikipedia, Wikidata, academic sources Verifiable — no LLM hallucinations Surprising — passes the "dinner party test"… See the full description on the dataset page: https://huggingface.co/datasets/percepteyeAI/10001-Science-Facts.texttext-generation10K<n<100K1 likes81 downloads7mo agoHugging Face21StarpowerTechnology /Dense-Information-Science-Physics-Dataset Dense Information With Multiple Fine-tuned Variations This dataaset has multiple for each input to learn how to express the same answer in different ways Dataset Structure The dataset contains two columns: Column Description input A science or quantum-physics question output A conversational answer to the question Example: { "input": "What is quantum entanglement?", "output": "Quantum entanglement is when two quantum systems share one… See the full description on the dataset page: https://huggingface.co/datasets/StarpowerTechnology/Dense-Information-Science-Physics-Dataset.texttext-generation1K<n<10K0 likes74 downloads13d agoHugging Face22TheJackBright /verisci-verified-science-math-code VeriSci Verified Science Math Code Verifier-grounded dataset for the Adaption AutoScientist Challenge Part 2, targeting the Science category with secondary Math and Code coverage. Summary VeriSci trains models to solve scientific computations, finite-difference PDE updates, numerical ODE steps, unit-checked mechanics, thermodynamics, circuits, chemistry stoichiometry, molarity, unit conversion, vector decomposition, two-point linear modeling, small Python… See the full description on the dataset page: https://huggingface.co/datasets/TheJackBright/verisci-verified-science-math-code.texttext-generation1K<n<10K0 likes71 downloads2mo agoHugging Face23marin-community /openthoughts4-science-26041-prompts-qwen3-4b-n8-flattened-logprobs-k16 OpenThoughts-4 Science SDG: Qwen3-4B (n=8, top-16 logprobs) Synthetic generations from Qwen/Qwen3-4B on the Marin OpenThoughts-4 science SDG prompt set. Each prompt is sampled n=8 times, and for every generated token the dataset stores the chosen-token log probability plus the top-16 log probabilities over the vocabulary, enabling distillation, KL-style fine-tuning, reranking, and uncertainty analysis. Generation setup Field Value Generator model… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-science-26041-prompts-qwen3-4b-n8-flattened-logprobs-k16.tabulartext-generation100K<n<1M0 likes70 downloads5mo agoHugging Face24expertdata-factory /science-cot-datasetgated ExpertData Science — Scientific Reasoning Expert-Annotated · Rights-Cleared · Ground-Truth Verified · PII-Clean Each record captures a complete experimental or theoretical reasoning chain: Hypothesis → Methodology → Causal Chain → Validated Conclusion. Extracted from peer-reviewed papers across physics, biology, materials science, astrophysics, and neuroscience using structured scientific-reasoning extraction. This dataset is produced by the ExpertData-Factory pipeline (Mine →… See the full description on the dataset page: https://huggingface.co/datasets/expertdata-factory/science-cot-dataset.textquestion-answeringn<1K0 likes64 downloads7mo agoHugging Face25ytu-ce-cosmos /tubitak-science-olympiad-tr TUBITAK Science Olympiad Dataset This dataset contains multiple-choice and open-ended scientific questions sourced from the TUBITAK (The Scientific and Technological Research Council of Turkey) Science Olympiads spanning various years. It is intended to serve as a benchmark for evaluating the advanced analytical, mathematical, and computational reasoning capabilities of Large Language Models (LLMs) in the Turkish language. The dataset comprises approximately 2700 problems across… See the full description on the dataset page: https://huggingface.co/datasets/ytu-ce-cosmos/tubitak-science-olympiad-tr.imagequestion-answering1K<n<10K13 likes64 downloads6mo agoHugging Face26batuhanozkose /Rehber-CoT-Science 🧬 Rehber-CoT-Science: Turkish Scientific Reasoning Dataset Turkish Scientific Computational Reasoning (Chain-of-Thought) Dataset Multi-step scientific problem-solving dataset with verifiable Python code and detailed explanations Dataset • Author 📌 Changelog Eski sürümlere erişim: Branch menüsünden v1 seçebilirsiniz. Version Date Changes v2.0 24.12.2025 ✨ Yeni explained_answer alanı eklendi, Statistics domain eklendi, 712 örneğe genişletildi… See the full description on the dataset page: https://huggingface.co/datasets/batuhanozkose/Rehber-CoT-Science.textquestion-answering1K<n<10K4 likes63 downloads9mo agoHugging Face27flamiinngo /applied-science-qa Applied Science QA 2,472 question and answer pairs across seven scientific disciplines, written by practitioners and scored by their peers on the Stack Exchange network. Used to fine-tune a model for the Adaption Labs AutoScientist Challenge, Part 2, Science category. The resulting model beat its own base model 60 to 40 on Adaption's held-out Science test set. What is in it Column Description question The question title plus the opening of the asker's… See the full description on the dataset page: https://huggingface.co/datasets/flamiinngo/applied-science-qa.question-answering1K<n<10K0 likes62 downloads1mo agoHugging Face28agentlans /nvidia-Nemotron-Science-Math NVIDIA Nemotron Science and Math Reasoning This is an unofficial, curated collection derived from NVIDIA's open-source Nemotron datasets. It is designed specifically to train language models in complex scientific and mathematical reasoning by providing structured chain-of-thought (CoT) examples. To ensure efficiency, the shortest available CoT sequence was chosen for each question, filtering out redundant variations while preserving the core logical progression toward the final… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/nvidia-Nemotron-Science-Math.texttext-generation100K<n<1M0 likes59 downloads5mo agoHugging Face29graf /bvg_science_qwen_4b_not_easy BVG Qwen3-4B science not-easy prompts This is the exact two-split dataset artifact used by the BVG Qwen3-4B science experiments. It is published as a DatasetDict with train and validation splits. Split Rows SHA-256 of canonical JSON rows train 11,121 e985f63334809b43e8cffa971829bff6c2159fe08d79630b2fbbda9d22bc0831 validation 997 c2ea31a3e676b2c28c72daaddb9023c9163cd6671c7cf3afd2e305f7fc206480 The active experiment TOMLs consume the complete train split. Their… See the full description on the dataset page: https://huggingface.co/datasets/graf/bvg_science_qwen_4b_not_easy.texttext-generation10K<n<100K0 likes59 downloads13d agoHugging Face30jeffmeloy /sonnet3.5_science_conversationsThis dataset features sharegpt structured dialogues focused on a variety of advanced scientific topics. The content reflects a high level of scientific expertise, providing in-depth information on complex subjects. texttext-generation1K<n<10K23 likes58 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.