CoolFace
22 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nvidia /Nemotron-SFT-Science-v2 Dataset Description: Nemotron-Science-v2 is a science reasoning dataset with synthetic (synthetic MCQ, RQA) and non-synthetic vendor problems and LLM-generated solutions. It comprises three domains (Physics, Biology, and Chemistry), two question formats (multiple-choice questions [MCQ] and open questions [OpenQ]), and three generation setups: chain-of-thought (CoT) reasoning without tools, Python tool usage, and search tools usage with the Tavily API. The solutions were… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-SFT-Science-v2.texttext-generation1M<n<10M16 likes6.3k downloads4mo agoHugging Face02nvidia /Nemotron-RL-Science-v1 Dataset Description: Nemotron-RL-Science-v1 is a reinforcement learning (RL) dataset for science reasoning. Each example provides a problem, a reference answer, and a verifiable RL environment configuration (the agent prompt, the agent/verifier reference, and the answer-extraction template) so that a policy model can be trained with verifiable rewards. It covers three domains (Physics, Biology, and Chemistry), the open-question (OpenQ) format, and two generation setups:… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Science-v1.texttext-generation100K<n<1M13 likes525 downloads4mo agoHugging Face03Royal-lobster /10001-Science-Facts 10,001 Science Facts 10,000+ obscure, surprising, and verifiable science facts The kind that make you go "wait, really?" 🔗 GitHub Repository • 📁 Download by Category 🤔 What is this? A curated dataset of 10,003 science facts across 32 categories — from quantum physics to parasites to the history of food. Every fact is: Sourced — from Wikipedia, Wikidata, academic sources Verifiable — no LLM hallucinations Surprising — passes the "dinner party test"… See the full description on the dataset page: https://huggingface.co/datasets/Royal-lobster/10001-Science-Facts.texttext-generation10K<n<100K1 likes237 downloads8mo agoHugging Face04stindardlogic /science-qa-sft-100k Science QA SFT (100K) 100,000 science Q&A examples with step-by-step explanations for SFT fine-tuning. Covers physics, chemistry, biology, astronomy, and earth science at beginner through advanced difficulty. Motivation Models trained on general text often give superficially plausible but mechanistically wrong answers to science questions — stating the right conclusion without understanding the underlying reasoning. This dataset trains models to explain why an… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/science-qa-sft-100k.texttext-generation100K<n<1M0 likes237 downloads2mo agoHugging Face05ianncity /GLM-5.2-Science GLM-5.2 · Science-50000x 50,000x traces distilled from GLM-5.2 on High reasoning Physics · Chemistry · Biology Token Count: 160M Theres prompt overlap with my Kimi K2.5 dataset science subset, which I think those prompts are getting used in alot of places now You can use this dataset for any purpose and you dont need to credit me, preferably dont claim it as your own. hi - ianncity texttext-generation10K<n<100K19 likes204 downloads2mo agoHugging Face06percepteyeAI /10001-Science-Facts 10,001 Science Facts 10,000+ obscure, surprising, and verifiable science facts The kind that make you go "wait, really?" 🔗 GitHub Repository • 📁 Download by Category 🤔 What is this? A curated dataset of 10,003 science facts across 32 categories — from quantum physics to parasites to the history of food. Every fact is: Sourced — from Wikipedia, Wikidata, academic sources Verifiable — no LLM hallucinations Surprising — passes the "dinner party test"… See the full description on the dataset page: https://huggingface.co/datasets/percepteyeAI/10001-Science-Facts.texttext-generation10K<n<100K1 likes79 downloads7mo agoHugging Face07jeffmeloy /sonnet3.5_science_conversationsThis dataset features sharegpt structured dialogues focused on a variety of advanced scientific topics. The content reflects a high level of scientific expertise, providing in-depth information on complex subjects. texttext-generation1K<n<10K23 likes70 downloads2y agoHugging Face08batuhanozkose /Rehber-CoT-Science 🧬 Rehber-CoT-Science: Turkish Scientific Reasoning Dataset Turkish Scientific Computational Reasoning (Chain-of-Thought) Dataset Multi-step scientific problem-solving dataset with verifiable Python code and detailed explanations Dataset • Author 📌 Changelog Eski sürümlere erişim: Branch menüsünden v1 seçebilirsiniz. Version Date Changes v2.0 24.12.2025 ✨ Yeni explained_answer alanı eklendi, Statistics domain eklendi, 712 örneğe genişletildi… See the full description on the dataset page: https://huggingface.co/datasets/batuhanozkose/Rehber-CoT-Science.textquestion-answering1K<n<10K4 likes64 downloads9mo agoHugging Face09agentlans /nvidia-Nemotron-Science-Math NVIDIA Nemotron Science and Math Reasoning This is an unofficial, curated collection derived from NVIDIA's open-source Nemotron datasets. It is designed specifically to train language models in complex scientific and mathematical reasoning by providing structured chain-of-thought (CoT) examples. To ensure efficiency, the shortest available CoT sequence was chosen for each question, filtering out redundant variations while preserving the core logical progression toward the final… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/nvidia-Nemotron-Science-Math.texttext-generation100K<n<1M0 likes63 downloads5mo agoHugging Face10TheJackBright /verisci-verified-science-math-code VeriSci Verified Science Math Code Verifier-grounded dataset for the Adaption AutoScientist Challenge Part 2, targeting the Science category with secondary Math and Code coverage. Summary VeriSci trains models to solve scientific computations, finite-difference PDE updates, numerical ODE steps, unit-checked mechanics, thermodynamics, circuits, chemistry stoichiometry, molarity, unit conversion, vector decomposition, two-point linear modeling, small Python… See the full description on the dataset page: https://huggingface.co/datasets/TheJackBright/verisci-verified-science-math-code.texttext-generation1K<n<10K0 likes61 downloads2mo agoHugging Face11stindardlogic /data-science-workflows-sft-100k Data Science Workflows SFT (100K) 100,000 ShareGPT conversations demonstrating expert-level data science practice across data cleaning, EDA, ML pipelines, feature engineering, SQL analytics, statistical analysis, model evaluation, visualization, and production deployment. Motivation Data science is one of the most in-demand technical skills — companies need models that can reason through real analytical problems with the rigor of a senior data scientist. Models… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/data-science-workflows-sft-100k.texttext-generation100K<n<1M0 likes38 downloads2mo agoHugging Face12AngelWarmSmile123 /deep-physics-science-zh Deep Physics & Science Dialogue Dataset (Chinese) 深度物理科学对话数据集 Dataset Description High-quality Chinese physics and science dialogues covering quantum gravity, theory of everything, relativity, quantum mechanics, and entropy/information theory. 高质量中文物理科学对话,涵盖量子引力理论、万物理论、相对论、量子力学、熵与信息论等硬核科学议题。 Dataset Structure Format: JSONL (JSON Lines) Fields: instruction: User message / question input: Additional context (if any) output: AI response… See the full description on the dataset page: https://huggingface.co/datasets/AngelWarmSmile123/deep-physics-science-zh.texttext-generation1K<n<10K1 likes37 downloads3mo agoHugging Face13Somtharu181coder /science_behavioral_and_domain_diversity_dataset Nepali Science SFT Dataset — Clean Candidate A high-quality Nepali Science Supervised Fine-Tuning (SFT) dataset containing short question–answer instruction-following examples written primarily in Nepali Devanagari script. This release is the clean candidate produced after structural validation, language checks, duplicate analysis, and Unicode-contamination filtering. Dataset Overview Property Value Dataset file clean_candidate.jsonl Records 29,320… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/science_behavioral_and_domain_diversity_dataset.texttext-generation10K<n<100K0 likes33 downloads1mo agoHugging Face14pthinc /BCE-Prettybird-Middle-Science-v0.1 BCE-Prettybird-Middle-Science-v0.1 - 101000 Science Q&A Dataset for Instruction-Based Learning We are excited to introduce a comprehensive math-physics-chemistry-biology dataset containing 100500 instruction-based question-answer pairs, designed to support research in science reasoning, problem-solving, and AI training. Generated using Python’s math libraries (e.g., math, numpy, sympy), the dataset covers a diverse range of difficulty levels—from basic arithmetic and algebra to… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Middle-Science-v0.1.texttext-classification100K<n<1M0 likes31 downloads5mo agoHugging Face15pthinc /BCE-Prettybird-Nano-Science-v0.1 BCE-Prettybird-Nano-Science-v0.1 - 500 Science Q&A Dataset for Instruction-Based Learning We are excited to introduce a comprehensive math-physics-chemistry-biology dataset containing 500 instruction-based question-answer pairs, designed to support research in science reasoning, problem-solving, and AI training. Generated using Python’s math libraries (e.g., math, numpy, sympy), the dataset covers a diverse range of difficulty levels—from basic arithmetic and algebra to advanced… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Science-v0.1.texttext-classificationn<1K0 likes29 downloads6mo agoHugging Face16issdandavis /scbe-life-science-research-training-demo Status: experimental. Experiment-specific slice. Primary public dataset: scbe-aethermoore-training-data. SCBE Research Training Package This package was generated from live pubmed pulls for the query protein structure prediction and is meant for lightweight Hugging Face dataset and SFT experiments. Files papers.jsonl: normalized raw research records sft_train.jsonl: train split for instruction-style tasks sft_validation.jsonl: validation split… See the full description on the dataset page: https://huggingface.co/datasets/issdandavis/scbe-life-science-research-training-demo.texttext-generationn<1K0 likes25 downloads2mo agoHugging Face17RafaelUI /ru_scienceThe dataset is based on russian scientific articles. Data filtering was easy, there may be garbage text texttext-generation10K<n<100K2 likes25 downloads3mo agoHugging Face18Rishidar /autoscientist-science-dataset AutoScientist adapted dataset — science Adaption Labs AutoScientist v5 adapted fine-tuning data for the science category. science_adapted.jsonl — prompt/completion pairs used for QLoRA SFT. science_v5_raw.csv — full Adaption output (prompt, completion, enhanced_prompt, chosen, rejected, reasoning_trace, embeddings) used for DPO. Paired weights: Rishidar/autoscientist-science-qlora (Kaggle mirror rishidard/autoscientist-science-qlora). texttext-generationn<1K0 likes18 downloads3mo agoHugging Face19Hamzasajjad38 /data-science-chatbot 📊 Data Science Chatbot Dataset (2000 Samples) 🚀 A high-quality instruction-style dataset designed for fine-tuning Large Language Models (LLMs) on Data Science concepts. This dataset contains ~2000 curated question-answer pairs in ChatML format, enabling models to learn how to explain, define, and discuss core data science topics in a clear and beginner-friendly way. 🎯 Objective The goal of this dataset is to: Train LLMs to act as a Data Science Tutor Provide clear… See the full description on the dataset page: https://huggingface.co/datasets/Hamzasajjad38/data-science-chatbot.texttext-generation1K<n<10K0 likes14 downloads5mo agoHugging Face20tojpaj /science-factcheck-indic Health-Science Fact-Check (Hindi/Punjabi) Native Hindi and Punjabi text from ai4bharat/IndicCorpV2, adapted with AutoScientist into substantive domain responses written by a health-science fact-checker judging claims and showing the reasoning. Rows 667 Unique source texts 667 Absolute quality score 9.0/10 (grade A) Source score before adaptation 9.0/10 (grade A) Percentile 33.0 Relative change +0.0% Median response length 520 chars… See the full description on the dataset page: https://huggingface.co/datasets/tojpaj/science-factcheck-indic.texttext-generationn<1K0 likes11 downloads2mo agoHugging Face21MIldoc /rus_science_for_gpt_oss_20b rus_science_for_gpt_oss_20b Русскоязычный датасет для дообучения LLM под научно-академический ассистент. Описание ~32 272 примера в формате JSONL. Тематика: научные тексты, академический стиль, описание таблиц/методик, введения, пояснения, переформулировки. Каждая строка содержит полный контекст диалога и готовые ответы ассистента. Формат полей reasoning_language: язык рассуждений ("Russian"). developer: инструкция для ассистента (роль/стиль/задача). user:… See the full description on the dataset page: https://huggingface.co/datasets/MIldoc/rus_science_for_gpt_oss_20b.texttext-generation10K<n<100K1 likes9 downloads9mo agoHugging Face22SAgarwal34 /grade3-science-explanations-v4r7 Grade-Level Science Explanations v4r7 The final 485-record supervised fine-tuning dataset for the grade-level science explainer. Each record maps a unique elementary-science question to a concise, mechanism-complete explanation. Training uses the minimal prompt Explain: {phrasing} so the reading behavior must be learned from examples rather than supplied through prompt instructions. Files File Records Purpose gold_v4_r7.jsonl 485 Final training split… See the full description on the dataset page: https://huggingface.co/datasets/SAgarwal34/grade3-science-explanations-v4r7.tabulartext-generationn<1K0 likes5 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.