CoolFace
28 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Hugodonotexit /math-code-science-deepseek-r1-en R1 Dataset Collection Aggregated high-quality English prompts and model-generated responses from DeepSeek R1 and DeepSeek R1-0528. Dataset Summary The R1 Dataset Collection combines multiple public DeepSeek-generated instruction-response corpora into a single, cleaned, English-only JSONL file. Each example consists of a <|user|> prompt and a <|assistant|> response in one "text" field. This release includes: ~21,000 examples from the DeepSeek-R1-0528 Distilled Custom… See the full description on the dataset page: https://huggingface.co/datasets/Hugodonotexit/math-code-science-deepseek-r1-en.textquestion-answering1M<n<10M5 likes1.1k downloads1y agoHugging Face02LLaMAX /BenchMAX_Science Dataset Sources Paper: BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models Link: https://huggingface.co/papers/2502.07346 Repository: https://github.com/CONE-MT/BenchMAX Dataset Description BenchMAX_Science is a dataset of BenchMAX, sourcing from GPQA, which evaluates the natural science reasoning capability in multilingual scenarios. We extend the original English dataset to 16 non-English languages. The data is first translated by Google… See the full description on the dataset page: https://huggingface.co/datasets/LLaMAX/BenchMAX_Science.textquestion-answering1K<n<10K2 likes957 downloads2y agoHugging Face03simpleG2023 /chinese-materials-science-open-intelligence 🔬 Chinese Materials Science & Metallurgy Open Intelligence Dataset Curated open intelligence dataset providing English research briefs, authoritative DOIs, executive summaries, and high-resolution micrographs of breakthrough Chinese scientific research in Materials Science, Metallurgy, Advanced Alloys, and Mining Engineering. [!IMPORTANT] Data Completeness & Research Authenticity Notice: Included in this Hugging Face Open Dataset: English structured abstracts, core… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-materials-science-open-intelligence.tabulartext-retrieval1K<n<10K0 likes391 downloads18h agoHugging Face04seonjeongh /science_reasoning science_reasoning Mistral-7B의 과학 지식·추론 능력 향상을 위해 6개 공개 과학 객관식 QA 데이터셋을 통일 포맷으로 변환하고, ARC-Challenge test와의 오염을 제거한 데이터셋입니다. 원본 데이터셋 allenai/sciq allenai/openbookqa (main) allenai/qasc allenai/quartz allenai/ai2_arc (ARC-Easy / ARC-Challenge) nguyen-brat/worldtree 전처리 포맷 통일: 각 데이터셋의 서로 다른 스키마를 unique_id, orig_id, source, question, choices, answer, support 필드로 변환. support는 근거 문단/문장으로, 데이터셋별 원본 필드(support/fact/para/cot)에서 구성하거나 없으면 빈 문자열.… See the full description on the dataset page: https://huggingface.co/datasets/seonjeongh/science_reasoning.textmultiple-choice10K<n<100K0 likes333 downloads2mo agoHugging Face05Royal-lobster /10001-Science-Facts 10,001 Science Facts 10,000+ obscure, surprising, and verifiable science facts The kind that make you go "wait, really?" 🔗 GitHub Repository • 📁 Download by Category 🤔 What is this? A curated dataset of 10,003 science facts across 32 categories — from quantum physics to parasites to the history of food. Every fact is: Sourced — from Wikipedia, Wikidata, academic sources Verifiable — no LLM hallucinations Surprising — passes the "dinner party test"… See the full description on the dataset page: https://huggingface.co/datasets/Royal-lobster/10001-Science-Facts.texttext-generation10K<n<100K1 likes238 downloads8mo agoHugging Face06stindardlogic /science-qa-sft-100k Science QA SFT (100K) 100,000 science Q&A examples with step-by-step explanations for SFT fine-tuning. Covers physics, chemistry, biology, astronomy, and earth science at beginner through advanced difficulty. Motivation Models trained on general text often give superficially plausible but mechanistically wrong answers to science questions — stating the right conclusion without understanding the underlying reasoning. This dataset trains models to explain why an… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/science-qa-sft-100k.texttext-generation100K<n<1M0 likes234 downloads2mo agoHugging Face07ianncity /GLM-5.2-Science GLM-5.2 · Science-50000x 50,000x traces distilled from GLM-5.2 on High reasoning Physics · Chemistry · Biology Token Count: 160M Theres prompt overlap with my Kimi K2.5 dataset science subset, which I think those prompts are getting used in alot of places now You can use this dataset for any purpose and you dont need to credit me, preferably dont claim it as your own. hi - ianncity texttext-generation10K<n<100K19 likes176 downloads2mo agoHugging Face08HacksHaven /science-on-a-sphere-prompt-completions Dataset Card for Science On a Sphere QA Dataset Dataset Details Dataset Description This dataset comprises question-and-answer (QA) pairs generated from NOAA's Science On a Sphere (SOS) website, including support documentation and the dataset catalog. Each entry contains a prompt and a corresponding completion, designed to support educational and research use cases in Earth science. This dataset includes a custom dataset_script.py and a consolidated file… See the full description on the dataset page: https://huggingface.co/datasets/HacksHaven/science-on-a-sphere-prompt-completions.textquestion-answering1K<n<10K0 likes131 downloads1y agoHugging Face09nhminh107 /VietEmbed-RAG-Science VietEmbed-RAG Science VietEmbed-RAG Science is a Vietnamese retrieval dataset containing 68,567 query-document examples across seven scientific and technical domains. Each record consists of: A Vietnamese query (anchor) A relevant passage (positive) A semantically related but non-answering passage (hard_negative) Topic and domain metadata The dataset is designed for training and domain adaptation of Vietnamese text embedding, semantic retrieval, and Retrieval-Augmented… See the full description on the dataset page: https://huggingface.co/datasets/nhminh107/VietEmbed-RAG-Science.textsentence-similarity10K<n<100K1 likes123 downloads18d agoHugging Face10KSU-HW-SEC /food-science-llm-protocol Food Science LLM Text-Mining Protocol Pipeline and derived data accompanying: Guo X, Fu W. Data Mining and Text Mining Using Large Language Models. In: Li Y, Zhang D, Guo Z (eds), AI in Food Science: Methods and Protocols. Methods and Protocols in Food Science. Springer. The chapter prints one protocol as 26 numbered steps with abbreviated code listings. This repository is the executable form of that protocol. Every step has a corresponding function here, and every number in… See the full description on the dataset page: https://huggingface.co/datasets/KSU-HW-SEC/food-science-llm-protocol.documenttoken-classificationn<1K0 likes84 downloads1mo agoHugging Face11percepteyeAI /10001-Science-Facts 10,001 Science Facts 10,000+ obscure, surprising, and verifiable science facts The kind that make you go "wait, really?" 🔗 GitHub Repository • 📁 Download by Category 🤔 What is this? A curated dataset of 10,003 science facts across 32 categories — from quantum physics to parasites to the history of food. Every fact is: Sourced — from Wikipedia, Wikidata, academic sources Verifiable — no LLM hallucinations Surprising — passes the "dinner party test"… See the full description on the dataset page: https://huggingface.co/datasets/percepteyeAI/10001-Science-Facts.texttext-generation10K<n<100K1 likes81 downloads7mo agoHugging Face12TheJackBright /verisci-verified-science-math-code VeriSci Verified Science Math Code Verifier-grounded dataset for the Adaption AutoScientist Challenge Part 2, targeting the Science category with secondary Math and Code coverage. Summary VeriSci trains models to solve scientific computations, finite-difference PDE updates, numerical ODE steps, unit-checked mechanics, thermodynamics, circuits, chemistry stoichiometry, molarity, unit conversion, vector decomposition, two-point linear modeling, small Python… See the full description on the dataset page: https://huggingface.co/datasets/TheJackBright/verisci-verified-science-math-code.texttext-generation1K<n<10K0 likes71 downloads2mo agoHugging Face13batuhanozkose /Rehber-CoT-Science 🧬 Rehber-CoT-Science: Turkish Scientific Reasoning Dataset Turkish Scientific Computational Reasoning (Chain-of-Thought) Dataset Multi-step scientific problem-solving dataset with verifiable Python code and detailed explanations Dataset • Author 📌 Changelog Eski sürümlere erişim: Branch menüsünden v1 seçebilirsiniz. Version Date Changes v2.0 24.12.2025 ✨ Yeni explained_answer alanı eklendi, Statistics domain eklendi, 712 örneğe genişletildi… See the full description on the dataset page: https://huggingface.co/datasets/batuhanozkose/Rehber-CoT-Science.textquestion-answering1K<n<10K4 likes63 downloads9mo agoHugging Face14agentlans /nvidia-Nemotron-Science-Math NVIDIA Nemotron Science and Math Reasoning This is an unofficial, curated collection derived from NVIDIA's open-source Nemotron datasets. It is designed specifically to train language models in complex scientific and mathematical reasoning by providing structured chain-of-thought (CoT) examples. To ensure efficiency, the shortest available CoT sequence was chosen for each question, filtering out redundant variations while preserving the core logical progression toward the final… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/nvidia-Nemotron-Science-Math.texttext-generation100K<n<1M0 likes59 downloads5mo agoHugging Face15JosephLee /science_textbook_elementary_kortextquestion-answering10K<n<100K3 likes52 downloads3y agoHugging Face16pthinc /BCE-Prettybird-Middle-Science-v0.1 BCE-Prettybird-Middle-Science-v0.1 - 101000 Science Q&A Dataset for Instruction-Based Learning We are excited to introduce a comprehensive math-physics-chemistry-biology dataset containing 100500 instruction-based question-answer pairs, designed to support research in science reasoning, problem-solving, and AI training. Generated using Python’s math libraries (e.g., math, numpy, sympy), the dataset covers a diverse range of difficulty levels—from basic arithmetic and algebra to… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Middle-Science-v0.1.texttext-classification100K<n<1M0 likes35 downloads5mo agoHugging Face17Puidii /aalen_university_faculty_computer_science Dataset Card This dataset contains question-answer pairs from all study programmes of the Faculty of Computer Science at the University of Aalen, Germany. The training dataset is automatically generated by ChatGPT. The validation dataset was manually created. It was collected to train an answer-Q&A chatbot based on LLM fine-tuning. All used scripts and examples can be found in the linked GitHub repository (https://github.com/pattplatt/llm_dataset_creation_and_finetuning).… See the full description on the dataset page: https://huggingface.co/datasets/Puidii/aalen_university_faculty_computer_science.textquestion-answering1K<n<10K0 likes34 downloads2y agoHugging Face18AethronPhantom /nexa-science-multitask-balanced Nexa Science Multitask Balanced This dataset is a curated, instruction-formatted scientific multitask mixture for: claim verification (<TASK:VERIFY>) abstract-grounded biomedical QA (<TASK:QA>) retrieval relevance re-ranking (<TASK:RERANK>) Format Each row is JSONL with: {task, instruction, input, output, meta} Splits Included train_balanced_short.jsonl val_balanced_short.jsonl stats_balanced_short.json Notes QA in this balanced release is… See the full description on the dataset page: https://huggingface.co/datasets/AethronPhantom/nexa-science-multitask-balanced.texttext-classification10K<n<100K0 likes34 downloads7mo agoHugging Face19Somtharu181coder /science_behavioral_and_domain_diversity_dataset Nepali Science SFT Dataset — Clean Candidate A high-quality Nepali Science Supervised Fine-Tuning (SFT) dataset containing short question–answer instruction-following examples written primarily in Nepali Devanagari script. This release is the clean candidate produced after structural validation, language checks, duplicate analysis, and Unicode-contamination filtering. Dataset Overview Property Value Dataset file clean_candidate.jsonl Records 29,320… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/science_behavioral_and_domain_diversity_dataset.texttext-generation10K<n<100K0 likes34 downloads1mo agoHugging Face20pthinc /BCE-Prettybird-Nano-Science-v0.1 BCE-Prettybird-Nano-Science-v0.1 - 500 Science Q&A Dataset for Instruction-Based Learning We are excited to introduce a comprehensive math-physics-chemistry-biology dataset containing 500 instruction-based question-answer pairs, designed to support research in science reasoning, problem-solving, and AI training. Generated using Python’s math libraries (e.g., math, numpy, sympy), the dataset covers a diverse range of difficulty levels—from basic arithmetic and algebra to advanced… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Science-v0.1.texttext-classificationn<1K0 likes29 downloads6mo agoHugging Face21issdandavis /scbe-life-science-research-training-demo Status: experimental. Experiment-specific slice. Primary public dataset: scbe-aethermoore-training-data. SCBE Research Training Package This package was generated from live pubmed pulls for the query protein structure prediction and is meant for lightweight Hugging Face dataset and SFT experiments. Files papers.jsonl: normalized raw research records sft_train.jsonl: train split for instruction-style tasks sft_validation.jsonl: validation split… See the full description on the dataset page: https://huggingface.co/datasets/issdandavis/scbe-life-science-research-training-demo.texttext-generationn<1K0 likes27 downloads2mo agoHugging Face22pur4nj41y /bro-science-benchmark BroScienceBench A literature-grounded benchmark for evaluating how large language models handle strength-training misinformation ("bro-science"). 246 items across 42 myth clusters and 7 categories, each pairing an evidence-based answer against a documented gym myth and a plausible distractor. Code, evaluation harness, analysis scripts, figures, and the full datasheet: https://github.com/puranjayh/bro-science-benchmark ⚠️ For engineering evaluation and research only. Not medical… See the full description on the dataset page: https://huggingface.co/datasets/pur4nj41y/bro-science-benchmark.textquestion-answeringn<1K0 likes25 downloads2mo agoHugging Face23ov1n /science-sinhala-gce-olevel-2023-mcq Dataset Details This dataset contains 40 Science MCQ questions and answers in Sinhala language of the GCE Ordinary Level Science paper 2023. textquestion-answeringn<1K0 likes19 downloads2y agoHugging Face24Marco711 /ScienceQA-Weather-R1 Introduction Dataset Summary This dataset is the out-of-domain (OOD) evaluation set used in "Weather-R1: Logically Consistent Reinforcement Fine-Tuning for Multimodal Reasoning in Meteorology". It is curated from the "Weather and climate" category of ScienceQA and contains 324 English, multimodal multiple-choice questions to test cross-domain generalization. Supported Tasks Multi-modal Multiple Choice Languages English Dataset Overview For… See the full description on the dataset page: https://huggingface.co/datasets/Marco711/ScienceQA-Weather-R1.imagemultiple-choicen<1K0 likes17 downloads8mo agoHugging Face25psdn-ai /science-qa-samplesgated Science Q&A Samples This sample shows structured science question-answer pairs for reviewing subject coverage, difficulty labeling, and answer format before scoping a larger educational dataset. What This Shows Q&A examples across science and math subjects Metadata for topic, difficulty, curriculum alignment, and question type A view of how text and asset-backed questions are represented Dataset Specifications Field Value Modality Text… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/science-qa-samples.textquestion-answeringn<1K0 likes15 downloads3mo agoHugging Face26JosephLee /science_textbook_elementary_kor_seedtextquestion-answering1K<n<10K3 likes12 downloads3y agoHugging Face27ov1n /sinhala-political-science-gce-alevel-2021-questionstextquestion-answeringn<1K0 likes11 downloads2y agoHugging Face28Hamzasajjad38 /data-science-chatbot 📊 Data Science Chatbot Dataset (2000 Samples) 🚀 A high-quality instruction-style dataset designed for fine-tuning Large Language Models (LLMs) on Data Science concepts. This dataset contains ~2000 curated question-answer pairs in ChatML format, enabling models to learn how to explain, define, and discuss core data science topics in a clear and beginner-friendly way. 🎯 Objective The goal of this dataset is to: Train LLMs to act as a Data Science Tutor Provide clear… See the full description on the dataset page: https://huggingface.co/datasets/Hamzasajjad38/data-science-chatbot.texttext-generation1K<n<10K0 likes11 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.