CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01TalentZHOU /hle_material_science HLE Material Science: A Specialized Benchmark for Materials Science A Materials Science Subset of Humanity's Last Exam (HLE) Overview HLE Material Science is a carefully curated materials science subset derived from the Humanity's Last Exam (HLE) dataset, containing 106 high-quality expert-level questions covering 25+ materials science subfields, with 97% of questions rated as high confidence. This dataset is designed to evaluate large language models'… See the full description on the dataset page: https://huggingface.co/datasets/TalentZHOU/hle_material_science.textquestion-answeringn<1K1 likes845 downloads8mo agoHugging Face02dvilasuero /natural-science-reasoning Natural Sciences Reasoning: the "smolest" reasoning dataset A smol-scale open dataset for reasoning tasks using Hugging Face Inference Endpoints. While intentionally limited in scale, this resource prioritizes: Reproducible pipeline for reasoning tasks using a variety of models (Deepseek V3, Deepsek-R1, Llama70B-Instruct, etc.) Knowledge sharing for domains other than Math and Code reasoning In this repo, you can find: The prompts and the pipeline (see the config file). The… See the full description on the dataset page: https://huggingface.co/datasets/dvilasuero/natural-science-reasoning.texttext-generationn<1K40 likes686 downloads2y agoHugging Face03laion /llama-nemotron-science-reasoning-on-canonical-think-full Llama-Nemotron science reasoning — Delphi canonical-think (COMPLETE, no length filter) The complete reasoning:on science split of nvidia/Llama-Nemotron-Post-Training-Dataset, converted once into the canonical Delphi chat-template thinking format. 708,920 rows. Unlike the cold-start warmup slice open-athena/llama-nemotron-science-reasoning-on-le3000tok-100k (and its -canonical-think variant), this build applies no length cap and no subsample — every long-CoT science example is… See the full description on the dataset page: https://huggingface.co/datasets/laion/llama-nemotron-science-reasoning-on-canonical-think-full.texttext-generation100K<n<1M0 likes381 downloads21d agoHugging Face04169Pi /Science-QnA Science-QnA The Science-QnA is a large-scale, high-quality science-focused dataset (~5.63M rows) curated using synthetic data generation through distillation techniques and select open-source resources. Designed to train and evaluate reasoning-capable models in science domains with emphasis on conceptual understanding, numerical problem-solving, and exam-style Q&A patterns across Physics, Chemistry, Biology, and Mathematics. Summary • Domain: Science, Physics… See the full description on the dataset page: https://huggingface.co/datasets/169Pi/Science-QnA.texttext-generation1M<n<10M3 likes317 downloads7mo agoHugging Face05islamlab /islamic-sciences islamlab — The Islamic Sciences Corpus The Islamic sciences other than Qur'an and hadith, as their authors wrote them: 4,022 works by scholars who died between the 0st and the 14th Hijri century, cut along their own chapter and biographical-entry boundaries into 1,864,389 units (3.41 billion characters of Arabic), each carrying the volume and page it sits on so a quotation can be cited rather than merely produced. Scope is Ahl al-Sunnah wa'l-Jamāʿah, and the gate is the author… See the full description on the dataset page: https://huggingface.co/datasets/islamlab/islamic-sciences.tabulartext-generation1M<n<10M2 likes308 downloads1mo agoHugging Face06marin-community /openthoughts4-science-26041-prompts-qwen3-30b-a3B-thinking-2507-n8-flattened-logprobs-k16 OpenThoughts-4 Science SDG: Qwen3-30B-A3B-Thinking-2507 (n=8, top-16 logprobs) Synthetic generations from Qwen/Qwen3-30B-A3B-Thinking-2507 on the Marin OpenThoughts-4 science SDG prompt set. Each prompt is sampled n=8 times, and for every generated token the dataset stores the chosen-token log probability plus the top-16 log probabilities over the vocabulary, enabling distillation, KL-style fine-tuning, reranking, and uncertainty analysis. Generation setup Field… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-science-26041-prompts-qwen3-30b-a3B-thinking-2507-n8-flattened-logprobs-k16.tabulartext-generation100K<n<1M0 likes285 downloads5mo agoHugging Face07stonelight /hle_material_science HLE Material Science: A Specialized Benchmark for Materials Science A Materials Science Subset of Humanity's Last Exam (HLE) Overview HLE Material Science is a carefully curated materials science subset derived from the Humanity's Last Exam (HLE) dataset, containing 106 high-quality expert-level questions covering 25+ materials science subfields, with 97% of questions rated as high confidence. This dataset is designed to evaluate large language models'… See the full description on the dataset page: https://huggingface.co/datasets/stonelight/hle_material_science.textquestion-answeringn<1K0 likes232 downloads3mo agoHugging Face08marin-community /openthoughts4-science-26041-prompts-qwen3-32b-n8-flattened-logprobs-k16 OpenThoughts-4 Science SDG: Qwen3-32B (n=8, top-16 logprobs) Synthetic generations from Qwen/Qwen3-32B on the Marin OpenThoughts-4 science SDG prompt set. Each prompt is sampled n=8 times, and for every generated token the dataset stores the chosen-token log probability plus the top-16 log probabilities over the vocabulary, enabling distillation, KL-style fine-tuning, reranking, and uncertainty analysis. Generation setup Field Value Generator model… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-science-26041-prompts-qwen3-32b-n8-flattened-logprobs-k16.tabulartext-generation100K<n<1M0 likes216 downloads5mo agoHugging Face09Lots-of-LoRAs /task047_miscellaneous_answering_science_questions Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task047_miscellaneous_answering_science_questions Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task047_miscellaneous_answering_science_questions.texttext-generationn<1K0 likes148 downloads2y agoHugging Face10tilikumotp /Global-Ocean-Science-Corpus 🌊 Global-Ocean-Science-Corpus (v2.0 Curated & Cleaned) A Highly Curated, Large-Scale Pre-Training & RAG Corpus for Deep Ocean Sciences, Marine Biology, and Oceanography Language Note: This dataset is a 100% English-language scientific corpus (language: "en") aggregating peer-reviewed literature, deep-sea exploration dossiers, and technical oceanographic reports from leading global marine institutes. Global-Ocean-Science-Corpus, derin okyanus bilimleri, deniz biyolojisi… See the full description on the dataset page: https://huggingface.co/datasets/tilikumotp/Global-Ocean-Science-Corpus.tabulartext-generation10K<n<100K1 likes116 downloads2d agoHugging Face11Lots-of-LoRAs /task701_mmmlu_answer_generation_high_school_computer_science Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task701_mmmlu_answer_generation_high_school_computer_science Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task701_mmmlu_answer_generation_high_school_computer_science.texttext-generationn<1K1 likes80 downloads2y agoHugging Face12Ik45 /data-science-en-id Data Science EN-ID Parallel Corpus (Scientific Domain) Dataset Description This dataset is a curated English-Indonesian (EN-ID) parallel corpus specifically designed for the Scientific and Data Science domains. It was developed to support the training of Machine Translation (NMT) models and Large Language Models (LLMs) to better handle technical terminology, academic structures, and formal scientific language. Primary Languages: English (EN) and Indonesian (ID) Domain:… See the full description on the dataset page: https://huggingface.co/datasets/Ik45/data-science-en-id.texttext-generation10M<n<100M0 likes72 downloads6mo agoHugging Face13ytu-ce-cosmos /tubitak-science-olympiad-tr TUBITAK Science Olympiad Dataset This dataset contains multiple-choice and open-ended scientific questions sourced from the TUBITAK (The Scientific and Technological Research Council of Turkey) Science Olympiads spanning various years. It is intended to serve as a benchmark for evaluating the advanced analytical, mathematical, and computational reasoning capabilities of Large Language Models (LLMs) in the Turkish language. The dataset comprises approximately 2700 problems across… See the full description on the dataset page: https://huggingface.co/datasets/ytu-ce-cosmos/tubitak-science-olympiad-tr.imagequestion-answering1K<n<10K13 likes70 downloads6mo agoHugging Face14islamlab /islamic-sciences-training islamlab — Islamic Sciences Training Sets Training data derived from islamlab/islamic-sciences: text for domain adaptation, retrieval pairs with hard negatives, and citation questions whose answers are read out of the corpus rather than written by a model. Nothing here is generated. Questions come from a fixed set of templates and every answer is a field already present in the corpus. That buys a narrow dataset in exchange for one that cannot teach a model a fact the sources do… See the full description on the dataset page: https://huggingface.co/datasets/islamlab/islamic-sciences-training.texttext-generation1M<n<10M0 likes68 downloads1mo agoHugging Face15HSiTori /scienceQA Filter: no image && hint != '' texttext-generation1K<n<10K0 likes63 downloads3y agoHugging Face16graf /bvg_science_qwen_4b_not_easy BVG Qwen3-4B science not-easy prompts This is the exact two-split dataset artifact used by the BVG Qwen3-4B science experiments. It is published as a DatasetDict with train and validation splits. Split Rows SHA-256 of canonical JSON rows train 11,121 e985f63334809b43e8cffa971829bff6c2159fe08d79630b2fbbda9d22bc0831 validation 997 c2ea31a3e676b2c28c72daaddb9023c9163cd6671c7cf3afd2e305f7fc206480 The active experiment TOMLs consume the complete train split. Their… See the full description on the dataset page: https://huggingface.co/datasets/graf/bvg_science_qwen_4b_not_easy.texttext-generation10K<n<100K0 likes60 downloads16d agoHugging Face17marin-community /openthoughts4-science-26041-prompts-qwen3-4b-n8-flattened-logprobs-k16 OpenThoughts-4 Science SDG: Qwen3-4B (n=8, top-16 logprobs) Synthetic generations from Qwen/Qwen3-4B on the Marin OpenThoughts-4 science SDG prompt set. Each prompt is sampled n=8 times, and for every generated token the dataset stores the chosen-token log probability plus the top-16 log probabilities over the vocabulary, enabling distillation, KL-style fine-tuning, reranking, and uncertainty analysis. Generation setup Field Value Generator model… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-science-26041-prompts-qwen3-4b-n8-flattened-logprobs-k16.tabulartext-generation100K<n<1M0 likes57 downloads5mo agoHugging Face18expertdata-factory /science-cot-datasetgated ExpertData Science — Scientific Reasoning Expert-Annotated · Rights-Cleared · Ground-Truth Verified · PII-Clean Each record captures a complete experimental or theoretical reasoning chain: Hypothesis → Methodology → Causal Chain → Validated Conclusion. Extracted from peer-reviewed papers across physics, biology, materials science, astrophysics, and neuroscience using structured scientific-reasoning extraction. This dataset is produced by the ExpertData-Factory pipeline (Mine →… See the full description on the dataset page: https://huggingface.co/datasets/expertdata-factory/science-cot-dataset.textquestion-answeringn<1K0 likes54 downloads7mo agoHugging Face19omniomni /omni-science GitHub   Website   Paper (Coming Soon) Dataset Details This dataset is a combination of corpora of text from scientific Wikipedia articles and scientific papers across major fields of science. This dataset contains continued-pretrain data. Sources This dataset was sourced from the following open-sourced datasets: Science zeroshot/arxiv-biology legacy-datasets/wikipedia bisectgroup/PubMed_TA bluuebunny/biorxiv_abstract_embedding_mxbai_large_v1_milvus… See the full description on the dataset page: https://huggingface.co/datasets/omniomni/omni-science.texttext-generation100K<n<1M1 likes50 downloads1y agoHugging Face20robworks-software /k12-science-standards [!WARNING] Deprecated - use k12-science-standards-expanded instead. This dataset is superseded: every instruction in this set also appears there, plus 1,123 more and nine additional metadata columns. Nothing here is unique to it. It stays online so existing references keep resolving, but it will not be updated. New work should point at robworks-software/k12-science-standards-expanded. K-12 Science Standards (generated instruction data) 6,787 instruction/input/output records… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-science-standards.texttext-classification1K<n<10K0 likes49 downloads2mo agoHugging Face21Laz4rz /wikipedia_science_chunked_small_rag_512 ScienceWikiSmallChunk Processed version of millawell/wikipedia_field_of_science, prepared to be used in small context length RAG systems. Chunk length is tokenizer dependent, but each chunk should be around 512 tokens. Longer wikipedia pages have been split into smaller entries, with title added as a prefix. There is also 256 tokens dataset available: Laz4rz/wikipedia_science_chunked_small_rag_256 If you wish to prepare some other chunk length: use… See the full description on the dataset page: https://huggingface.co/datasets/Laz4rz/wikipedia_science_chunked_small_rag_512.texttext-generation1M<n<10M4 likes40 downloads2y agoHugging Face22AgentsSci /IEEE2026_BigData_MAS-4-Science-Matching SciAgentTrace An execution-layer trace resource for scientific-agent workload characterization. A protocol fixes who reasons, what each role can see, when feedback returns, and when a workflow stops. Those choices determine the sequence of model requests that produces an answer, so protocol design is also workload design. Two workflows that consume similar token totals can issue very different request sequences. SciAgentTrace records that difference. The matched core runs the… See the full description on the dataset page: https://huggingface.co/datasets/AgentsSci/IEEE2026_BigData_MAS-4-Science-Matching.tabulartext-generation1M<n<10M0 likes40 downloads2d agoHugging Face23omniomni /omni-science-instruct GitHub   Website   Paper (Coming Soon) Dataset Details This dataset is a subset of a general STEM dataset, containing only science-related question-answer pairs. This dataset contains only chat-based data. Sources This dataset was sourced from the following open-sourced dataset: Science TIGER-Lab/WebInstructSub texttext-generation100K<n<1M1 likes34 downloads1y agoHugging Face24Ik45 /wikipedia_dataset_science_en_id Wikipedia Dataset Science (English - Indonesian) Dataset Description This dataset contains 122,433 aligned sentence pairs extracted from Wikipedia science articles in English and Indonesian. It is highly suitable for Natural Language Processing (NLP) tasks such as machine translation, cross-lingual alignment, and fine-tuning Large Language Models (LLMs) to better understand scientific terminology in Indonesian. Language(s): English (en) and Indonesian (id) Domain:… See the full description on the dataset page: https://huggingface.co/datasets/Ik45/wikipedia_dataset_science_en_id.texttranslation100K<n<1M0 likes34 downloads6mo agoHugging Face25HerrHruby /synthetic-science-v2-sample Synthetic Scientific Research Threads — v2 (sample) A synthetic continual-learning benchmark: each episode is a coherent sequence of short fictional scientific research documents about a single made-up entity, with per-document QA anchors. Later documents build on, revise, or supersede earlier ones. Designed to stress test-time / meta-learning approaches where a model must adapt to a stream of documents and answer questions grounded in what it has just seen. This is a sample… See the full description on the dataset page: https://huggingface.co/datasets/HerrHruby/synthetic-science-v2-sample.tabularquestion-answeringn<1K0 likes32 downloads3mo agoHugging Face26Laz4rz /wikipedia_science_chunked_small_rag_256 ScienceWikiSmallChunk Processed version of millawell/wikipedia_field_of_science, prepared to be used in small context length RAG systems. Chunk length is tokenizer dependent, but each chunk should be around 256 tokens. Longer wikipedia pages have been split into smaller entries, with title added as a prefix. There is also 512 tokens dataset available: Laz4rz/wikipedia_science_chunked_small_rag_512 If you wish to prepare some other chunk length: use… See the full description on the dataset page: https://huggingface.co/datasets/Laz4rz/wikipedia_science_chunked_small_rag_256.texttext-generation1M<n<10M3 likes29 downloads2y agoHugging Face27mkurman /synthlabs-GLM-5.2-Science GLM-5.2 Science Synth Reasoning Synthetic reasoning traces for science questions from the GLM-5.2 science dataset. Each record contains a complex scientific question with SYNTH-style reasoning and a generated answer. Dataset Summary 33,014 records (605 dupes + 726 incomplete/truncated removed from 34,345 source) 33,014 reasoning turns (99.9% format compliance) Average 3,094 chars per reasoning trace Models Used Model Records… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/synthlabs-GLM-5.2-Science.tabulartext-generation10K<n<100K0 likes29 downloads2mo agoHugging Face28RecursiveMAS /Mixture-Science RecursiveMAS Mixture-Science Project Page | Code | Paper We introduce RecursiveMAS, a multi-agent framework that scales agent collaboration through latent-space recursion. This dataset contains training examples for the Mixture-Style setting. Dataset Details Item Description Dataset RecursiveMAS/Mixture-Science Original file Mixture-Science.json Collaboration style Mixture-Style Used for science specialist inner agent training Split train Rows… See the full description on the dataset page: https://huggingface.co/datasets/RecursiveMAS/Mixture-Science.texttext-generation1K<n<10K0 likes27 downloads3mo agoHugging Face29caiyuchen /OPV-Science OPV Science: original experiment data Prepared for Learning to Steer, Steering to See. Original four-domain split (physics, chemistry, biology, material). Original system/user prompts and answer-tag instructions are preserved. Split Rows train 3,901 validation 434 Provenance and terms The data is derived from the following sources; their original terms and attribution obligations continue to apply. No new blanket license is asserted over the… See the full description on the dataset page: https://huggingface.co/datasets/caiyuchen/OPV-Science.texttext-generation1K<n<10K0 likes26 downloads3d agoHugging Face30Lots-of-LoRAs /task688_mmmlu_answer_generation_college_computer_science Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task688_mmmlu_answer_generation_college_computer_science Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task688_mmmlu_answer_generation_college_computer_science.texttext-generationn<1K1 likes25 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.