CoolFace
23 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01scientifi-papers /scientific-papers Scientific Papers - Raw Full Text ~57 million scientific papers with full text, extracted from multiple large-scale academic paper collections. This dataset provides raw full text suitable for pre-training, fine-tuning, or building search indices over scientific literature. Subsets Subset Papers Size Source papers-2 ~18.5M ~358 GB S2ORC papers collection (untitled subset) papers-3 ~27.4M ~198 GB S2ORC scientific-papers collection pes2o ~8.2M ~106 GB… See the full description on the dataset page: https://huggingface.co/datasets/scientifi-papers/scientific-papers.texttext-generation10M<n<100M2 likes5.2k downloads4mo agoHugging Face02AethronPhantom /Scientific_Research_Tokenized NexaSci Scientific Research Tokenized This dataset repository now holds the active NexaSci scientific pretraining reservoir, the NexaMat controller fine-tuning pack, and archived legacy reservoir builds. The current production reservoir is the 10B-token Apache Arrow release under nexasci_reservoir_v3_10b_prod_rust/. Current Status The active large-scale training artifact is: nexasci_reservoir_v3_10b_prod_rust/ It was produced from the NexaSci 10B data-engineering campaign… See the full description on the dataset page: https://huggingface.co/datasets/AethronPhantom/Scientific_Research_Tokenized.texttext-generation100K<n<1M7 likes1.2k downloads4mo agoHugging Face03onkanat /scientific-systems-sft-dpo-distilled Scientific and Systems Distillation Golden Dataset (SFT & DPO) Bu veri seti, temel ve uygulamalı bilimler ile yüksek başarımlı hesaplama (HPC) ve Linux sistem mühendisliği alanlarında yapay zeka modellerini ince ayar (fine-tuning) ve tercih hizalama (preference alignment) süreçlerine tabi tutmak amacıyla tasarlanmış, 2.551 adet ileri düzey teknik prompt ve bunlara karşılık gelen yüksek kaliteli model çıktılarından derlenmiş zengin bir sentetik veri kümesidir. 🚀 Veri… See the full description on the dataset page: https://huggingface.co/datasets/onkanat/scientific-systems-sft-dpo-distilled.texttext-generation10K<n<100K0 likes674 downloads2mo agoHugging Face04AxiomSetLabs /stem-scientific-code-sample AxiomSet Labs STEM Scientific-Code Sample A 30-task sample of STEM reasoning and scientific-code problems across five domains. Domains Biology: 6 tasks Chemistry: 6 tasks Materials Science: 6 tasks Mathematics: 6 tasks Physics: 6 tasks Each task contains two subproblems and one main problem, including prompts, scientific background, testing templates, and reference solutions. Files data/sample.jsonl — one task per line; used by the Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/AxiomSetLabs/stem-scientific-code-sample.texttext-generationn<1K0 likes224 downloads2mo agoHugging Face05P0u4a /Scientific-Reasoning Spark-234K: Skeleton-Guided Scientific Reasoning from Large-Scale Literature 🎉 Accepted to EMNLP 2026 Findings! Spark-234K is a scientific reasoning dataset containing 234K question-answer pairs synthesized from frontier scientific literature. Instead of directly generating QA pairs from full papers, SPARK first distills each paper into a compact reasoning skeleton—preserving its central claim, supporting evidence, quantitative relations, assumptions, and boundary… See the full description on the dataset page: https://huggingface.co/datasets/P0u4a/Scientific-Reasoning.texttext-generation100K<n<1M0 likes122 downloads9d agoHugging Face06Malikeh1375 /basic_mathematical-scientific-notation-parallel Mathematical and Scientific Notation Parallel Corpus Dataset Description This dataset is designed for tokenizer robustness testing in mathematical and scientific contexts. It contains identical mathematical content expressed in four different notation styles, allowing researchers to isolate tokenization effects from semantic differences when evaluating language models. Purpose Tokenizer Comparison: Compare how different tokenizers (BPE, SentencePiece… See the full description on the dataset page: https://huggingface.co/datasets/Malikeh1375/basic_mathematical-scientific-notation-parallel.texttext-generationn<1K0 likes75 downloads1y agoHugging Face07yousefg /scientific_studiestexttext-generation1M<n<10M4 likes64 downloads3y agoHugging Face08Malikeh1375 /complex_mathematical-scientific-notation-parallel Mathematical and Scientific Notation Parallel Corpus Dataset Description This dataset is designed for tokenizer robustness testing in mathematical and scientific contexts. It contains identical mathematical content expressed in four different notation styles, allowing researchers to isolate tokenization effects from semantic differences when evaluating language models. Purpose Tokenizer Comparison: Compare how different tokenizers (BPE, SentencePiece… See the full description on the dataset page: https://huggingface.co/datasets/Malikeh1375/complex_mathematical-scientific-notation-parallel.texttext-generationn<1K1 likes56 downloads1y agoHugging Face09JoaoCoelho /scientific_papers_citation_scores Dataset Summary This dataset comprises an array of scientific papers, each paper is associated with a series of scores. These scores quantify the number of citations each paper has received. The data regarding the papers and their citations were sourced from OpenCitations, a comprehensive and accessible online database of scholarly citations (available at https://opencitations.net/). How are these scores calculated? Imagine a tree where papers are nodes and citations… See the full description on the dataset page: https://huggingface.co/datasets/JoaoCoelho/scientific_papers_citation_scores.tabulartext-generation100K<n<1M2 likes44 downloads3y agoHugging Face10abhinavdread /CoT-Scientific-RAG-Reasoning CoT-Scientific-RAG-Reasoning This dataset is designed for fine-tuning Large Language Models (specifically Qwen-series) to perform complex reasoning over scientific and technical documents using Chain-of-Thought (CoT). Dataset Description The dataset contains instructions and scientific contexts (Medical Imaging, Autonomous Driving, VLA Frameworks) where the model is required to generate a reasoning trace before providing the final answer. Format: JSONL Logic: All outputs… See the full description on the dataset page: https://huggingface.co/datasets/abhinavdread/CoT-Scientific-RAG-Reasoning.texttext-generation1K<n<10K0 likes41 downloads9mo agoHugging Face11Agnuxo /scientific-papers-dataset Scientific Papers Dataset Scientific papers, whitepapers and documentation. Part of the Agnuxo Ecosystem by Francisco Angulo de Lafuente. texttext-generationn<1K0 likes38 downloads5mo agoHugging Face12alrobles /ecocoder-scientific-reasoning ecocoder-scientific-reasoning Chain-of-Thought (CoT) traces for fine-tuning LLMs on ecological scientific reasoning + code generation. Each trace follows: [CONTEXT] (ecological problem) → [REASONING] (step-by-step scientific thinking) → [CODE] (runnable R/Python implementation). Dataset Summary Split Traces Train 1,268 Val 159 Test 159 Total 1,586 73 unique ecological methods across 18 categories Languages: ~60% R, ~40% Python… See the full description on the dataset page: https://huggingface.co/datasets/alrobles/ecocoder-scientific-reasoning.texttext-generation1K<n<10K0 likes36 downloads4mo agoHugging Face13ExponentialScience /DLT-Scientific-Literature DLT-Scientific-Literature Paper | GitHub Dataset Description Dataset Summary DLT-Scientific-Literature is a specialized corpus of academic publications focused on Distributed Ledger Technology (DLT). This dataset is part of the larger DLT-Corpus collection, designed to support NLP research, language model development, and innovation studies in the DLT domain. The dataset contains 37,440 scientific documents with 564 million tokens, spanning publications from… See the full description on the dataset page: https://huggingface.co/datasets/ExponentialScience/DLT-Scientific-Literature.tabulartext-generation10K<n<100K0 likes31 downloads4mo agoHugging Face14abhi26 /openpipe-dpo-scientific-reasoning Openpipe Dpo Scientific Reasoning This dataset contains 100 high-quality examples for Direct Preference Optimization (DPO) training, formatted for OpenPipe fine-tuning, focused on scientific reasoning and analysis. Dataset Description This dataset was generated using an enhanced DSPy-based pipeline that creates structured reasoning traces for scientific questions. Each example follows the OpenAI chat completion format required by OpenPipe: OpenAI Chat Format: Standard… See the full description on the dataset page: https://huggingface.co/datasets/abhi26/openpipe-dpo-scientific-reasoning.texttext-generationn<1K0 likes30 downloads1y agoHugging Face15leaschuessler /scientific_question-generationData originated from: https://openstax.org.” https://openstax.org/details/books/chemistry-2e textsummarization1K<n<10K0 likes27 downloads1y agoHugging Face16ArabicNLPWorld /arabic-russian-scientific-translationsgated Arabic–Russian Scientific Translation Corpus Description This dataset provides parallel translations of scientific and medical texts from Arabic (original) and English (source) into Russian, generated by two state‑of‑the‑art language models: Gemma 3:4B (Google) LLaMA 3.1:8B (Meta) The corpus is built from four established Arabic–English corpora (see Sources below) and is intended for machine translation, model evaluation, and linguistic research.… See the full description on the dataset page: https://huggingface.co/datasets/ArabicNLPWorld/arabic-russian-scientific-translations.texttranslation10K<n<100K0 likes26 downloads3mo agoHugging Face17Miladsaeedi70 /scientific-multitask-instructions Scientific Multitask Instructions A multi-task scientific instruction-following dataset created for supervised fine-tuning and preference-optimization experiments. Dataset summary The dataset contains 1,576 conversational scientific examples across eight task types. Split Examples Train 1,260 Validation 158 Test 158 Total 1,576 Task distribution Task Examples Scientific question answering 256 Summarization 220… See the full description on the dataset page: https://huggingface.co/datasets/Miladsaeedi70/scientific-multitask-instructions.texttext-generation1K<n<10K0 likes24 downloads2mo agoHugging Face18abhi26 /dpo-scientific-reasoning Dpo Scientific Reasoning This dataset contains 100 high-quality examples for Direct Preference Optimization (DPO) training, focused on scientific reasoning and analysis. Dataset Description This dataset was generated using an enhanced DSPy-based pipeline that creates structured reasoning traces for scientific questions. Each example includes: Separated content fields: System prompt, user question, and full context as individual columns Chosen responses: High-quality… See the full description on the dataset page: https://huggingface.co/datasets/abhi26/dpo-scientific-reasoning.texttext-generationn<1K0 likes23 downloads1y agoHugging Face19Shekswess /scientific-research Description Topic: Scientific Research Papers Domains: Biology, Physics, Chemistry Number of Entries: 1,000 Dataset Type: Raw Dataset Model Used: Meta Llama4 Maverick 17B Instruct Language: English texttext-generation1K<n<10K4 likes16 downloads1y agoHugging Face20AgentsSci /scientific-agent-protocol-traces SciAgentTrace Matched cross-domain dataset of scientific-agent protocols. The central comparison contains the same 6,653 problems under two actor models and four protocols: 53,224 trajectories in 40 complete model--benchmark--protocol groups. The broader table-first package contains 68,892 trajectories. Begin with trajectories, outcomes, or matched_outcomes, then follow stable identifiers to messages and compressed raw traces. Repository:… See the full description on the dataset page: https://huggingface.co/datasets/AgentsSci/scientific-agent-protocol-traces.tabulartext-generation1M<n<10M0 likes16 downloads2mo agoHugging Face21Agnuxo /scientific-reasoning-dataset Scientific Reasoning Dataset Synthetic reasoning traces based on Agnuxo research. Part of the Agnuxo Ecosystem by Francisco Angulo de Lafuente. texttext-generationn<1K0 likes15 downloads5mo agoHugging Face22abhi26 /openpipe-chat-complete-scientific-reasoning Openpipe Chat Complete Scientific Reasoning This dataset contains 100 high-quality examples for chat completion fine-tuning, formatted for OpenPipe, focused on scientific reasoning and analysis. Dataset Description This dataset was generated using an enhanced DSPy-based pipeline that creates structured reasoning traces for scientific questions. Each example follows the OpenAI chat completion format required by OpenPipe: OpenAI Chat Format: Standard messages array with… See the full description on the dataset page: https://huggingface.co/datasets/abhi26/openpipe-chat-complete-scientific-reasoning.texttext-generationn<1K0 likes8 downloads1y agoHugging Face23rausch /scientific_corpus_cleanedgated Dataset Card for Scientific Corpus (Cleaned) This corpus contains ≈11 M English scientific documents cleaned via the DataTrove pipeline. It was used to continue pretraining T5-base (EN‑T5-Sci) before sliding-window materialization. Each document is provided as a row in one of 75 Parquet shards together with extensive per-document QA metadata. Dataset Details Uses Direct Use Continued pretraining / domain adaptation of encoder-decoder LMs on… See the full description on the dataset page: https://huggingface.co/datasets/rausch/scientific_corpus_cleaned.tabulartext-generation1K<n<10K2 likes2 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.