datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
scientific-papers
Scientific Papers - Raw Full Text
~57 million scientific papers with full text, extracted from multiple large-scale academic paper collections. This dataset provides raw full text suitable for pre-training, fine-tuning, or building search indices over scientific literature.
Subsets
Subset
Papers
Size
Source
papers-2
~18.5M
~358 GB
S2ORC papers collection (untitled subset)
papers-3
~27.4M
~198 GB
S2ORC scientific-papers collection
pes2o
~8.2M
~106 GB… See the full description on the dataset page: https://huggingface.co/datasets/scientifi-papers/scientific-papers.Scientific_Research_Tokenized
NexaSci Scientific Research Tokenized
This dataset repository now holds the active NexaSci scientific pretraining reservoir, the NexaMat controller fine-tuning pack, and archived legacy reservoir builds. The current production reservoir is the 10B-token Apache Arrow release under nexasci_reservoir_v3_10b_prod_rust/.
Current Status
The active large-scale training artifact is:
nexasci_reservoir_v3_10b_prod_rust/
It was produced from the NexaSci 10B data-engineering campaign… See the full description on the dataset page: https://huggingface.co/datasets/AethronPhantom/Scientific_Research_Tokenized.scientific-systems-sft-dpo-distilled
Scientific and Systems Distillation Golden Dataset (SFT & DPO)
Bu veri seti, temel ve uygulamalı bilimler ile yüksek başarımlı hesaplama (HPC) ve Linux sistem mühendisliği alanlarında yapay zeka modellerini ince ayar (fine-tuning) ve tercih hizalama (preference alignment) süreçlerine tabi tutmak amacıyla tasarlanmış, 2.551 adet ileri düzey teknik prompt ve bunlara karşılık gelen yüksek kaliteli model çıktılarından derlenmiş zengin bir sentetik veri kümesidir.
🚀 Veri… See the full description on the dataset page: https://huggingface.co/datasets/onkanat/scientific-systems-sft-dpo-distilled.stem-scientific-code-sample
AxiomSet Labs STEM Scientific-Code Sample
A 30-task sample of STEM reasoning and scientific-code problems across five domains.
Domains
Biology: 6 tasks
Chemistry: 6 tasks
Materials Science: 6 tasks
Mathematics: 6 tasks
Physics: 6 tasks
Each task contains two subproblems and one main problem, including prompts, scientific background, testing templates, and reference solutions.
Files
data/sample.jsonl — one task per line; used by the Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/AxiomSetLabs/stem-scientific-code-sample.Scientific-Reasoning
Spark-234K: Skeleton-Guided Scientific Reasoning from Large-Scale Literature
🎉 Accepted to EMNLP 2026 Findings!
Spark-234K is a scientific reasoning dataset containing 234K question-answer pairs synthesized from frontier scientific literature. Instead of directly generating QA pairs from full papers, SPARK first distills each paper into a compact reasoning skeleton—preserving its central claim, supporting evidence, quantitative relations, assumptions, and boundary… See the full description on the dataset page: https://huggingface.co/datasets/P0u4a/Scientific-Reasoning.basic_mathematical-scientific-notation-parallel
Mathematical and Scientific Notation Parallel Corpus
Dataset Description
This dataset is designed for tokenizer robustness testing in mathematical and scientific contexts. It contains identical mathematical content expressed in four different notation styles, allowing researchers to isolate tokenization effects from semantic differences when evaluating language models.
Purpose
Tokenizer Comparison: Compare how different tokenizers (BPE, SentencePiece… See the full description on the dataset page: https://huggingface.co/datasets/Malikeh1375/basic_mathematical-scientific-notation-parallel.scientific_studiescomplex_mathematical-scientific-notation-parallel
Mathematical and Scientific Notation Parallel Corpus
Dataset Description
This dataset is designed for tokenizer robustness testing in mathematical and scientific contexts. It contains identical mathematical content expressed in four different notation styles, allowing researchers to isolate tokenization effects from semantic differences when evaluating language models.
Purpose
Tokenizer Comparison: Compare how different tokenizers (BPE, SentencePiece… See the full description on the dataset page: https://huggingface.co/datasets/Malikeh1375/complex_mathematical-scientific-notation-parallel.scientific_papers_citation_scores
Dataset Summary
This dataset comprises an array of scientific papers, each paper is associated with a series of scores.
These scores quantify the number of citations each paper has received.
The data regarding the papers and their citations were sourced from OpenCitations, a comprehensive and accessible online database of scholarly citations (available at https://opencitations.net/).
How are these scores calculated?
Imagine a tree where papers are nodes and citations… See the full description on the dataset page: https://huggingface.co/datasets/JoaoCoelho/scientific_papers_citation_scores.CoT-Scientific-RAG-Reasoning
CoT-Scientific-RAG-Reasoning
This dataset is designed for fine-tuning Large Language Models (specifically Qwen-series) to perform complex reasoning over scientific and technical documents using Chain-of-Thought (CoT).
Dataset Description
The dataset contains instructions and scientific contexts (Medical Imaging, Autonomous Driving, VLA Frameworks) where the model is required to generate a reasoning trace before providing the final answer.
Format: JSONL
Logic: All outputs… See the full description on the dataset page: https://huggingface.co/datasets/abhinavdread/CoT-Scientific-RAG-Reasoning.scientific-papers-dataset
Scientific Papers Dataset
Scientific papers, whitepapers and documentation.
Part of the Agnuxo Ecosystem by Francisco Angulo de Lafuente.
ecocoder-scientific-reasoning
ecocoder-scientific-reasoning
Chain-of-Thought (CoT) traces for fine-tuning LLMs on ecological scientific reasoning + code generation.
Each trace follows: [CONTEXT] (ecological problem) → [REASONING] (step-by-step scientific thinking) → [CODE] (runnable R/Python implementation).
Dataset Summary
Split
Traces
Train
1,268
Val
159
Test
159
Total
1,586
73 unique ecological methods across 18 categories
Languages: ~60% R, ~40% Python… See the full description on the dataset page: https://huggingface.co/datasets/alrobles/ecocoder-scientific-reasoning.DLT-Scientific-Literature
DLT-Scientific-Literature
Paper | GitHub
Dataset Description
Dataset Summary
DLT-Scientific-Literature is a specialized corpus of academic publications focused on Distributed Ledger Technology (DLT). This dataset is part of the larger DLT-Corpus collection, designed to support NLP research, language model development, and innovation studies in the DLT domain.
The dataset contains 37,440 scientific documents with 564 million tokens, spanning publications from… See the full description on the dataset page: https://huggingface.co/datasets/ExponentialScience/DLT-Scientific-Literature.openpipe-dpo-scientific-reasoning
Openpipe Dpo Scientific Reasoning
This dataset contains 100 high-quality examples for Direct Preference Optimization (DPO) training, formatted for OpenPipe fine-tuning, focused on scientific reasoning and analysis.
Dataset Description
This dataset was generated using an enhanced DSPy-based pipeline that creates structured reasoning traces for scientific questions. Each example follows the OpenAI chat completion format required by OpenPipe:
OpenAI Chat Format: Standard… See the full description on the dataset page: https://huggingface.co/datasets/abhi26/openpipe-dpo-scientific-reasoning.scientific_question-generationData originated from:
https://openstax.org.”
https://openstax.org/details/books/chemistry-2e
arabic-russian-scientific-translations
Arabic–Russian Scientific Translation Corpus
Description
This dataset provides parallel translations of scientific and medical texts from Arabic (original) and English (source) into Russian, generated by two state‑of‑the‑art language models:
Gemma 3:4B (Google)
LLaMA 3.1:8B (Meta)
The corpus is built from four established Arabic–English corpora (see Sources below) and is intended for machine translation, model evaluation, and linguistic research.… See the full description on the dataset page: https://huggingface.co/datasets/ArabicNLPWorld/arabic-russian-scientific-translations.scientific-multitask-instructions
Scientific Multitask Instructions
A multi-task scientific instruction-following dataset created for
supervised fine-tuning and preference-optimization experiments.
Dataset summary
The dataset contains 1,576 conversational scientific examples across
eight task types.
Split
Examples
Train
1,260
Validation
158
Test
158
Total
1,576
Task distribution
Task
Examples
Scientific question answering
256
Summarization
220… See the full description on the dataset page: https://huggingface.co/datasets/Miladsaeedi70/scientific-multitask-instructions.dpo-scientific-reasoning
Dpo Scientific Reasoning
This dataset contains 100 high-quality examples for Direct Preference Optimization (DPO) training, focused on scientific reasoning and analysis.
Dataset Description
This dataset was generated using an enhanced DSPy-based pipeline that creates structured reasoning traces for scientific questions. Each example includes:
Separated content fields: System prompt, user question, and full context as individual columns
Chosen responses: High-quality… See the full description on the dataset page: https://huggingface.co/datasets/abhi26/dpo-scientific-reasoning.scientific-research
Description
Topic: Scientific Research Papers
Domains: Biology, Physics, Chemistry
Number of Entries: 1,000
Dataset Type: Raw Dataset
Model Used: Meta Llama4 Maverick 17B Instruct
Language: English
scientific-agent-protocol-traces
SciAgentTrace
Matched cross-domain dataset of scientific-agent protocols. The central
comparison contains the same 6,653 problems under two actor models and four
protocols: 53,224 trajectories in 40 complete model--benchmark--protocol
groups. The broader table-first package contains 68,892
trajectories. Begin with trajectories, outcomes, or matched_outcomes, then
follow stable identifiers to messages and compressed raw traces.
Repository:… See the full description on the dataset page: https://huggingface.co/datasets/AgentsSci/scientific-agent-protocol-traces.scientific-reasoning-dataset
Scientific Reasoning Dataset
Synthetic reasoning traces based on Agnuxo research.
Part of the Agnuxo Ecosystem by Francisco Angulo de Lafuente.
openpipe-chat-complete-scientific-reasoning
Openpipe Chat Complete Scientific Reasoning
This dataset contains 100 high-quality examples for chat completion fine-tuning, formatted for OpenPipe, focused on scientific reasoning and analysis.
Dataset Description
This dataset was generated using an enhanced DSPy-based pipeline that creates structured reasoning traces for scientific questions. Each example follows the OpenAI chat completion format required by OpenPipe:
OpenAI Chat Format: Standard messages array with… See the full description on the dataset page: https://huggingface.co/datasets/abhi26/openpipe-chat-complete-scientific-reasoning.scientific_corpus_cleaned
Dataset Card for Scientific Corpus (Cleaned)
This corpus contains ≈11 M English scientific documents cleaned via the DataTrove pipeline. It was used to continue pretraining T5-base (EN‑T5-Sci) before sliding-window materialization. Each document is provided as a row in one of 75 Parquet shards together with extensive per-document QA metadata.
Dataset Details
Uses
Direct Use
Continued pretraining / domain adaptation of encoder-decoder LMs on… See the full description on the dataset page: https://huggingface.co/datasets/rausch/scientific_corpus_cleaned.
