scientific-reasoning
Scientific-Reasoning
Spark-234K: Skeleton-Guided Scientific Reasoning from Large-Scale Literature
🎉 Accepted to EMNLP 2026 Findings!
Spark-234K is a scientific reasoning dataset containing 234K question-answer pairs synthesized from frontier scientific literature. Instead of directly generating QA pairs from full papers, SPARK first distills each paper into a compact reasoning skeleton—preserving its central claim, supporting evidence, quantitative relations, assumptions, and boundary… See the full description on the dataset page: https://huggingface.co/datasets/P0u4a/Scientific-Reasoning.CoT-Scientific-RAG-Reasoning
CoT-Scientific-RAG-Reasoning
This dataset is designed for fine-tuning Large Language Models (specifically Qwen-series) to perform complex reasoning over scientific and technical documents using Chain-of-Thought (CoT).
Dataset Description
The dataset contains instructions and scientific contexts (Medical Imaging, Autonomous Driving, VLA Frameworks) where the model is required to generate a reasoning trace before providing the final answer.
Format: JSONL
Logic: All outputs… See the full description on the dataset page: https://huggingface.co/datasets/abhinavdread/CoT-Scientific-RAG-Reasoning.ecocoder-scientific-reasoning
ecocoder-scientific-reasoning
Chain-of-Thought (CoT) traces for fine-tuning LLMs on ecological scientific reasoning + code generation.
Each trace follows: [CONTEXT] (ecological problem) → [REASONING] (step-by-step scientific thinking) → [CODE] (runnable R/Python implementation).
Dataset Summary
Split
Traces
Train
1,268
Val
159
Test
159
Total
1,586
73 unique ecological methods across 18 categories
Languages: ~60% R, ~40% Python… See the full description on the dataset page: https://huggingface.co/datasets/alrobles/ecocoder-scientific-reasoning.openpipe-dpo-scientific-reasoning
Openpipe Dpo Scientific Reasoning
This dataset contains 100 high-quality examples for Direct Preference Optimization (DPO) training, formatted for OpenPipe fine-tuning, focused on scientific reasoning and analysis.
Dataset Description
This dataset was generated using an enhanced DSPy-based pipeline that creates structured reasoning traces for scientific questions. Each example follows the OpenAI chat completion format required by OpenPipe:
OpenAI Chat Format: Standard… See the full description on the dataset page: https://huggingface.co/datasets/abhi26/openpipe-dpo-scientific-reasoning.dpo-scientific-reasoning
Dpo Scientific Reasoning
This dataset contains 100 high-quality examples for Direct Preference Optimization (DPO) training, focused on scientific reasoning and analysis.
Dataset Description
This dataset was generated using an enhanced DSPy-based pipeline that creates structured reasoning traces for scientific questions. Each example includes:
Separated content fields: System prompt, user question, and full context as individual columns
Chosen responses: High-quality… See the full description on the dataset page: https://huggingface.co/datasets/abhi26/dpo-scientific-reasoning.adaption-experimentiq-scientific-reasoning-instruction-dataset-v1
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-ExperimentIQ – Scientific Reasoning Instruction Dataset
This dataset contains diverse samples focused on scientific methodology, experimental design, and data analysis across chemistry and biology. It includes tasks such as defining core concepts, correcting procedural errors, optimizing reaction conditions using Bayesian principles, and interpreting measurement accuracy. The content… See the full description on the dataset page: https://huggingface.co/datasets/Manan2802/adaption-experimentiq-scientific-reasoning-instruction-dataset-v1.
