datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
stem-scientific-code-sample
AxiomSet Labs STEM Scientific-Code Sample
A 30-task sample of STEM reasoning and scientific-code problems across five domains.
Domains
Biology: 6 tasks
Chemistry: 6 tasks
Materials Science: 6 tasks
Mathematics: 6 tasks
Physics: 6 tasks
Each task contains two subproblems and one main problem, including prompts, scientific background, testing templates, and reference solutions.
Files
data/sample.jsonl — one task per line; used by the Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/AxiomSetLabs/stem-scientific-code-sample.scientific_papers-archive
Dataset Card for "scientific_papers"
Dataset Summary
Scientific papers datasets contains two sets of long and structured documents.
The datasets are obtained from ArXiv and PubMed OpenAccess repositories.
Both "arxiv" and "pubmed" have two features:
article: the body of the document, paragraphs separated by "/n".
abstract: the abstract of the document, paragraphs separated by "/n".
section_names: titles of sections, separated by "/n".
Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/scillm/scientific_papers-archive.scientific_papers_DANCERscientific-question-outcomes
Scientific Question Outcomes
980 astronomy research questions, frozen at five historical cutoffs, each
labelled with what the following five years of literature actually did with
it.
Systems that propose research questions are usually evaluated by asking a
person or a model how good the questions sound. This dataset supplies the
alternative: questions frozen using only pre-cutoff literature, and outcome
labels drawn from the literature published afterwards. It is, to our… See the full description on the dataset page: https://huggingface.co/datasets/huiluckylucky/scientific-question-outcomes.scientific-exaggeration-detection
Dataset Card for Scientific Exaggeration Detection
Dataset Summary
Public trust in science depends on honest and factual communication of scientific papers. However, recent studies have demonstrated a tendency of news media to misrepresent scientific papers by exaggerating their findings. Given this, we present a formalization of and study into the problem of exaggeration detection in science communication. While there are an abundance of scientific papers and popular… See the full description on the dataset page: https://huggingface.co/datasets/copenlu/scientific-exaggeration-detection.groundtruthdouvras-scientific-ci-evidence-graph
Douvras Scientific CI Evidence Graph v0.1
Synthetic protocol dataset for linking a claim to its paper, repository,
dataset, seed and reproduced metric. It contains 30 records from six toy paper
instances (20 train, 5 validation and 5 frozen test), split by paper_id.
The labels distinguish REPRODUCED, PARTIAL, FAILED and INCONCLUSIVE.
Shortcuts and leakage fail closed. No real paper, code, dataset or result is
included, and this release is not a reproduction benchmark.
CoT-Scientific-RAG-Reasoning
CoT-Scientific-RAG-Reasoning
This dataset is designed for fine-tuning Large Language Models (specifically Qwen-series) to perform complex reasoning over scientific and technical documents using Chain-of-Thought (CoT).
Dataset Description
The dataset contains instructions and scientific contexts (Medical Imaging, Autonomous Driving, VLA Frameworks) where the model is required to generate a reasoning trace before providing the final answer.
Format: JSONL
Logic: All outputs… See the full description on the dataset page: https://huggingface.co/datasets/abhinavdread/CoT-Scientific-RAG-Reasoning.douvras-tamesis-scientific-validation
Douvras Tamesis Scientific Validation Dataset
Registros sintéticos para classificar a força e o estado de uma evidência científica: suportada,
inconclusiva, falsificada, replicação inconclusiva, não verificada ou alegação proibida.
Preserva resultados negativos e não transforma um proxy em descoberta. Não contém dados científicos
brutos, dados clínicos ou resultados novos; não prova qualquer hipótese Tamesis.
scientific-papers-dataset
Scientific Papers Dataset
Scientific papers, whitepapers and documentation.
Part of the Agnuxo Ecosystem by Francisco Angulo de Lafuente.
scientific-calculation-test!!!当前数据集仅为了方便测试使用,不保证题目答案正确!!!
!!!如想用于科学研究,请留意后续正式发布!!!
ecocoder-scientific-reasoning
ecocoder-scientific-reasoning
Chain-of-Thought (CoT) traces for fine-tuning LLMs on ecological scientific reasoning + code generation.
Each trace follows: [CONTEXT] (ecological problem) → [REASONING] (step-by-step scientific thinking) → [CODE] (runnable R/Python implementation).
Dataset Summary
Split
Traces
Train
1,268
Val
159
Test
159
Total
1,586
73 unique ecological methods across 18 categories
Languages: ~60% R, ~40% Python… See the full description on the dataset page: https://huggingface.co/datasets/alrobles/ecocoder-scientific-reasoning.scientific-posttrain-raman-eval
Scientific Post-Training Raman Evaluation Resolver v1
This repository is the canonical provenance manifest for the
raman-bioprocess Harbor task. It intentionally contains no Raman spectra or
labels because the upstream RamanBench mirror requires users to respect each
original dataset's terms and prohibits unapproved redistribution.
manifest.json pins the public Hugging Face source, exact commit, four Parquet
file hashes, deterministic split, and target definitions. During the… See the full description on the dataset page: https://huggingface.co/datasets/aashay96/scientific-posttrain-raman-eval.IN-Scientific
📥 IN-Scientific
IN-Scientific: An Open Multimodal Interleaved Dataset for Scientific Knowledge Representation
This project is a subproject of the 📌PIN project, focusing on the development of the largest scientific document multimodal dataset, which integrates both text and images.
📑: https://arxiv.org/abs/2406.13923
🤗: https://huggingface.co/datasets/m-a-p/PIN-14M
Dataset statistics
Source
Content Images (#)
Content Images (Size GB)
Documents (#)
Documents… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/IN-Scientific.adaption-scientific-research
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-scientific research
This dataset contains multi-turn dialogues where users pose scientific and technical questions across domains like physics, climate science, and machine learning. Assistants respond with conceptual explanations and executable Python code snippets to demonstrate calculations or simulate scenarios. Each sample follows a 'Before/After' structure, comparing initial… See the full description on the dataset page: https://huggingface.co/datasets/Azfarhashmi/adaption-scientific-research.persian-scientific-qa-triplets
Dataset Description
This dataset contains 1,016 high-quality triplets specifically designed for fine-tuning sentence-embedding models for retrieval tasks in the Persian scientific domain. Each triplet consists of a (query, positive, negative) tuple, making it ideal for contrastive learning.
Query: A relevant question about a scientific topic.
Positive: The ground-truth abstract that correctly answers the question.
Negative: A "hard negative" abstract that is semantically… See the full description on the dataset page: https://huggingface.co/datasets/safora/persian-scientific-qa-triplets.Scientific_Dialogue-ShareGPTConverted, deslopped, min-hash deduplicated, rejection filtered, grammar corrected using: https://github.com/The-Chaotic-Neutrals/ShareGPT-Formaxxing
Scientific-dataset-on-articles-small-thinkScientific Dataset on Articles (Small)
Это специализированный русскоязычный датасет небольшого объема (~1.2 тыс. строк), содержащий текстовую информацию, извлеченную из научных журналов по биологии (энтомология, арахнология, палеонтология), биографических очерков ученых-исследователей, а также дополненную тематическими материалами из открытых источников интернета.
Описание датасета
Датасет спроектирован для задач извлечения знаний (Information Extraction), ответов на вопросы по научным текстам… See the full description on the dataset page: https://huggingface.co/datasets/HoundyWoundy/Scientific-dataset-on-articles-small-think.adaption-experimentiq-scientific-reasoning-instruction-dataset-v1
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-ExperimentIQ – Scientific Reasoning Instruction Dataset
This dataset contains diverse samples focused on scientific methodology, experimental design, and data analysis across chemistry and biology. It includes tasks such as defining core concepts, correcting procedural errors, optimizing reaction conditions using Bayesian principles, and interpreting measurement accuracy. The content… See the full description on the dataset page: https://huggingface.co/datasets/Manan2802/adaption-experimentiq-scientific-reasoning-instruction-dataset-v1.scientific-reasoning-dataset
Scientific Reasoning Dataset
Synthetic reasoning traces based on Agnuxo research.
Part of the Agnuxo Ecosystem by Francisco Angulo de Lafuente.
Scientific-dataset-on-articles-smallScientific Dataset on Articles (Small)
Это специализированный русскоязычный датасет небольшого объема (~1.5 тыс. строк), содержащий текстовую информацию, извлеченную из научных журналов по биологии (энтомология, арахнология, палеонтология), биографических очерков ученых-исследователей, а также дополненную тематическими материалами из открытых источников интернета.
Описание датасета
Датасет спроектирован для задач извлечения знаний (Information Extraction), ответов на вопросы по научным текстам… See the full description on the dataset page: https://huggingface.co/datasets/HoundyWoundy/Scientific-dataset-on-articles-small.cleand_moremilk_CoT_Reasoning_Scientific_Discovery_and_Research元データ: https://huggingface.co/datasets/moremilk/CoT_Reasoning_Scientific_Discovery_and_Research
使用したコード: https://github.com/LLMTeamAkiyama/0-data_prepare/tree/master/src/CoT_Reasoning_Scientific_Discovery_and_Research
データ件数: 3,733
平均トークン数: 1,193
最大トークン数: 2,489
合計トークン数: 4,453,517
ファイル形式: JSONL
ファイル分割数: 1
合計ファイルサイズ: 23.2 MB
加工内容:
メタデータ列の解析と新列生成: metadata列(辞書型)を解析し、その中のreasoningをthought列に、difficultyをdifficulty列に展開しました。解析に失敗した行は除外されました。また、元のmetadata列は削除されました。
難易度によるフィルタリング:… See the full description on the dataset page: https://huggingface.co/datasets/LLMTeamAkiyama/cleand_moremilk_CoT_Reasoning_Scientific_Discovery_and_Research.AI-taste-scientificdrtulu_v2_stepfun_scientific_knowledge_0415adaption-experimentiq-scientific-reasoning-instruction-dataset
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-ExperimentIQ – Scientific Reasoning Instruction Dataset
This dataset contains diverse samples focused on scientific methodology, experimental design, and data analysis across chemistry and biology. It includes tasks such as defining core concepts, correcting procedural errors, optimizing reaction conditions using Bayesian principles, and interpreting measurement accuracy. The content… See the full description on the dataset page: https://huggingface.co/datasets/Manan2802/adaption-experimentiq-scientific-reasoning-instruction-dataset.scientific-papers-sft-v1
Scientific Papers SFT Queue
Instruction-tuning style records derived from scientific paper passages for summarization and keyword extraction tasks.
Splits in this repo
pilot/sft_review_queue_10.jsonl: 10-item pilot split for quick validation
sft_review_queue_500.jsonl: 500-item review queue
Schema
Each JSONL row has:
task: summary or keywords
instruction: task instruction
input: source passage
draft_output: initial target output
review_status: review state… See the full description on the dataset page: https://huggingface.co/datasets/YanJo199/scientific-papers-sft-v1.sft_ablations_scientific_minimax_v1scientific_papers_datasetjson_1000_Scientific_Paperbwb-scientific-academic-intel
Scientific & Academic Intelligence
BWB Data Store. 220 records. Omega schema.
Full dataset: https://data.bankingwithbilly.com/dataset/scientific-academic
CC BY 4.0 — Banking With Billy
