datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
scientific_data_2026_curated_trackssim-datasets
SIM-Datasets: A Unified Symbolic Regression Benchmark
A standardized benchmark collection designed for the Scientific Intelligent Modelling (SIM) toolkit, providing comprehensive datasets for symbolic regression research and applications.
Overview
SIM-Datasets serves as a unified benchmark for symbolic regression tasks, offering standardized datasets with consistent formatting and evaluation protocols. This collection is specifically curated to support the Scientific… See the full description on the dataset page: https://huggingface.co/datasets/scientific-intelligent-modelling/sim-datasets.scientific-stuff-1scientific-stuff-2scientific-stuff-4scientific_data_2026_images
Image data for Hardo, Li, and Bakshi, 2026
Dataset summary
This repository contains the metadata-complete trench image store for the paper "An annotated timelapse imaging dataset on dormancy exit dynamics of Escherichia coli cells in Mother Machine".
The data consists of one xarray-compatible Zarr v2 store:
20260307_SB7_exit_snake_V4_1_with_metadata.trenches.zarr
The store occupies approximately 20 GB on disk.
This store contains extracted mother-machine trench movies… See the full description on the dataset page: https://huggingface.co/datasets/ghardo/scientific_data_2026_images.LimAgents_limitation_data_scientific_papers_with_cited_papers
LimAgents Data
This dataset contains scientific paper metadata and extracted limitation information prepared for use with LLM Agents.The data comes from NeurIPS 2021–2022 papers and related OpenReview reviews, enriched with Cited in and Cited by information.
Dataset Structure
The repository contains two main directories:
1. NeurIPS_21_22_Lim_OPR_with_cited_in_by_papers
This directory includes one JSON file per paper. Each file contains:
title: Original paper… See the full description on the dataset page: https://huggingface.co/datasets/iaadlab/LimAgents_limitation_data_scientific_papers_with_cited_papers.autonlp-data-Scientific_Title_Generator
AutoNLP Dataset for project: Scientific_Title_Generator
Table of content
Dataset Description
Languages
Dataset Structure
Data Instances
Data Fields
Data Splits
Dataset Descritpion
This dataset has been automatically processed by AutoNLP for project Scientific_Title_Generator.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as… See the full description on the dataset page: https://huggingface.co/datasets/AryanLala/autonlp-data-Scientific_Title_Generator.scientific-challenges-and-directions-dataset
Dataset Card for scientific-challenges-and-directions
Dataset Summary
The scientific challenges and directions dataset is a collection of 2894 sentences and their surrounding contexts, from 1786 full-text papers in the CORD-19 corpus, labeled for classification of challenges and directions by expert annotators with biomedical and bioNLP backgrounds.
At a high level, our labels are defined as follows:
Challenge: A sentence mentioning a problem, difficulty, flaw… See the full description on the dataset page: https://huggingface.co/datasets/DanL/scientific-challenges-and-directions-dataset.table-extraction-scientific-datasets
Benchmarking Table Extraction from Heterogeneous Scientific PDF Documents
Dataset Sources
Repository: GitHub
Code archive: Zenodo (DOI 10.5281/zenodo.20486311)
Paper: 10.1145/3770855.3817462
Extended version: arXiv:2511.16134
Interactive demo: table-extraction-benchmark-explorer.streamlit.app
(source)
Dataset Details
PubTables (pubtables/*.tar.gz)
Subset of PubTables-Test dataset, enriched with HTML table ground… See the full description on the dataset page: https://huggingface.co/datasets/marijanic/table-extraction-scientific-datasets.datapoints_round1_dpsk_scientific_computing_shard1_daytona_n100k1scientific-sft-grpo-datadatapoints_round1_dpsk_scientific_computing_shard2_daytona_n50k1Scientific-dataset-on-articles-small-thinkScientific Dataset on Articles (Small)
Это специализированный русскоязычный датасет небольшого объема (~1.2 тыс. строк), содержащий текстовую информацию, извлеченную из научных журналов по биологии (энтомология, арахнология, палеонтология), биографических очерков ученых-исследователей, а также дополненную тематическими материалами из открытых источников интернета.
Описание датасета
Датасет спроектирован для задач извлечения знаний (Information Extraction), ответов на вопросы по научным текстам… See the full description on the dataset page: https://huggingface.co/datasets/HoundyWoundy/Scientific-dataset-on-articles-small-think.scientific-papers-dataset
Scientific Papers Dataset
Scientific papers, whitepapers and documentation.
Part of the Agnuxo Ecosystem by Francisco Angulo de Lafuente.
cpugym-v9-scientific-dataset
CPUGym V9 Scientific Dataset
This dataset package contains the durable telemetry and evaluation evidence used to freeze the V9 release:
root dashboard SQLite database
live p3 and p2r SQLite databases
Azure evaluation artifacts
release manifest and scientific report
monitoring logs and V9 deployment documentation
compact-scientific-lm-datascientific-literature-research-assistant-datascientific_data_2026_masks
Mask data for Hardo, Li, and Bakshi, 2026
Dataset summary
This repository contains the released multi-hypothesis instance segmentation masks for the paper "An annotated timelapse imaging dataset on dormancy exit dynamics of Escherichia coli cells in Mother Machine".
The release consists of one Zarr store:
20260307_SB7_exit_snake_V4_1.segmentation_masks_multi_epoch_uint8_masks_only.zarr
The store occupies approximately 4.6 GB on disk and contains only the… See the full description on the dataset page: https://huggingface.co/datasets/ghardo/scientific_data_2026_masks.Scientific-dataset-on-articles-smallScientific Dataset on Articles (Small)
Это специализированный русскоязычный датасет небольшого объема (~1.5 тыс. строк), содержащий текстовую информацию, извлеченную из научных журналов по биологии (энтомология, арахнология, палеонтология), биографических очерков ученых-исследователей, а также дополненную тематическими материалами из открытых источников интернета.
Описание датасета
Датасет спроектирован для задач извлечения знаний (Information Extraction), ответов на вопросы по научным текстам… See the full description on the dataset page: https://huggingface.co/datasets/HoundyWoundy/Scientific-dataset-on-articles-small.scientific-paragraphs-categorization
A Multi-lingual Dataset of Classified Paragraphs from Open Access Scientific
We present a dataset of 833k paragraphs extracted from CC-BY licensed
scientific publications, classified into four categories: acknowledgments, data
mentions, software/code mentions, and clinical trial mentions. The paragraphs
are primarily in English and French, with additional European languages
represented. Each paragraph is annotated with language identification (using
fastText) and scientific domain… See the full description on the dataset page: https://huggingface.co/datasets/dataesr/scientific-paragraphs-categorization.scientific-sft-grpo-data-trialScientific-Laboratory-Instrument-Readout-Recognition-Dataset
Scientific Laboratory Instrument Readout Recognition Dataset
In modern scientific research, laboratory instruments are key sources of data acquisition, while manual recording of instrument readouts is inefficient and prone to errors. Traditional solutions such as manual transcription and basic OCR technology often prove inadequate when dealing with complex backgrounds, reflections, and multiple fonts of scientific instrument readings. This dataset aims to solve the problem of… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Scientific-Laboratory-Instrument-Readout-Recognition-Dataset.adaption-experimentiq-scientific-reasoning-instruction-dataset-v1
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-ExperimentIQ – Scientific Reasoning Instruction Dataset
This dataset contains diverse samples focused on scientific methodology, experimental design, and data analysis across chemistry and biology. It includes tasks such as defining core concepts, correcting procedural errors, optimizing reaction conditions using Bayesian principles, and interpreting measurement accuracy. The content… See the full description on the dataset page: https://huggingface.co/datasets/Manan2802/adaption-experimentiq-scientific-reasoning-instruction-dataset-v1.scientific_papers_summary_datasetscientific-stuff-3scientific-reasoning-dataset
Scientific Reasoning Dataset
Synthetic reasoning traces based on Agnuxo research.
Part of the Agnuxo Ecosystem by Francisco Angulo de Lafuente.
scientific-style-gen-data
Scientific Style Generation Datasets
This repository contains the dataset collection accompanying the Master's thesis "Improving Few-Shot Capabilities of LLMs for Style-Conditioned Text Generation".
GitHub Repository: MadnessOverflow/scientific-style-gen
Licensing & Attribution
Overall Repository License: CC BY-NC-SA 4.0 (governed by the most restrictive sub-dataset).
Sub-Dataset Licensing Breakdown
Abstract-Only Datasets (CC0 1.0 Universal /… See the full description on the dataset page: https://huggingface.co/datasets/MadnessOverflow/scientific-style-gen-data.armanc_scientific_papers_arxiv_datasetadaption-experimentiq-scientific-reasoning-instruction-dataset
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-ExperimentIQ – Scientific Reasoning Instruction Dataset
This dataset contains diverse samples focused on scientific methodology, experimental design, and data analysis across chemistry and biology. It includes tasks such as defining core concepts, correcting procedural errors, optimizing reaction conditions using Bayesian principles, and interpreting measurement accuracy. The content… See the full description on the dataset page: https://huggingface.co/datasets/Manan2802/adaption-experimentiq-scientific-reasoning-instruction-dataset.
