datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
chinese-materials-science-open-intelligence
🔬 Chinese Materials Science & Metallurgy Open Intelligence Dataset
Curated open intelligence dataset providing English research briefs, authoritative DOIs, executive summaries, and high-resolution micrographs of breakthrough Chinese scientific research in Materials Science, Metallurgy, Advanced Alloys, and Mining Engineering.
[!IMPORTANT]
Data Completeness & Research Authenticity Notice:
Included in this Hugging Face Open Dataset: English structured abstracts, core… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-materials-science-open-intelligence.realistic-bpe5-science-math-10bDataPRM-ScienceAgentBenchSO-Python_QA-Data_Science_and_Machine_Learning_classscience-hierarchography
Science Hierarchography: Hierarchical Organization of Science Literature
Dataset for the paper: Science hierarchography: Hierarchical organization of science literature
Dataset Description
This dataset contains academic research papers with structured annotations for building hierarchical taxonomies of scientific literature.
Configurations
Config
Papers
SciPile
1,938
SciPileLarge
9,978
Usage
from datasets import load_dataset
# Load… See the full description on the dataset page: https://huggingface.co/datasets/MuhanGao/science-hierarchography.frontier-science-bench
OpenEvolve FrontierScience Benchmark Report
Executive Summary
This report presents a comprehensive benchmark comparing three approaches for solving FrontierScience problems using the nvidia/nemotron-3-nano-30b-a3b model:
Zero-shot Code Generation - Single LLM call to generate Python code
Multi-turn Feedback - Up to 10 retry attempts with error feedback
OpenEvolve - Evolutionary code optimization with 10 iterations
Key Results
Method
Solved
Accuracy… See the full description on the dataset page: https://huggingface.co/datasets/batuhanaktas/frontier-science-bench.Got_Science_28k3rd-Degree-Burn__L-3.1-Science-Writer-8B-details
Dataset Card for Evaluation run of 3rd-Degree-Burn/L-3.1-Science-Writer-8B
Dataset automatically created during the evaluation run of model 3rd-Degree-Burn/L-3.1-Science-Writer-8B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/3rd-Degree-Burn__L-3.1-Science-Writer-8B-details.us-science-policy-bills
US Science Policy Bills
A stance-annotated corpus of 8,191 public-health-related bills from the United States federal Congress and 41 state legislatures, collected and classified by SAFE Action (Science and Freedom for Everyone Action Fund), a 501(c)(4) social welfare organization building open-source democratic tools for science-based policy.
To our knowledge this is the first openly published, stance-annotated dataset of state-level public health legislation. Raw legislative… See the full description on the dataset page: https://huggingface.co/datasets/SAFE-Action/us-science-policy-bills.grade3-science-explanations-v4r7
Grade-Level Science Explanations v4r7
The final 485-record supervised fine-tuning dataset for the grade-level science
explainer. Each record maps a unique elementary-science question to a concise,
mechanism-complete explanation. Training uses the minimal prompt
Explain: {phrasing} so the reading behavior must be learned from examples rather
than supplied through prompt instructions.
Files
File
Records
Purpose
gold_v4_r7.jsonl
485
Final training split… See the full description on the dataset page: https://huggingface.co/datasets/SAgarwal34/grade3-science-explanations-v4r7.synthetic.qa.science.llm💰 ¿Necesitas el dataset completo?
10M muestras por $50K.
Contacto: jary20@gmail.com
