datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Scientific_Research_Tokenized
NexaSci Scientific Research Tokenized
This dataset repository now holds the active NexaSci scientific pretraining reservoir, the NexaMat controller fine-tuning pack, and archived legacy reservoir builds. The current production reservoir is the 10B-token Apache Arrow release under nexasci_reservoir_v3_10b_prod_rust/.
Current Status
The active large-scale training artifact is:
nexasci_reservoir_v3_10b_prod_rust/
It was produced from the NexaSci 10B data-engineering campaign… See the full description on the dataset page: https://huggingface.co/datasets/AethronPhantom/Scientific_Research_Tokenized.scientific-verification
Scientific Verification Benchmark: NMC Cathodes
Dataset summary
The benchmark contains 50 scientific claims about NMC (lithium nickel manganese cobalt oxide) battery cathodes. Each claim is answered by Claude Opus 5, GPT 5.6 Luna and Gemini 3.1 Pro using a set of 20 open-access papers, producing 150 scored answers. The accompanying reference set contains 1,991 experiment-grounded measurements curated from 227 open-access papers, with experimental conditions and… See the full description on the dataset page: https://huggingface.co/datasets/GenData-Research/scientific-verification.scientific-literature-research-assistant-dataadaption-scientific-research
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-scientific research
This dataset contains multi-turn dialogues where users pose scientific and technical questions across domains like physics, climate science, and machine learning. Assistants respond with conceptual explanations and executable Python code snippets to demonstrate calculations or simulate scenarios. Each sample follows a 'Before/After' structure, comparing initial… See the full description on the dataset page: https://huggingface.co/datasets/Azfarhashmi/adaption-scientific-research.scientific-research
Description
Topic: Scientific Research Papers
Domains: Biology, Physics, Chemistry
Number of Entries: 1,000
Dataset Type: Raw Dataset
Model Used: Meta Llama4 Maverick 17B Instruct
Language: English
cleand_moremilk_CoT_Reasoning_Scientific_Discovery_and_Research元データ: https://huggingface.co/datasets/moremilk/CoT_Reasoning_Scientific_Discovery_and_Research
使用したコード: https://github.com/LLMTeamAkiyama/0-data_prepare/tree/master/src/CoT_Reasoning_Scientific_Discovery_and_Research
データ件数: 3,733
平均トークン数: 1,193
最大トークン数: 2,489
合計トークン数: 4,453,517
ファイル形式: JSONL
ファイル分割数: 1
合計ファイルサイズ: 23.2 MB
加工内容:
メタデータ列の解析と新列生成: metadata列(辞書型)を解析し、その中のreasoningをthought列に、difficultyをdifficulty列に展開しました。解析に失敗した行は除外されました。また、元のmetadata列は削除されました。
難易度によるフィルタリング:… See the full description on the dataset page: https://huggingface.co/datasets/LLMTeamAkiyama/cleand_moremilk_CoT_Reasoning_Scientific_Discovery_and_Research.
