datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
biology
CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society
Github: https://github.com/lightaime/camel
Website: https://www.camel-ai.org/
Arxiv Paper: https://arxiv.org/abs/2303.17760
Dataset Summary
Biology dataset is composed of 20K problem-solution pairs obtained using gpt-4. The dataset problem-solutions pairs generating from 25 biology topics, 25 subtopics for each topic and 32 problems for each "topic,subtopic" pairs.
We provide… See the full description on the dataset page: https://huggingface.co/datasets/camel-ai/biology.wikipedia-biology
Dataset Card for wikipedia-biology
Dataset Summary
The dataset consists of text from 87045 Wikipedia articles created by processing all articles in the Wikipedia categories Branches of biology, Biological concepts, Eukaryote biology and Biology terminology, as well as their subcategories recursively till a depth of 4. It was originally created for the purpose of unlearning the domain of biology, although it may be used for other purposes such as biology fine-tuning.
It… See the full description on the dataset page: https://huggingface.co/datasets/jd5697/wikipedia-biology.agieval-gaokao-biology
Dataset Card for "agieval-gaokao-biology"
Dataset taken from https://github.com/microsoft/AGIEval and processed as in that repo, following dmayhem93/agieval-* datasets on the HF hub.
This dataset contains the contents of the Gaokao Biology subtask of AGIEval, as accessed in https://github.com/ruixiangcui/AGIEval/commit/5c77d073fda993f1652eaae3cf5d04cc5fd21d40 .
Citation:
@misc{zhong2023agieval,
title={AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models}… See the full description on the dataset page: https://huggingface.co/datasets/hails/agieval-gaokao-biology.trustworthy-biology-agents-traces
Trustworthy Biology Agents — Run Traces
Raw execution traces from 1,329 agent runs across three coding agents on three
biology benchmarks — BiomniBench-DA, BixBench, and CompBioBench. This is the scrubbed
trace bundle for the study in
manu-tej/ai-scientists; the write-up
lives in that repo's RESULTS.md.
The motivating question is not only whether an agent reaches the right answer, but
whether it behaves like a trustworthy analyst when the task is ambiguous,
under-specified, or… See the full description on the dataset page: https://huggingface.co/datasets/amanutej/trustworthy-biology-agents-traces.agieval-gaokao-biology
Dataset Card for "agieval-gaokao-biology"
Dataset taken from https://github.com/microsoft/AGIEval and processed as in that repo.
MIT License
Copyright (c) Microsoft Corporation.
Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell… See the full description on the dataset page: https://huggingface.co/datasets/dmayhem93/agieval-gaokao-biology.mmlu_college_biologyBiology
🔭 Overview
R2MED: First Reasoning-Driven Medical Retrieval Benchmark
R2MED is a high-quality, high-resolution synthetic information retrieval (IR) dataset designed for medical scenarios. It contains 876 queries with three retrieval tasks, five medical scenarios, and twelve body systems.
Dataset
#Q
#D
Avg. Pos
Q-Len
D-Len
Biology
103
57359
3.6
115.2
83.6
Bioinformatics77
47473
2.9
273.8
150.5
Medical Sciences
88
34810
2.8
107.1
122.7
MedXpertQA-Exam
97… See the full description on the dataset page: https://huggingface.co/datasets/R2MED/Biology.biology_textbookoh_v1.2_sin_camel_biology_diversitypes2o_biology_780m_refinedweb_220m_before_20190101biology-german
License & Attribution
MTEB-format derivative of SolaireOfTheSun/Biology_German_DHBW (German biology Q&A). Query = question; corpus = answer. Licensed under BigScience OpenRAIL-M (same as source).
OpenSciReasoning-Biology-20K
OpenSciReasoning-Biology-20K
Three-domain release derived from nvidia/OpenScienceReasoning-2 for
domain-specific reasoner training and cross-domain transfer experiments.
Each row preserves the stable source_row_id and has exactly one mutually
exclusive domain value: BIOLOGY. Domain acceptance was checked from the
question and choices with two independent question-only verifiers; answer and
source-ID gates were also replayed.
The audit records list any remaining source-output… See the full description on the dataset page: https://huggingface.co/datasets/TerryJCZhang/OpenSciReasoning-Biology-20K.Handwritten-Biology-Notes-Dataset
English Handwritten Biology Notes Dataset
This dataset contains high-resolution images of handwritten biology notes written in English. The collection includes labeled diagrams, definitions, explanations of biological processes, and annotated sketches. It supports AI research in handwriting recognition, diagram understanding, and document interpretation within the field of life sciences.
Contact
For queries or collaborations related to this dataset, contact:… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Handwritten-Biology-Notes-Dataset.oh_v1.2_sin_camel_biology_diversityscience_biologydataset-CoT-Molecular-Biology-71task686_mmmlu_answer_generation_college_biology
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task686_mmmlu_answer_generation_college_biology
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task686_mmmlu_answer_generation_college_biology.jinyang-omentum-rmats-outlier-fixed-biology-results-v1
jinyang-omentum-rmats-outlier-fixed-biology-results-v1
Biology of the k=2 subgroups from jinyang-omentum-rmats-outlier-fixed-clustering (cluster1 n=116 / cluster2 n=50, 166 outlier-excluded omentum samples; expression analyses on the 160-sample TPM-labeled intersection). CAVEAT (stated plainly, not buried): cluster1/cluster2 are collinear with sequencing platform to within one sample -- cluster 2 contains ZERO NovaSeq-6000 samples and cluster 1 is 113/116 NovaSeq-6000. The… See the full description on the dataset page: https://huggingface.co/datasets/depinwang/jinyang-omentum-rmats-outlier-fixed-biology-results-v1.camel-ai-biology2-000-Physics-Chemistry-And-Biology-Questions-With-Image-Explanations-Dataset
2,000 Physics, Chemistry, and Biology Questions with Image Explanations Dataset
This collection contains 2,000 high-quality physics, chemistry, and biology questions, featuring a core modality of original images paired with text explanations. The data covers various formats, including multiple-choice, fill-in-the-blanks, experimental, and calculation questions. Each record provides the original question image, precise OCR-extracted text, and detailed step-by-step textual… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/2-000-Physics-Chemistry-And-Biology-Questions-With-Image-Explanations-Dataset.neet-biology-qa
NEET Biology Questions Dataset
A comprehensive collection of NEET (National Eligibility cum Entrance Test) Biology questions designed to help students prepare for medical entrance examinations.
About This Dataset
This dataset contains 793 carefully crafted multiple-choice questions covering essential Biology topics that appear in NEET exams. Each question follows the standard NEET format with four answer choices and one correct answer.
What's Inside
Questions:… See the full description on the dataset page: https://huggingface.co/datasets/sweatSmile/neet-biology-qa.task699_mmmlu_answer_generation_high_school_biology
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task699_mmmlu_answer_generation_high_school_biology
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task699_mmmlu_answer_generation_high_school_biology.camel_ai_biology_instruction_datasetontolearner-biology_and_life_sciences
Biology And Life Sciences Domain Ontologies
Overview
The biology and life sciences domain encompasses the structured representation and categorization of knowledge related to biological entities, processes, and systems, ranging from molecular and cellular levels to complex organisms and ecosystems. This domain plays a critical role in facilitating data interoperability, integration, and retrieval across diverse biological disciplines, thereby advancing research and… See the full description on the dataset page: https://huggingface.co/datasets/SciKnowOrg/ontolearner-biology_and_life_sciences.jinyang-omentum-pds-subtype-biology-results-v1
jinyang-omentum-pds-subtype-biology-results-v1
Whole-cohort Omental k=2 NMF subtype biology (S1 n=118 / S2 n=50, 168 samples; expression analyses on the 162-sample TPM-labeled intersection). Figures: survival-anchor KM, splicing/expression volcanos, ORA dot plots (6), ssGSEA splicing heatmap, per-gene splicing-vs-expression concordance scatter.
Dataset Info
Rows: 11
Columns: 2
Columns
Column
Type
Description
figure_name
Value('string')… See the full description on the dataset page: https://huggingface.co/datasets/depinwang/jinyang-omentum-pds-subtype-biology-results-v1.nishy-al-biology-adaptive-dataset
Nishy A/L Biology Adaptive Dataset
This repository contains Biology MCQ datasets prepared for the Nishy adaptive tutoring and assessment system for Sri Lankan G.C.E. A/L Biology.
Repository structure
Master audited dataset
biology_master_1500_final_audited.json
Final audited master collection containing 1500 MCQs.
V4 paper-level split
v4_train_base_1182.json
v4_validation_base_90.json
v4_test_final_76.json
These files represent… See the full description on the dataset page: https://huggingface.co/datasets/Nishy11/nishy-al-biology-adaptive-dataset.arxiv-biology
Dataset Curators
The original data is maintained by ArXiv
Licensing Information
The data is under the Creative Commons CC0 1.0 Universal Public Domain Dedication
Citation Information
@misc{clement2019arxiv,
title={On the Use of ArXiv as a Dataset},
author={Colin B. Clement and Matthew Bierbaum and Kevin P. O'Keeffe and Alexander A. Alemi},
year={2019},
eprint={1905.00075},
archivePrefix={arXiv},
primaryClass={cs.IR}
}
reason-qa-biology-finetune-preview
Reasoning · Biology · Finetuning · Preview (Synthetic)
A public, single-generator preview of a larger private biology reasoning corpus.
This dataset has been created with gpt-oss-20b output and uses a simplified three-field format.
The full set spans many generator models, two reasoning styles (linear and
branching), and a richer schema (metadata, instruction, thinking, reasoning, answer).
Synthetic question-reasoning-answer data for domain finetuning on biology and
biochemistry… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/reason-qa-biology-finetune-preview.biology-scienceqa
Dataset Card for "biology-scienceqa"
More Information needed
HLE_SFT_biology
HLE_SFT_biology
データセットの説明
このデータセットは、以下の分割(split)ごとに整理された処理済みデータを含みます。
ToT_biology: 1 JSON files, 1 Parquet files
データセット構成
各 split は JSON 形式と Parquet 形式の両方で利用可能です:
JSONファイル: 各 split 用サブフォルダ内の元データ(ToT_biology/)
Parquetファイル: split名をプレフィックスとした最適化データ(data/ToT_biology_*.parquet)
各 JSON ファイルには、同名の split プレフィックス付き Parquet ファイルが対応しており、大規模データセットの効率的な処理が可能です。
使い方
from datasets import load_dataset
# 特定の split を読み込む
ToT_biology_data =… See the full description on the dataset page: https://huggingface.co/datasets/neko-llm/HLE_SFT_biology.
