datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
biology
CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society
Github: https://github.com/lightaime/camel
Website: https://www.camel-ai.org/
Arxiv Paper: https://arxiv.org/abs/2303.17760
Dataset Summary
Biology dataset is composed of 20K problem-solution pairs obtained using gpt-4. The dataset problem-solutions pairs generating from 25 biology topics, 25 subtopics for each topic and 32 problems for each "topic,subtopic" pairs.
We provide… See the full description on the dataset page: https://huggingface.co/datasets/camel-ai/biology.MedThinkVQA
MedThinkVQA
MedThinkVQA is an expert-annotated benchmark for multi-image diagnostic reasoning in radiology. Unlike prior medical VQA benchmarks that typically contain at most one image per case, MedThinkVQA requires models to extract evidence from each image, integrate cross-view information, and perform differential-diagnosis reasoning.
Links
GitHub: https://github.com/benluwang/MedThinkVQA
Leaderboard: https://benluwang.github.io/MedThinkVQA/
Submission Guide:… See the full description on the dataset page: https://huggingface.co/datasets/bio-nlp-umass/MedThinkVQA.biodiversity_heritage_library_filtered
Biodiversity Heritage Library
Description
The Biodiversity Heritage Library (BHL) is an open-access digital library for biodiversity literature and archives.
This dataset contains over 15 million public domain books and documents from the BHL collection.
These works were collected using the bulk data download interface provided by the BHL and were filtered based on their associated license metadata.
We use the optical character recognition (OCR)-generated text distributed… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/biodiversity_heritage_library_filtered.BioMatrix-SFT
BioMatrix-SFT
This is the supervised fine-tuning (SFT) / instruction-tuning corpus used to train BioMatrix, a multimodal foundation model that natively integrates 1D sequences, 3D structures, and natural language for both molecules and proteins within a single decoder-only architecture.
📄 Paper: BioMatrix: Towards a Comprehensive Biological Foundation Model Spanning the Modality Matrix of Sequences, Structures, and Language
💻 Code: https://github.com/QizhiPei/BioMatrix… See the full description on the dataset page: https://huggingface.co/datasets/QizhiPei/BioMatrix-SFT.biodiversity_heritage_library
Biodiversity Heritage Library
Description
The Biodiversity Heritage Library (BHL) is an open-access digital library for biodiversity literature and archives.
This dataset contains over 42 million public domain books and documents from the BHL collection.
These works were collected using the bulk data download interface provided by the BHL and were filtered based on their associated license metadata.
We use the optical character recognition (OCR)-generated text… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/biodiversity_heritage_library.BioBERT_ItalianFrom this repository you can download the BioBERT_Italian dataset.
BioBERT_Italian is the Italian translation of the original BioBERT dataset, composed by millions of abstracts of PubMed papers.
Due to the unavailability of an Italian equivalent for the millions of abstracts and full-text scientific papers used by English, BERT-based biomedical models, we leveraged machine translation to obtain an Italian biomedical corpus based on PubMed abstracts and train BioBIT.
Corpus statistics:
Total… See the full description on the dataset page: https://huggingface.co/datasets/IVN-RIN/BioBERT_Italian.TransCorpus-bio
TransCorpus-bio
TransCorpus-bio is a large-scale, parallel biomedical corpus consisting of PubMed abstracts (title + abstract), translated with the TransCorpus Toolkit using NLLB-200. It is designed to enable high-quality multi-lingual biomedical language modeling and downstream NLP research.
This dataset was restructured from five separate single-language repositories into one dataset with a config (tab in the dataset viewer) per language, and with each row carrying its source… See the full description on the dataset page: https://huggingface.co/datasets/jknafou/TransCorpus-bio.biographies
📚 Synthetic Biographies
Synthetic Biographies is a dataset designed to facilitate research in factual recall and representation learning in language models. It comprises synthetic biographies of fictional individuals, each associated with sampled attributes like birthplace, university, and employer. The dataset is intended to support training and evaluating small language models (LLMs), particularly in their ability to store and extract factual knowledge.
🧾 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/alex-karev/biographies.bioleaflets-biomedical-ner
Dataset Card for BioLeaflets Dataset
Dataset Summary
BioLeaflets is a biomedical dataset for Data2Text generation. It is a corpus of 1,336 package leaflets of medicines authorised in Europe, which were obtained by scraping the European Medicines Agency (EMA) website.
Package leaflets are included in the packaging of medicinal products and contain information to help patients use the product safely and appropriately.
This dataset comprises the large majority (∼ 90%) of… See the full description on the dataset page: https://huggingface.co/datasets/ruslan/bioleaflets-biomedical-ner.bioinstruct
Dataset Card for BioInstruct
GitHub repo: https://github.com/bio-nlp/BioInstruct
Dataset Summary
BioInstruct is a dataset of 25k instructions and demonstrations generated by OpenAI's GPT-4 engine in July 2023.
This instruction data can be used to conduct instruction-tuning for language models (e.g. Llama) and make the language model follow biomedical instruction better.
Improvements of Llama on 9 common BioMedical tasks are shown in the result section.
Taking… See the full description on the dataset page: https://huggingface.co/datasets/bio-nlp-umass/bioinstruct.task686_mmmlu_answer_generation_college_biology
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task686_mmmlu_answer_generation_college_biology
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task686_mmmlu_answer_generation_college_biology.biomed-fr-v3-enriched-softmin-standard
biomed-fr-v3-enriched-softmin-standard
This dataset is a quality-upsampled version of rntc/biomed-fr-v3-enriched using soft-min bottleneck sampling.
Preprocessing Method
Soft-min calculation:
Formula: s = (mean(q_k^p))^(1/p) where q_k are the 4 quality scores
Parameter p = -2.0
Weight computation:
Ratio preference (5 vs 1): R = 10
Gamma exponent: γ = 1.43 (computed as log(R)/log(5))
Weight formula: w = s^γ
Floor: w = max(w, median(w) × 0.05)
Resampling:
Target size:… See the full description on the dataset page: https://huggingface.co/datasets/rntc/biomed-fr-v3-enriched-softmin-standard.dbpedia-biomedical
DBpedia Categories
Dataset Description
Category relationships from DBpedia (English)
Original Source: https://downloads.dbpedia.org/repo/dbpedia/generic/categories/2022.12.01/categories_lang=en_articles.ttl.bz2
Dataset Summary
This dataset contains RDF triples from DBpedia Categories converted to HuggingFace dataset format
for easy use in machine learning pipelines.
Format: Originally turtle, converted to HuggingFace Dataset
Size: 3.0 GB (extracted)
Entities:… See the full description on the dataset page: https://huggingface.co/datasets/CleverThis/dbpedia-biomedical.biostars_qa
Dataset Summary
This dataset contains 4803 question/answer pairs extracted from the BioStars website. The site focuses on bioinformatics, computational genomics, and biological data analysis.
Dataset Structure
Data Fields
The data contains INSTRUCTION, RESPONSE, SOURCE, and METADATA fields. The format is described for LAION-AI/Open-Assistant
Dataset Creation
Curation Rationale
Questions were included if they were an accepted answer and the… See the full description on the dataset page: https://huggingface.co/datasets/cannin/biostars_qa.bio-safety-peft-lora
CBRN Safety Alignment & PEFT-LoRA Fine-Tuning Dataset
This repository contains the synthetic instruction-tuning dataset (.jsonl) designed for parameter-efficient fine-tuning (PEFT-LoRA) of edge language models (specifically Qwen/Qwen2.5-1.5B-Instruct).
The dataset is curated to evaluate and modify model logit distributions, persona attributions, and dual-use safety boundaries regarding Chemical, Biological, Radiological, and Nuclear (CBRN) risk scenarios.
🤖 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/devsgnr/bio-safety-peft-lora.BioPhys-Bridge
BioPhys-Bridge is a physics-grounded scientific reasoning dataset for AI-for-Science agents. The dataset subtype, Sci-Evo, represents each record as a Physics-Grounded Scientific Evolution Case linking:
physical model -> quantitative evidence -> biological mechanism -> agent decision
Each case is built from open-access scientific literature and includes evidence-linked text, tables, formulas, figure/caption blocks, normalized quantitative measurements, biophysical model fields, biological… See the full description on the dataset page: https://huggingface.co/datasets/qyxu1994/BioPhys-Bridge.BiochemForge
BiochemForge
BiochemForge is a provenance-first biology, chemistry, and biochemistry post-training mixture for
mechanistic explanation, quantitative derivation, experimental inference, and consistency between
reasoning and final answers.
Dataset summary
Slice
Records
Purpose
SFT train
99,773
Supervised post-training
SFT validation
2,052
Model selection and early stopping
SFT test
1,093
Internal held-out evaluation
Solver-verified records
27,657… See the full description on the dataset page: https://huggingface.co/datasets/0xKitkat/BiochemForge.BioProBench
BioProBench: Comprehensive Dataset and Benchmark in Biological Protocol Understanding and Reasoning
BioProBench is the first large-scale, integrated multi-task benchmark for biological protocol understanding and reasoning, specifically designed for large language models (LLMs). It moves beyond simple QA to encompass a comprehensive suite of tasks critical for procedural text comprehension.
Biological protocols are the fundamental bedrock of reproducible and safe life… See the full description on the dataset page: https://huggingface.co/datasets/bowenxian/BioProBench.task699_mmmlu_answer_generation_high_school_biology
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task699_mmmlu_answer_generation_high_school_biology
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task699_mmmlu_answer_generation_high_school_biology.ALIA-es-biomedical-synthetic-instructions
Dataset Introduction
The ALIA Spanish Biomedical Synthetic Instructions Corpus is a synthetic instruction-tuning resource in Spanish created under the ALIA project using the Magpie methodology. It was designed to train and evaluate language models in biomedical and healthcare tasks with controlled formats and large-scale supervision.
It contains:
639,456 instances
961,073,205 tokens
14 task modalities (clinical diagnosis, patient education, ethical reasoning, document… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-biomedical-synthetic-instructions.BioTool
BioTool
BioTool is a large-scale, function-calling benchmark and training corpus for the
biomedical domain. It pairs natural-language biomedical questions with the
correct tool call (function name + JSON arguments) that answers them, drawn
from 127 tools spanning the three flagship public APIs:
NCBI E-utilities (einfo, esearch, esummary, efetch, elink, ecitmatch) plus BLAST
UniProt REST (uniprotkb, uniref, uniparc, proteomes, taxonomy, keywords, human_diseases, …)
Ensembl REST… See the full description on the dataset page: https://huggingface.co/datasets/gxx27/BioTool.bcs-biostatistics-study
BCS Medical Dataset — biostatistika
Medicinski studijski materijal na bosanskom/hrvatskom/srpskom, obrađen automatizovanim
inbox pipeline-om (ekstrakcija, OCR, chunking, AI generacija s determinističkom validacijom).
Struktura
Fajl
Sadržaj
ispitna.jsonl
postojeća ispitna pitanja (stari testovi/zbornici): question, options, answer
qna.jsonl
AI-generirani QnA parovi (validacija V1-V4)
flashcards.jsonl / flashcards.csv
kartice front/back za učenje… See the full description on the dataset page: https://huggingface.co/datasets/adobug/bcs-biostatistics-study.ALIA-es-biomedical-pairs
Dataset Introduction
The ALIA Spanish Biomedical Retrieval Pairs Corpus contains tabular instances designed to train and evaluate retrieval-oriented models (e.g., dense retrievers / embedding encoders) using passage-grounded query data produced with a Qwen3-style prompting workflow integrated in the ALIA encoders pipeline.
It preserves provenance to the original document and chunk while exposing controls such as question type and difficulty (ranging from high_school to phd level).… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-biomedical-pairs.ALIA-es-biomedical
Dataset Introduction
The ALIA Spanish Biomedical Corpus constitutes a strategic data infrastructure designed to support research and innovation in the biomedical domain. By ensuring systematic access to multiple official medical repositories in a single consolidated dataset, it provides a robust foundation for Spanish-language BioNLP. With over 6 million instances and more than 4 billion tokens, it represents a relevant comprehensive corpus of biomedical and clinical-related… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-biomedical.reason-qa-biology-finetune-preview
Reasoning · Biology · Finetuning · Preview (Synthetic)
A public, single-generator preview of a larger private biology reasoning corpus.
This dataset has been created with gpt-oss-20b output and uses a simplified three-field format.
The full set spans many generator models, two reasoning styles (linear and
branching), and a richer schema (metadata, instruction, thinking, reasoning, answer).
Synthetic question-reasoning-answer data for domain finetuning on biology and
biochemistry… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/reason-qa-biology-finetune-preview.biomed-fr-v3-enriched-softmin-leger
biomed-fr-v3-enriched-softmin-leger
This dataset is a quality-upsampled version of rntc/biomed-fr-v3-enriched using soft-min bottleneck sampling.
Preprocessing Method
Soft-min calculation:
Formula: s = (mean(q_k^p))^(1/p) where q_k are the 4 quality scores
Parameter p = -0.7
Weight computation:
Ratio preference (5 vs 1): R = 5
Gamma exponent: γ = 1.00 (computed as log(R)/log(5))
Weight formula: w = s^γ
Floor: w = max(w, median(w) × 0.05)
Resampling:
Target size: Same… See the full description on the dataset page: https://huggingface.co/datasets/rntc/biomed-fr-v3-enriched-softmin-leger.BioManufacturingBench
BioManufacturingBench v1.0.0
BioManufacturingBench v1.0.0 is a 2,000-item benchmark for evidence-grounded
biomanufacturing reasoning. It covers evidence extraction, mass-balance calculation,
process diagnosis, microscopy count-range estimation, strict output formatting, and
abstention. Every primary score is computed by a deterministic rule; no score uses an
LLM judge. Public records are deliberately answer-free so the benchmark remains useful
for future evaluation.… See the full description on the dataset page: https://huggingface.co/datasets/capicu-ai/BioManufacturingBench.ToT-Biology
The ToT-Biology dataset emphasizes mechanistic understanding and explanatory biological reasoning, rather than just providing correct answers. It aims to train AI models in interpretability and logical deduction within the biological realm. Spanning a wide range of biological complexities, it starts with foundational concepts in cell biology, genetics, and ecology, and progresses to advanced areas like systems biology, synthetic biology, and computational biophysics. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/mattwesney/ToT-Biology.bioreason-pro-sft-reasoning-documents
BioReason-Pro SFT Reasoning Documents
Complete, self-contained training documents reconstructed from the BioReason-Pro SFT data, ready for LLM pre-training.
The upstream dataset wanglab/bioreason-pro-sft-reasoning-data
ships the assistant side of each training example (reasoning, final_answer) alongside the raw
biological context columns, but not the assembled prompt. The prompt cannot be recovered from the
data card alone, because two of its three parts were non-textual… See the full description on the dataset page: https://huggingface.co/datasets/timodonnell/bioreason-pro-sft-reasoning-documents.biographical
Biographical Dataset for Relation Extraction (RE)
Overview
This dataset is a reconstructed version of the Biographical Dataset, specifically designed for relation extraction (RE) tasks. It serves as a valuable resource for digital humanities (DH) and historical research, enabling the study of relationships within biographical data. The dataset is generated by automatically aligning sentences from Wikipedia articles with structured data sourced from platforms like… See the full description on the dataset page: https://huggingface.co/datasets/Despina/biographical.
