datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
biology
CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society
Github: https://github.com/lightaime/camel
Website: https://www.camel-ai.org/
Arxiv Paper: https://arxiv.org/abs/2303.17760
Dataset Summary
Biology dataset is composed of 20K problem-solution pairs obtained using gpt-4. The dataset problem-solutions pairs generating from 25 biology topics, 25 subtopics for each topic and 32 problems for each "topic,subtopic" pairs.
We provide… See the full description on the dataset page: https://huggingface.co/datasets/camel-ai/biology.task686_mmmlu_answer_generation_college_biology
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task686_mmmlu_answer_generation_college_biology
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task686_mmmlu_answer_generation_college_biology.task699_mmmlu_answer_generation_high_school_biology
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task699_mmmlu_answer_generation_high_school_biology
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task699_mmmlu_answer_generation_high_school_biology.nishy-al-biology-adaptive-dataset
Nishy A/L Biology Adaptive Dataset
This repository contains Biology MCQ datasets prepared for the Nishy adaptive tutoring and assessment system for Sri Lankan G.C.E. A/L Biology.
Repository structure
Master audited dataset
biology_master_1500_final_audited.json
Final audited master collection containing 1500 MCQs.
V4 paper-level split
v4_train_base_1182.json
v4_validation_base_90.json
v4_test_final_76.json
These files represent… See the full description on the dataset page: https://huggingface.co/datasets/Nishy11/nishy-al-biology-adaptive-dataset.reason-qa-biology-finetune-preview
Reasoning · Biology · Finetuning · Preview (Synthetic)
A public, single-generator preview of a larger private biology reasoning corpus.
This dataset has been created with gpt-oss-20b output and uses a simplified three-field format.
The full set spans many generator models, two reasoning styles (linear and
branching), and a richer schema (metadata, instruction, thinking, reasoning, answer).
Synthetic question-reasoning-answer data for domain finetuning on biology and
biochemistry… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/reason-qa-biology-finetune-preview.ToT-Biology
The ToT-Biology dataset emphasizes mechanistic understanding and explanatory biological reasoning, rather than just providing correct answers. It aims to train AI models in interpretability and logical deduction within the biological realm. Spanning a wide range of biological complexities, it starts with foundational concepts in cell biology, genetics, and ecology, and progresses to advanced areas like systems biology, synthetic biology, and computational biophysics. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/mattwesney/ToT-Biology.BiologyConcepts-Instruct-v1
BiologyConcepts-Instruct-v1
BiologyConcepts-Instruct-v1 is a synthetic biology instruction dataset designed for supervised fine-tuning of language models on fundamental and advanced biology concepts. It covers diverse domains including cell biology, genetics and heredity, anatomy and physiology, ecology and evolution, and energy metabolism through clear explanations, biological intuition, worked examples, and educational discussions. The dataset is suitable for training… See the full description on the dataset page: https://huggingface.co/datasets/kd13/BiologyConcepts-Instruct-v1.computational_biology_dataset
Additional Information
This dataset contains medicine problems generated using the CAMEL framework. Each entry includes:
A question
A detailed rationale explaining the solution approach
The llm_answer
hsc-biology-bangla-dataset
🌿 HSC Biology Bangla Dataset (Plant Physiology)
The Ultimate Resource for Bengali STEM NLP
This dataset is a large-scale collection of 10,000 instruction-response pairs meticulously generated from core HSC (Higher Secondary Certificate) Biology curriculum content. It focuses specifically on Plant Physiology (উদ্ভিদ শারীরতত্ত্ব), one of the most significant chapters for Bangladeshi students and medical aspirants.
✨ Key Highlights
Native Language… See the full description on the dataset page: https://huggingface.co/datasets/3amthoughts/hsc-biology-bangla-dataset.biology-ptbr
Tradução do Camel Biology dataset para Portuguese (PT-BR) usando NLLB 3.3b.
CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society
Github: https://github.com/lightaime/camel
Website: https://www.camel-ai.org/
Arxiv Paper: https://arxiv.org/abs/2303.17760
Dataset Summary
Biology dataset is composed of 20K problem-solution pairs obtained using gpt-4. The dataset problem-solutions pairs generating from 25 biology topics, 25… See the full description on the dataset page: https://huggingface.co/datasets/botbotrobotics/biology-ptbr.NCERT_Biology_11threwrite-questions-nonsensical-biology
nonsensical_biology.csv - Question Rewriting Dataset
This dataset contains question rewriting outputs from the file nonsensical_biology.csv.
Dataset Structure
The dataset contains the following columns:
custom_id: Unique identifier for each question
style: Rewriting style applied (e.g., "gibberish")
index: Numerical index
original: Original question text
rewritten: Rewritten version of the question
options: Multiple choice options (list format)
correct: Index of the… See the full description on the dataset page: https://huggingface.co/datasets/NLie2/rewrite-questions-nonsensical-biology.NCERT_Biology_12thcamel-ai_biology-gemini-exp-1206-ShareGPT
camel-ai_biology-gemini-exp-1206-ShareGPT
camel-ai/biology with responses generated with gemini-exp-1206.
Generation Details
If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the sample was skipped.
If ["candidates"][0]["safety_ratings"] == "SAFETY" the sample was skipped.
If ["candidates"][0]["finish_reason"] != 1 the sample was skipped.
model = genai.GenerativeModel(
model_name,
safety_settings=[
{… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/camel-ai_biology-gemini-exp-1206-ShareGPT.ViEduQA-biology
Dataset Access Request
This dataset is not publicly available.If you wish to request access, please send an email to:
📧 lehuuloi.cs@gmail.com
Important:
Requests will only be considered if sent from an official organizational email address (e.g., from a university, research institute, company, or non-profit organization).
Your email request must include:
Full name
Name of organization and position/title
Intended purpose and scope of use for the dataset
Requests that… See the full description on the dataset page: https://huggingface.co/datasets/shnl/ViEduQA-biology.camel-ai_biology-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
camel-ai_biology-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
camel-ai/biology with responses regenerated with gemini-2.0-flash-thinking-exp-1219.
Generation Details
If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the sample was skipped.
If ["candidates"][0]["safety_ratings"] == "SAFETY" the sample was skipped.
If ["candidates"][0]["finish_reason"] != 1 the sample was skipped.
model = genai.GenerativeModel(
model_name… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/camel-ai_biology-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT.biology
CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society
Github: https://github.com/lightaime/camel
Website: https://www.camel-ai.org/
Arxiv Paper: https://arxiv.org/abs/2303.17760
Dataset Summary
Biology dataset is composed of 20K problem-solution pairs obtained using gpt-4. The dataset problem-solutions pairs generating from 25 biology topics, 25 subtopics for each topic and 32 problems for each "topic,subtopic" pairs.
We… See the full description on the dataset page: https://huggingface.co/datasets/dddjjjppp/biology.NCERT_Biology_11th
