datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mmlu_medical_geneticstask721_mmmlu_answer_generation_medical_genetics
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task721_mmmlu_answer_generation_medical_genetics
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task721_mmmlu_answer_generation_medical_genetics.mmlu-medical_genetics-arabicgenetics-population-transfer-integrity-v0.1What this dataset tests
Whether genetic findings survivetransfer from the discovery populationto a different target population.
Required outputs
discovery population
target population
transfer assumptions
transfer failures
clinical risk if applied
Typical failures
European PRS used globally
penetrance assumed invariant
tag SNPs treated as causal across ancestries
Suggested prompt wrapper
System
You enforce population transfer integrity in genetics.
You prevent unsafe… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/genetics-population-transfer-integrity-v0.1.mmlu-medical_genetics
Dataset Card for "mmlu-medical_genetics"
More Information needed
Spanish-MMLU-Medical-Genetics-Benchmark
💻 Dataset Usage
Run the following command to load the testing set:
from datasets import load_dataset
dataset = load_dataset("shuyuej/Spanish-MMLU-Medical-Genetics-Benchmark", split="test")
print(dataset)
geneticsQA-corpus
Genetics QA Dataset
This dataset contains questions and answers related to genetics.
Dataset Information
Number of samples: 35523
Columns: text
Pre-Chunked for easier use in Teaching RAG and Retrieval
Usage
This is a corpora without questions-answers, but only contexts.
Sources
This dataset is sourced together from multiple previous sources, which due to my poor management can't be cited.
If you can find the correct source, I'd really appreciate a… See the full description on the dataset page: https://huggingface.co/datasets/nirantk/geneticsQA-corpus.mmlu-medical-geneticsmmlu_medical_geneticsgenetics-qaplexbench-base
Dataset Card for Maize and Arabidopsis gene expression
Plant Gene expression data used for benchmarking sequence to gene expression prediction ML models.
Dataset Description
Species included are Maize and Arabidopsis thaliana. Dataset includes gene expression values for leaf and root tissues.
Within the tasks folder, datasets are broken down by species-task-tissue. Genomes in the genomes folders include the annotation and the GFF files associated with that specific… See the full description on the dataset page: https://huggingface.co/datasets/maize-genetics/plexbench-base.GeneticsLlama2Traingenetics-arxiv-wiki
Dataset Card for Dataset Name
Small genetics-related text dataset based on 23200 ArXiv abstact records and 111 Wikipedia pages.
Dataset Details
Dataset Description
Dataset was produced using the python scripts you will find in this GitHub repository.
It represents a collection of genetics-related text data taken from ArXiv abstracts dataset and Wikipedia.
Dataset holds a total of 23311 text records, 23200 of which belonging to categories q-bio.BM, q-bio.GN… See the full description on the dataset page: https://huggingface.co/datasets/as-cle-bert/genetics-arxiv-wiki.genetics-intervention-leap-integrity-v0.1What this dataset names
The jump from associationto action.
What it protects
PatientsCliniciansInstitutions
Why teams want this named
This is wherelegal exposure concentrates.
mmlu-medical-geneticsgeneticsQA-trainCMMLU-Genetics-Benchmark
💻 Dataset Usage
Run the following command to load the testing set (176 examples):
from datasets import load_dataset
dataset = load_dataset("shuyuej/CMMLU-Genetics-Benchmark", split="train")
print(dataset)
French-MMLU-Medical-Genetics-Benchmark
💻 Dataset Usage
Run the following command to load the testing set:
from datasets import load_dataset
dataset = load_dataset("shuyuej/French-MMLU-Medical-Genetics-Benchmark", split="test")
print(dataset)
mmlu-medical_genetics-rule-neg
Dataset Card for "mmlu-medical_genetics-rule-neg"
More Information needed
genetics-variant-evidence-integrity-v0.1What this dataset tests
Variant × evidence × claim.
Whether a genetic variant claim is grounded in real evidence.
Common error modes
pathogenic label without functional evidence
GWAS association treated as causality
in silico prediction treated as proof
Required outputs
evidence binding map
evidence gaps
claim overreach flags
confidence calibration
Suggested prompt wrapper
System
You bind variant claims to evidence.
You calibrate confidence to evidence strength.
User… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/genetics-variant-evidence-integrity-v0.1.mmlu-medical_genetics-original-neg
Dataset Card for "mmlu-medical_genetics-original-neg"
More Information needed
mmlu-medical_genetics-neg
Dataset Card for "mmlu-medical_genetics-neg"
More Information needed
mmlu-medical_genetics-neg-prepend
Dataset Card for "mmlu-medical_genetics-neg-prepend"
More Information needed
mmlu-medical_genetics-original-neg-prepend
Dataset Card for "mmlu-medical_genetics-original-neg-prepend"
More Information needed
mmlu-medical_genetics-dev
Dataset Card for "mmlu-medical_genetics-dev"
More Information needed
mmlu-medical_genetics-neg-prepend-fix
Dataset Card for "mmlu-medical_genetics-neg-prepend-fix"
More Information needed
mmlu-medical_genetics-neg-prepend-verbal
Dataset Card for "mmlu-medical_genetics-neg-prepend-verbal"
More Information needed
genetics-mechanism-attribution-integrity-v0.1What this dataset tests
Whether a mechanistic explanationis justified by biological evidence.
Common failure modes
expression change treated as causation
nearest gene storytelling
pathway enrichment used as proof
PRS upgraded to mechanism
Required outputs
proposed mechanism
evidence supporting mechanism
mechanism gaps
attribution overreach flags
confidence in mechanism
Why this completes the trinity
Variant–Evidence IntegrityIs the claim real… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/genetics-mechanism-attribution-integrity-v0.1.mmlu-medical_genetics-neg-answer
Dataset Card for "mmlu-medical_genetics-neg-answer"
More Information needed
mmlu-medical_genetics-rule-neg-prepend
Dataset Card for "mmlu-medical_genetics-rule-neg-prepend"
More Information needed
