datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BiomniBench-DA
BiomniBench-DA
BiomniBench-DA is the data-analysis instantiation of BiomniBench, a process-level evaluation framework for LLM agents on real-world biomedical research tasks. Each task is a multi-step data analysis derived from a high-impact biomedical publication; agents are graded on the full analytical trajectory against an expert-authored rubric, not only the final answer.
This repository releases 50 of the 100 BiomniBench-DA tasks; the remaining 50 are held out as a private… See the full description on the dataset page: https://huggingface.co/datasets/phylobio/BiomniBench-DA.BiomniBench-DA-sample
BiomniBench-DA-sample
A small representative sample of BiomniBench-DA, intended for reviewer inspection of dataset quality, structure, and per-task contents.
Full dataset: phylobio/BiomniBench-DA (50 released tasks; 50 additional held-out as a private contamination-resistant evaluation set).
How this sample was created
We selected 3 tasks from the 50-task release, one from each of three different disease areas, prioritizing small footprint so the sample can be… See the full description on the dataset page: https://huggingface.co/datasets/phylobio/BiomniBench-DA-sample.
