CoolFace
Datasetpublic

haydn-jones/labbench2-fixed

LABBench2 PMID-enriched public mirror This is a public, schema-compatible mirror of EdisonScientific/labbench2, pinned to upstream revision 27d12d72af24e3f70db8a99df63e567366cbdb80. Original columns and source URLs are unchanged. Two columns are added to every configuration: pmids: deduplicated PubMed identifiers resolved for the row's sources. source_pmids: aligned one-to-one with sources; unresolved or non-PubMed sources are null. LABBench2 LABBench2 is a… See the full description on the dataset page: https://huggingface.co/datasets/haydn-jones/labbench2-fixed.

sourceHugging Facecc-by-sa-4.0updated 1mo agoView on Hugging Face
0likes3.1kdownloads
Dataset Card

LABBench2 PMID-enriched public mirror

This is a public, schema-compatible mirror of `EdisonScientific/labbench2`, pinned to upstream revision 27d12d72af24e3f70db8a99df63e567366cbdb80. Original columns and source URLs are unchanged.

Two columns are added to every configuration:

  • —pmids: deduplicated PubMed identifiers resolved for the row's sources.
  • —source_pmids: aligned one-to-one with sources; unresolved or non-PubMed sources are null.

![arXiv](https://arxiv.org/abs/2501.XXXXX)

LABBench2

LABBench2``` is a benchmark for measuring real-world capabilities of AI systems performing scientific research tasks. It is an evolution of the [Language Agent Biology Benchmark (LAB-Bench)](https://arxiv.org/abs/2407.10362), comprising nearly 1,900 tasks that measure similar capabilities but in more realistic contexts.

This repository contains the dataset of benchmark tasks. We also provide a public evaluation harness for running any model or agent against the benchmark, which is available on GitHub.


Changelog

Notable changes to ``LABBench2`` will be documented here. We expect to update the datset only in the case of clear issues, and do not intend to meangingfully change the benchmark over time.

2026-03-13 - We corrected an inadvertent data issue with sourcequality tasks. This has resulted in an entirely new set of 150 tasks being incorporated into the dataset. Published results have been updated accordingly.