datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LabCraft-Eval
LabCraft-Eval
LabCraft-Eval is an Inspect AI evaluation environment for measuring how well AI
agents execute benign molecular-microbiology protocols inside a seeded
laboratory simulator with task-dependent stochasticity. It pairs task prompts
and tool-accessible lab operations with deterministic, multi-axis trajectory
scoring.
This Hugging Face dataset export is generated from the GitHub repository:
https://github.com/jang1563/LabCraft-Eval.git
Release
Release… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/LabCraft-Eval.sci-agent-verification-cascade
Scientific Agent Verification Cascade
Public evaluation fixtures and verified aggregate results for testing whether
scientific claims keep their source, meaning, uncertainty, and verification
requirements as they move between AI agents.
This dataset accompanies the
Scientific Agent Verification Cascade
codebase. Version 0.2.0
contains synthetic evaluation data and aggregate-only results. It contains no
raw hosted-model response, private holdout identifier,
source-record… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/sci-agent-verification-cascade.BioEval
BioEval
BioEval is an open-ended benchmark for evaluating biological reasoning in large
language models. Release v0.7.1 contains 12 components and two
cumulative task-set configurations:
Configuration
Split
Rows
Meaning
base
test
296
Canonical base benchmark
extended
test
400
The identical 296 base records plus 104 extended records
Configurations represent benchmark tiers, not train/test partitions. The 296
task IDs shared by base and extended have… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/BioEval.
