datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
douvras-scientific-ci-evidence-graph
Douvras Scientific CI Evidence Graph v0.1
Synthetic protocol dataset for linking a claim to its paper, repository,
dataset, seed and reproduced metric. It contains 30 records from six toy paper
instances (20 train, 5 validation and 5 frozen test), split by paper_id.
The labels distinguish REPRODUCED, PARTIAL, FAILED and INCONCLUSIVE.
Shortcuts and leakage fail closed. No real paper, code, dataset or result is
included, and this release is not a reproduction benchmark.
clinical-evidence-dependency-graph-reasoning-v0.1
Clinical Multi-Evidence State Integration v0.1
Overview
Clinical Multi-Evidence State Integration v0.1 is a structured clinical-reasoning benchmark designed to test whether an AI system can integrate multiple sequential evidence events into a coherent final clinical state.
The benchmark evaluates more than final-answer classification.
A system must determine:
how each evidence event affects each tracked clinical item;
whether an item should be confirmed… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-evidence-dependency-graph-reasoning-v0.1.
