datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
precision-evidence-bench
Precision Evidence Bench
Precision Evidence Bench is a Precision Medicine Benchmark from
Atropos Health, the world's largest creator of
real-world evidence (RWE) for clinical decision support. It evaluates how well
large language models (LLMs) answer clinical questions that are grounded in
patient context and inclusive of patient history: not "which treatment is
better in general", but "which treatment is better for this patient", with a
specific comorbidity, age, prior therapy… See the full description on the dataset page: https://huggingface.co/datasets/atroposhealth/precision-evidence-bench.hedgehog-precision-repair
hedgehog-precision-repair
Hedgehog — precision-repair round (complete merchant extraction).
Contents
train.jsonl (1180 rows)
validation.jsonl (116 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Original content for the Hedgehog extraction model (Michael Anthony Falabella).
gsm8k-question-precision-levels
GSM8K Question Precision Levels
198 GSM8K problems, each stated at three levels of question precision, all
sharing one correct answer.
Field
Meaning
level_2
Original GSM8K question, unchanged
level_1
level_2 with exactly one span removed from the question sentence
level_0
A short model-written rewrite of the task
answer
Numeric ground truth, the same for all three levels
solution
Original GSM8K worked solution
removal_type
Which kind of removal produced… See the full description on the dataset page: https://huggingface.co/datasets/busycaesar/gsm8k-question-precision-levels.
