datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
precision-evidence-bench
Precision Evidence Bench
Precision Evidence Bench is a Precision Medicine Benchmark from
Atropos Health, the world's largest creator of
real-world evidence (RWE) for clinical decision support. It evaluates how well
large language models (LLMs) answer clinical questions that are grounded in
patient context and inclusive of patient history: not "which treatment is
better in general", but "which treatment is better for this patient", with a
specific comorbidity, age, prior therapy… See the full description on the dataset page: https://huggingface.co/datasets/atroposhealth/precision-evidence-bench.llm-stop-point-precision
license: cc-by-4.0
language:
en
---Title
Stop-Point Precision Evaluation for LLMs
Summary
This dataset captures stop-point precision failures in large language models. It focuses on cases where a response is correct but should have terminated earlier, violating explicit constraints such as word count, sentence count, or binary-only answers.
What this dataset tests
• Whether a model knows when to stop
• Adherence to explicit response boundaries
• Overcompletion after correct answers
Why this… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/llm-stop-point-precision.precision_agriculture_soil_health_monitoringprecision_medicine_genomic_drug_responseprecision_agriculture_drone_hyperspectralprecision_oncology_synthetic_trialsprecision_oncology_genomic_variant_analysis
