datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CAA
Clinical Agent Annotator (CAA)
A clinician-in-the-loop benchmark for long-horizon medical LLM agents:
333 KG-grounded clinical tasks, multi-gate evaluation across diagnosis,
required tool use, parameterised actions, and must-ask history-taking
topics, plus the full harbor evaluation harness so you can re-run
every number locally.
This repository bundles three artifacts:
Task corpus — 333 clinician-approved tasks (and the 290-task
authoring set the case study trains on).… See the full description on the dataset page: https://huggingface.co/datasets/anon-caa-neurips/CAA.my-distiset-caa59c52
Dataset Card for my-distiset-caa59c52
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/zastixx/my-distiset-caa59c52/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/zastixx/my-distiset-caa59c52.
