datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CAA
Clinical Agent Annotator (CAA)
A clinician-in-the-loop benchmark for long-horizon medical LLM agents:
333 KG-grounded clinical tasks, multi-gate evaluation across diagnosis,
required tool use, parameterised actions, and must-ask history-taking
topics, plus the full harbor evaluation harness so you can re-run
every number locally.
This repository bundles three artifacts:
Task corpus — 333 clinician-approved tasks (and the 290-task
authoring set the case study trains on).… See the full description on the dataset page: https://huggingface.co/datasets/anon-caa-neurips/CAA.CA-AIN-V2
Nepali Source-Grounded Instruction Dataset
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/CA-AIN-V2.ca-adyayamhatertonevople_e_uvaCaa3RqZDMC
