datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CaReBench
CaReBench: A Fine-grained Benchmark for Video Captioning and Retrieval
Yifan Xu, Xinhao Li, Yichun Yang, Desen Meng, Rui Huang, Limin Wang
🤗 Model | 🤗 Data | 📑 Paper
📝 Introduction
🌟 CaReBench is a fine-grained benchmark comprising 1,000 high-quality videos with detailed human-annotated captions, including manually separated spatial and temporal descriptions for independent spatiotemporal bias evaluation.
📊 ReBias… See the full description on the dataset page: https://huggingface.co/datasets/MCG-NJU/CaReBench.CARE-Bench
CARE-Bench v1.1 public data
GitHub link to the project
The public release contains 439 cases and 925 labeled prefixes.
Split
Cases
Prefixes
Development
284
609
Validation
93
181
Public Test 1
62
135
The prefix labels are distributed as follows:
Label
Prefixes
Information needed
249
Self-care or monitor
162
Nonurgent care
342
Urgent care
172
The Dataset Viewer includes four configurations:
model_inputs: identifiers, split information… See the full description on the dataset page: https://huggingface.co/datasets/ningkko/CARE-Bench.carebench
CareBench
CareBench is a benchmark for process-aware evaluation of clinical LLM agents. Each case places an agent in a simulated patient encounter in which it must gather evidence through interaction, commit to a diagnosis, and propose management, while safety-critical red flags are tracked throughout the episode. Scoring is process-aware: in addition to diagnostic accuracy, the benchmark measures evidence coverage, premature closure, red-flag coverage, and unsafe commitments.… See the full description on the dataset page: https://huggingface.co/datasets/EliasHossain/carebench.CAREBench
CAREBench
CAREBench (Child AI Risk Evaluation) is a benchmark of 500 single-turn prompts for evaluating whether language models recognize and respond appropriately to upstream child-safety risks. Each prompt targets a real-world risk mechanism — grooming, sextortion, social manipulation, emotional dependency, AI anthropomorphization, therapist replacement, and more. LLM responses to these prompts are automatically graded via a LLM judge methodology calibrated against expert- and… See the full description on the dataset page: https://huggingface.co/datasets/handshake-ai-research/CAREBench.CaReBench
