CoolFace
5 results

carebench

MCG-NJU /CaReBench CaReBench: A Fine-grained Benchmark for Video Captioning and Retrieval Yifan Xu, Xinhao Li, Yichun Yang, Desen Meng, Rui Huang, Limin Wang 🤗 Model    |    🤗 Data   |    📑 Paper    📝 Introduction 🌟 CaReBench is a fine-grained benchmark comprising 1,000 high-quality videos with detailed human-annotated captions, including manually separated spatial and temporal descriptions for independent spatiotemporal bias evaluation. 📊 ReBias… See the full description on the dataset page: https://huggingface.co/datasets/MCG-NJU/CaReBench.text1K<n<10K0 likes233 downloads1y agoHugging Faceningkko /CARE-Bench CARE-Bench v1.1 public data GitHub link to the project The public release contains 439 cases and 925 labeled prefixes. Split Cases Prefixes Development 284 609 Validation 93 181 Public Test 1 62 135 The prefix labels are distributed as follows: Label Prefixes Information needed 249 Self-care or monitor 162 Nonurgent care 342 Urgent care 172 The Dataset Viewer includes four configurations: model_inputs: identifiers, split information… See the full description on the dataset page: https://huggingface.co/datasets/ningkko/CARE-Bench.text1K<n<10K0 likes129 downloads2mo agoHugging FaceEliasHossain /carebench CareBench CareBench is a benchmark for process-aware evaluation of clinical LLM agents. Each case places an agent in a simulated patient encounter in which it must gather evidence through interaction, commit to a diagnosis, and propose management, while safety-critical red flags are tracked throughout the episode. Scoring is process-aware: in addition to diagnostic accuracy, the benchmark measures evidence coverage, premature closure, red-flag coverage, and unsafe commitments.… See the full description on the dataset page: https://huggingface.co/datasets/EliasHossain/carebench.texttext-generationn<1K0 likes60 downloads2mo agoHugging Facehandshake-ai-research /CAREBenchgated CAREBench CAREBench (Child AI Risk Evaluation) is a benchmark of 500 single-turn prompts for evaluating whether language models recognize and respond appropriately to upstream child-safety risks. Each prompt targets a real-world risk mechanism — grooming, sextortion, social manipulation, emotional dependency, AI anthropomorphization, therapist replacement, and more. LLM responses to these prompts are automatically graded via a LLM judge methodology calibrated against expert- and… See the full description on the dataset page: https://huggingface.co/datasets/handshake-ai-research/CAREBench.text-generation10K<n<100K3 likes43 downloads3mo agoHugging Facemyang333 /CaReBenchtext1K<n<10K0 likes24 downloads2mo agoHugging Face