herb
Datasets
All datasets matching “herb”HERBench
HERBench: A Benchmark for Multi-Evidence Integration in Video Question Answering
A challenging benchmark for evaluating multi-evidence integration capabilities of vision-language models
🎉 HERBench has been accepted to CVPR 2026!
🆕 New: Lite-v2 config. We released a refreshed lite_v2 version of the
Lite split (1,971 questions / 68 videos) in which 9 of the 12 tasks were
regenerated and went through additional manual refinement for higher
quality, while TSO, SVA… See the full description on the dataset page: https://huggingface.co/datasets/DanBenAmi/HERBench.herb-ai-vaultherbHERB
Dataset Card for HERB
Dataset Description
HERB is a benchmark for evaluating LLM agents’ ability to perform Deep Search and Long Context Reasoning. It is generated using a synthetic data pipeline that simulates business workflows across product planning, development, and support stages, generating interconnected content with realistic noise and multi-hop questions with guaranteed ground-truth answers.
Directory Structure
data/
├── metadata/
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/HERB.Herbarium_FieldCoVUBench
Erase Persona, Forget Lore: Benchmarking Multimodal Copyright Unlearning in Large Vision Language Models
Abstract
Large Vision-Language Models (LVLMs), trained on web-scale data, risk memorizing and regenerating copyrighted visual content like characters and logos, creating significant challenges. Machine unlearning offers a path to mitigate these risks by removing specific content post-training, but evaluating its effectiveness, especially in the complex multimodal… See the full description on the dataset page: https://huggingface.co/datasets/herbwood27/CoVUBench.
