datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cyberseceval3-visual-prompt-injection
Dataset Card for CyberSecEval 3 - Visual Prompt Injection Benchmark
Dataset Details
Dataset Description
This dataset provides a multimodal benchmark for visual prompt injection, with text/image inputs. It is part of CyberSecEval 3, the third edition of Meta's flagship suite of security benchmarks for LLMs to measure cybersecurity risks and capabilities across multiple domains.
Language(s): English
License: MIT
Dataset Sources
Repository: Link… See the full description on the dataset page: https://huggingface.co/datasets/facebook/cyberseceval3-visual-prompt-injection.IBCBench
IBCBench: Image Bundle Composition Benchmark
IBCBench is the benchmark introduced in Weaving Visual Narratives: Agentic Image Bundle Composition Beyond Atomic Visual Matching, accepted to the EMNLP 2026 Main Conference.
Paper | GitHub
Overview
Image Bundle Composition (IBC) shifts image retrieval from independently ranking images to dynamically composing a compact, cohesive bundle whose images jointly satisfy relational, temporal, spatial, or narrative… See the full description on the dataset page: https://huggingface.co/datasets/CyberDancer/IBCBench.CyberClear
CyberClear:
CyberClear evaluates LLM agents on APT attack-chain provenance from security logs. This repository contains 20 preview instances, not the full benchmark or the complete evaluation set used in the paper.
The preview contains 10 single-step and 10 multi-step instances, spanning 1-9 attack steps (18 easy and 2 hard instances). Selection prioritizes coverage of chain lengths; this subset is not proportionally representative of the full benchmark and should not be used to… See the full description on the dataset page: https://huggingface.co/datasets/CyberClear/CyberClear.
