datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fragbench
FragBench (Public Tier)
Anonymous submission for NeurIPS 2026 Datasets and Benchmarks Track.
Author identity will be revealed at camera-ready.
Dataset Summary
FragBench is a benchmark for evaluating cross-session, fragmented attacks on
LLM agents that use tools via the Model Context Protocol (MCP). Each campaign
is decomposed into many small fragments distributed across sessions; a
defender must reconstruct the compositional intent. The public tier in this
repository… See the full description on the dataset page: https://huggingface.co/datasets/LidaSafety/fragbench.OncoBench
OncoBench
OncoBench is an oncology decision-reasoning benchmark for evaluating large language models and agentic systems on treatment recommendations, safety violations, risk recognition, missing-information handling, and abstention behavior.
This repository contains two benchmark subsets:
data/benchmark1000/benchmark1000_weak_labels.jsonl: Benchmark1000 weak-labeled oncology split for broader development and screening. This packaged file contains 998 JSONL records.… See the full description on the dataset page: https://huggingface.co/datasets/liderion/OncoBench.
