datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GA_Long_Context_Jailbreak_Benchmark
GA Long Context Bench
A benchmark of 1500 multi-turn conversations designed to stress guardrails in long contexts. Each dialog pairs a serialized agent trace with optional prompt injections or per-policy adjudications. Half of the rows embed malicious content deep inside long instructions, enabling evaluation of long-context systems.
Accompanying guardrail releases: GA Guard Core and GA Guard Lite. Check out public benchmarks and results in our blogpost.
[!Note]
Disclaimer: This… See the full description on the dataset page: https://huggingface.co/datasets/GeneralAnalysis/GA_Long_Context_Jailbreak_Benchmark.long-context-hindi20w_examples_LongContextfinal_samples_for_labeling_LongContext
