datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
history-anchor-100
History Anchor 100
*The benchmark behind the paper "History Anchors: How Prior Behavior Steers LLM Decisions Toward Unsafe Actions".*
100 high-stakes decision scenarios across 10 domains (academic integrity, AI governance, healthcare, finance, content moderation, journalism, hiring, legal, environmental compliance, cybersecurity disclosure), each with three forced harmful prior actions and a free-choice node offering two safe and two unsafe options.
Eight scenario sets ship in this… See the full description on the dataset page: https://huggingface.co/datasets/albertoRodriguez97/history-anchor-100.Anchor-benchmarks
🧠 Anchor Benchmarks
A curated long-term memory benchmark bundle for LLM and agent evaluation
Anchor Benchmarks packages three public long-term memory evaluation resources for studying factual recall, temporal reasoning, knowledge update, multi-hop inference, and multimodal conversational memory.
Quick Start ·
At a Glance ·
Benchmarks ·
Evaluation ·
Citation
[!IMPORTANT]
This repository is a benchmark bundle, not a new… See the full description on the dataset page: https://huggingface.co/datasets/zhangdw/Anchor-benchmarks.
