notmehul/memory-bench
memory-bench v1.0.0, public release A screened benchmark for organizational memory in agent harnesses: does a memory system keep a rule that was stated once, drop a fact that was superseded, and pick the right one when tiers conflict? 371 valid paired probe instances across 3 simulated organizations, drawn from 486 probes over 612 events. Scored as pair credit: an instance counts only if the base task and its counterfactual twin both pass, so anything answerable from priors… See the full description on the dataset page: https://huggingface.co/datasets/notmehul/memory-bench.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face