Danchi17/agent-memory-integrity-leaderboard
0
Agent-Memory Integrity Leaderboard
Recall leaderboards measure whether an agent's memory returns the right fact. None measure whether a correction sticks. This one does — a standing, open, cross-system leaderboard on the integrity axis:
- Value-obscuring revert — undo a correction from an unmarked "go back". A capability gap: only a system with an explicit revert channel can do it (inspeximus 0.75, mem0 0.20, Graphiti 0.00; CIs disjoint).
- Echo resurrection — restate the retired value and see if it comes back. A tie — all three defend. We lead with the cell we tie, because a leaderboard you only ever top is one you built to flatter yourself.
One shared, ground-truth-blind judge (gpt-4o-mini) reads each system's own recall surface — no home-field instrument. The harness is open; add your system by PR and it appears here.
- Benchmark: RAMR · DOI 10.5281/zenodo.20818291
- Method: integrity/METHODOLOGY.md
- Submit your system: integrity/SUBMISSION.md
- Machine-readable: `leaderboard.json`
MIT licensed. Numbers are verified against RAMR's canonical cross-system run before publishing.
