CoolFace
Apppublic

agent-memory-leaderboard/leaderboard

sourceHugging Faceupdated 1mo agoView on Hugging Face
790likes
App README

Agent Memory Leaderboard · 记忆之巅

A unified, open, and reproducible evaluation platform for long-term memory systems and memory-enabled agents.

Agent Memory Leaderboard (AML) compares research methods and commercial products under one evaluation contract. Candidate systems implement memory Add and Search; the official platform fixes Answer, Eval, datasets, models, configurations, result review, and publication.

First public release: The inaugural verified leaderboard is expected to be published on August 12, 2026. 首期发布: 首期经核验榜单预计将于 2026 年 8 月 12 日发布。

Evaluation structure

Results are separated along two independent dimensions. Textual and coding tasks use different metrics, while academic methods and commercial products are published in separate divisions.

Evaluation typeWhat it evaluatesPrimary ranking signal
Textual MemoryLong-horizon recall, composition, time, governance, personalization, execution, safety, and privacyOverall score across the fixed textual suite
Coding MemoryRetrieval and reuse of historical debugging and development experienceTask Solve (%)

Verified rows will bind each result to a fixed source, product, image, commit, or API version. No placeholder systems or unverified scores are published before release.

Open evaluation release

The public AML GitHub repository exposes per-benchmark evaluation contracts, shared runtime configuration, and documentation so that reported leaderboard releases can be inspected and reviewed.

To protect benchmark integrity and participant privacy, the repository deliberately excludes benchmark corpora, held-out questions, gold answers, private annotations, participant traces, production infrastructure, and credentials. Every public leaderboard row remains tied to a named method or product version and its complete evaluation contract.

Official links