ai-memory
MemoryAgentBench
🚧 Update
(Sep 29th, 2025) We updated our paper, where we removed some in-efficient and high-cost samples. We also added a sub-sample of DetectiveQA.
(July 7th, 2025) We released the initial version of our datasets.
(July 22nd, 2025) We modify the datasets slightly, adding the keypoints in LRU and change the uuid into qa_pair_ids. The question_ids is only used in Longmemeval task.
(July 26th, 2025) We fixed bug on qa_pair_ids.
(Aug.5th, 2025) We removed the… See the full description on the dataset page: https://huggingface.co/datasets/ai-hyz/MemoryAgentBench.MemoryArena-product-dbdaily-paper-2026-07-30-agent-memory-tiering-recall-cost
Memory Tiering Policies for Long-Running Autonomous Agent Harnesses: Mapping the Recall-Cost Frontier
TL;DR — On a production agent-memory corpus, semantic deduplication (token-Jaccard merging before ranking) outperforms both recency and frequency ordering by +17.4% AUC, reaching full recall coverage at 70% of the corpus budget. The shipped per-section item-count cap, not the character cap, is the binding constraint.
ThakiCloud AI Research · 2026-07-30 · 📝 Tech blog (KO)… See the full description on the dataset page: https://huggingface.co/datasets/thaki-AI/daily-paper-2026-07-30-agent-memory-tiering-recall-cost.lottojarvis-project-memorypublic_ai_memory_slice
Public AI Memory Slice
A scientific-domain benchmark for evaluating LLM agent memory systems on the AI / agent-memory research literature.
103 structured paper notes (~448K tokens) covering LLM agent memory, memory benchmarks, and adjacent cognitive-architecture / theory-formation work
81 full-text paper mirrors (~1.47M tokens), OCR extracted from open-access arXiv PDFs
66 main queries + 10 holdout queries with rubric-style ground truth, every must-have fact traceable to a verbatim… See the full description on the dataset page: https://huggingface.co/datasets/quantellence/public_ai_memory_slice.survival-ai-memory
