memory-benchmark
multisource-memory-benchmark
Multi-Source Memory Benchmark
Status — anonymous artefact for double-blind review (NeurIPS 2026 Evaluations & Datasets Track).
Author identities, organisations, and funders are intentionally withheld until the review period concludes.
A diagnostic testbed for selective question-answering (ANSWER / SKIP) over conflicting multi-source personal memory.
Each persona has five evidence streams projected from a single latent event table with known, controlled per-source distortions… See the full description on the dataset page: https://huggingface.co/datasets/anon-neuripsed26/multisource-memory-benchmark.mnemosyne-memory-lifecycle-benchmark
Mnemosyne Memory Lifecycle Benchmark Dataset v0.2.0
This deterministic synthetic benchmark contains 40 scenarios across eight
categories and five documented reference strategies. Strategy execution is
blind to expected labels; execution and evaluation are separate phases.
Current results are fixture results, not real-world memory accuracy. Simple
strategies remain competitive on straightforward current-state questions, while
typed lineage uniquely reconstructs the historical… See the full description on the dataset page: https://huggingface.co/datasets/solsticestudioai/mnemosyne-memory-lifecycle-benchmark.audrey-memory-benchmark-artifacts
Audrey Memory Benchmark Artifacts
This dataset publishes current Audrey benchmark artifacts for inspection and reproducibility.
Project: https://github.com/Evilander/Audrey
Package: https://www.npmjs.com/package/audrey
Live report Space: https://huggingface.co/spaces/Evilander/audrey-memory-benchmark-report
Current MemoryGym rerun after Audrey tag-isolation fix
Generated: 2026-05-01T04:38:49.849ZLocal time: April 30, 2026, 11:38 PM CentralSuite: MemoryGym… See the full description on the dataset page: https://huggingface.co/datasets/Evilander/audrey-memory-benchmark-artifacts.agent-memory-benchmark
Agent Memory Compression & Evaluation Benchmark
This dataset is a controlled evaluation testbed designed to benchmark long-term memory architectures for conversational AI agents. It stress-tests how agents handle long conversations with complex fact dynamics.
Dataset Structure
1. conversation.json
A 100-turn synthetic conversation (50 user, 50 assistant turns) containing embedded facts categorized under:
Simple Facts: Baseline retrieval details.… See the full description on the dataset page: https://huggingface.co/datasets/kushalicious/agent-memory-benchmark.braincore-memory-benchmark
BrainCore Memory Benchmark (MVP)
A lightweight, extensible harness for evaluating long-term memory systems in LLM-based agents.Status: MVP — not a SOTA leaderboard. Intended for rapid iteration and community extension.
Goals
Retrieval accuracy — Can the memory system recall the right fact when queried?
Temporal consistency — Does it respect the order and timing of events?
Contradiction handling — Can it resolve or flag updated / retracted facts?
Cost & latency — How… See the full description on the dataset page: https://huggingface.co/datasets/trentdoney/braincore-memory-benchmark.contextual_memory_benchmark_results
